Neural network for simulating data center hardware

A liquid cooling system with an absorption chiller and modular neural networks optimizes data center cooling, addressing inefficiencies and labor-intensive manual processes, enhancing efficiency and reducing environmental impact.

DE102025147255A1Pending Publication Date: 2026-05-21NVIDIA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
NVIDIA CORP
Filing Date
2025-11-14
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

The manual process of designing and testing data center cooling systems is time-consuming, labor-intensive, and can cause damage to cooling hardware, while existing cooling methods are inefficient for varying computational loads and heat demands.

Method used

A data center cooling system utilizing a liquid cooling system with an absorption chiller and modular neural networks to simulate and optimize cooling hardware, enabling real-time monitoring and control without physical testing.

Benefits of technology

Facilitates rapid design adjustments and efficient real-time monitoring, reducing costs and labor, while improving data center efficiency and reducing environmental impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The embodiments described herein provide one or more neural networks for simulating a data center cooling system. In at least one embodiment, the one or more neural networks can each be trained to simulate one or more types of data center hardware and can then be integrated to jointly simulate a data center cooling system with the one or more types of data center hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Area

[0001] This application relates to a machine learning system or machine learning systems for simulating data center hardware in a data center cooling system according to various embodiments described herein. background

[0002] Today's data center cooling systems are designed with significant manual effort. For example, engineers often need to select, assemble, and test various cooling hardware components, such as cooling distribution units (CDUs), racks, valves, filters, and the like, to determine how a cooling system should be designed and built. Engineers might select specific makes and models of different cooling hardware and physically install it in a test cooling system, which can then be run to gather various performance data, such as server temperature, coolant flow rate, and so on. This performance data can be collected by various sensors distributed throughout the test cooling system.Based on the collected test cooling performance data, engineers can then assess whether the selected cooling hardware and the individual design of the test cooling system can meet expectations. This entire manual process of installing and testing a test cooling system and cooling hardware is both time-consuming and labor-intensive, and can sometimes even cause damage to the cooling hardware during testing. Brief description of the drawings Fig. Figure 1 illustrates an exemplary data center cooling system according to an embodiment described herein; Fig. 2A-2D illustrate exemplary block diagrams that depict modular neural networks trained to simulate various data center hardware to collectively perform a task in Fig. 1 to simulate the data center system described herein, according to the embodiments described herein. Fig. Figures 3A-3B illustrate exemplary data transmission between neural module networks corresponding to different types of data center modules, to simulate the process described in Fig. 1 to enable the data center system described herein in accordance with the embodiments described herein. Fig. Figures 4A-4B illustrate exemplary control modules corresponding to modular neural networks corresponding to different types of data center modules, according to embodiments described herein. Fig. Figure 5 illustrates an exemplary logic flow diagram of a method for using one or more modular neural networks to jointly simulate a data center, according to at least one embodiment. Fig. Figure 6A is a simplified diagram illustrating a computing device implementing a digitally simulated data center cooling system located in Fig. 1-5 is described, according to an embodiment described herein. Fig. 6B illustrates a distributed system according to at least one embodiment; Fig. Figure 7 illustrates an exemplary data center according to at least one embodiment; Fig. Figure 8 illustrates a client-server network according to at least one embodiment; Fig. Figure 9 illustrates a computer network according to at least one embodiment; Fig. 10A illustrates a networked computer system according to at least one embodiment; Fig. 10B illustrates a networked computer system according to at least one embodiment; Fig. 10C illustrates a networked computer system according to at least one embodiment; Fig. Figure 11 illustrates one or more components of a system environment in which services can be offered as third-party network services, according to at least one embodiment. Fig. Figure 12 illustrates a cloud computing environment according to at least one embodiment. Fig. Figure 13 illustrates a set of functional abstraction layers provided by a cloud computing environment, according to at least one embodiment. Fig. 14 illustrates a chip-level supercomputer according to at least one embodiment; Fig. Figure 15 illustrates a rack module-level supercomputer according to at least one embodiment; Fig. Figure 16 illustrates a rack-level supercomputer according to at least one embodiment; Fig. 17 illustrates a supercomputer at the overall system level according to at least one embodiment; Fig. 18A illustrates inference and / or training logic according to at least one embodiment. Fig. 18B illustrates inference and / or training logic according to at least one embodiment. Fig. 19 illustrates the training and use of a neural network according to at least one embodiment. Fig. 20 illustrates an architecture of a system of a network according to at least one embodiment; Fig. 21 illustrates an architecture of a system of a network according to at least one embodiment; Fig. Figure 22 illustrates a control plane protocol stack according to at least one embodiment. Fig. Figure 23 illustrates a user-level protocol stack according to at least one embodiment. Fig. 24 illustrates components of a core network according to at least one embodiment; and Fig. Figure 25 illustrates components of a system for supporting network function virtualization (NFV) according to at least one embodiment; Fig. Figure 26 illustrates a processing system according to at least one embodiment. Fig. 27 illustrates a computer system according to at least one embodiment; Fig. 28 illustrates a system according to at least one embodiment; Fig. Figure 29 illustrates an exemplary integrated circuit according to at least one embodiment. Fig. 30 illustrates a computer system according to at least one embodiment; Fig. 31 illustrates an APU according to at least one embodiment; Fig. 32 illustrates a CPU according to at least one embodiment; Fig. Figure 33 illustrates an exemplary accelerator integration slice according to at least one embodiment; Fig. Figures 34A-34B illustrate exemplary graphics processors according to at least one embodiment; Fig. 35A illustrates a graphics kernel according to at least one embodiment; Fig. 35B illustrates a GPGPU according to at least one embodiment; Fig. Figure 36A illustrates a parallel processor according to at least one embodiment; Fig. Figure 36B illustrates a processing cluster according to at least one embodiment; Fig. 36C illustrates a graphics multiprocessor according to at least one embodiment; Fig. Figure 37 illustrates a software stack of a programming platform according to at least one embodiment. Fig. Figure 38 illustrates a CUDA implementation from a software stack of Fig. 37 according to at least one embodiment. Fig. Figure 39 illustrates a ROCm implementation of a software stack of Fig. 37 according to at least one embodiment. Fig. Figure 40 illustrates an OpenCL implementation of a software stack of Fig. 37 according to at least one embodiment. Fig. Figure 41 illustrates software supported by a programming platform according to at least one embodiment. Fig. 42 illustrates compiling code for execution on programming platforms of Fig. 37 - 40 according to at least one embodiment. Detailed description

[0003] The following description sets out numerous specific details to provide a more thorough understanding of at least one embodiment. However, it is evident to a person skilled in the art that the inventive concepts can be implemented without one or more of these specific details.

[0004] In at least one embodiment, air cooling of high-density servers alone may be inefficient or ineffective with regard to sudden high heat demands caused by changing computational loads in modern computing components. In at least one embodiment, as the requirements are subject to change or tend to range from a minimum to a maximum of various cooling demands, these demands must be met economically using a suitable cooling system. In at least one embodiment, a liquid cooling system can be used for moderate to high cooling demands. In at least one embodiment, the varying cooling requirements also reflect different thermal characteristics of the data center.In at least one embodiment, heat generated by the components, servers and racks is cumulatively referred to as a heat feature or a cooling requirement, since the cooling requirement must fully address the heat feature.

[0005] In at least one embodiment, a data center liquid cooling system is disclosed. In at least one embodiment, the data center cooling system addresses thermal features in associated computing or data center devices or components, such as graphics processing units (GPUs), switches, dual inline memory modules (DIMMs), or central processing units (CPUs). Furthermore, in at least one embodiment, an associated computing or data center device can be a processing card with one or more GPUs, switches, or CPUs on it. In at least one embodiment, each of the GPUs, switches, and CPUs can be a heat-generating feature of the computing device. In at least one embodiment, the GPU, CPU, or switch can have one or more cores, and each core can be a heat-generating feature.

[0006] In at least one embodiment, data center components can be designed for high computing demands in artificial intelligence / machine learning (AI / ML) and other high-performance computing (HPC) applications. In at least one embodiment, these data center components can be high-heat-density components requiring reliable and economical heat removal. Since these data center components can generate significant amounts of heat, waste heat utilization, as in the present heat recovery system, offers benefits for the data center cooling system. In at least one embodiment, when used in a liquid cooling system for the high-heat-density components, the heat recovery system can improve data center efficiency and contribute to free cooling within the data center.In at least one embodiment, such waste heat utilization reduces environmental impacts from data center cooling and reduces the carbon footprint of the data center.

[0007] In at least one embodiment, an absorption chiller can be positioned and calibrated to receive heat from fluid returning from the data center. In at least one embodiment, the fluid can be a coolant. In at least one embodiment, the absorption chiller comprises a working fluid that is different from the fluid returning to and being sent from the data center. In at least one embodiment, the working fluid is a mixed solution with lithium bromide as the absorber material and water as the carrier material. In at least one embodiment, the fluid returning from the data center can be a secondary coolant of a secondary cooling circuit. In at least one embodiment, the fluid can return from a cold plate or liquid-cooled immersion server in the data center.In at least one embodiment, the fluid can be introduced directly or indirectly (such as through a heat exchanger or a burner stage) into a generator vessel of a single- or multi-stage absorption chiller.

[0008] In at least one embodiment, the absorption chiller cools the fluid by enabling a generator vessel to transfer heat from the fluid to its contents, thereby removing at least some of the heat from the fluid. In at least one embodiment, this portion of the heat is transferred to the generator vessel, which in turn transfers the heat to the contents, such as the working fluid, within the generator vessel. In at least one embodiment, this portion of the heat supplements a burner in a burner stage used with the generator vessel to achieve temperatures that allow the separation of the absorption material from the working fluid in the absorption chiller. In at least one embodiment, the data center fluid can be recirculated to the data center for cooling data center components in the cold plate or liquid-cooled immersion server.In at least one embodiment, residual heat can be removed from the fluid by a coolant distribution unit (CDU) that connects the fluid to a primary coolant of a primary cooling circuit with an external chiller device.

[0009] In at least one embodiment, cooling the fluid via the absorption chiller enables the cooling of at least a portion of the working fluid in an evaporation zone that houses an evaporation coil within the absorption chiller. In at least one embodiment, the fluid can be further cooled in the evaporation zone after the CDU. In at least one embodiment, a separate cooling circuit can include a separate fluid (other than the data center fluid or the working fluid) for other cooling functions of the data center, including cooling staff rooms, heating, ventilation, and air conditioning (HVAC) units of the data center, or areas near the data center. In at least one embodiment, the absorption chiller can provide benefits in its own right or can provide benefits as an additional cooling feature for a data center cooling system.In at least one embodiment, the absorption chiller enables the removal of heat of more than 1 kW from fluid recirculated from data center components.

[0010] In at least one embodiment, an exemplary data center infrastructure 100 can be used, as shown in Fig. Figure 1A illustrates a cooling system that incorporates the improvements described herein. In at least one embodiment, the data center infrastructure 100 can be one or more rooms 1003 comprising racks 106 and auxiliary equipment to house one or more servers on one or more server trays. In at least one embodiment, the data center infrastructure 100 is supported by a cooling tower 1005 located outside the data center infrastructure 100. In at least one embodiment, the cooling tower 1005 dissipates heat from within the data center infrastructure 100 by acting on a primary cooling circuit 1007.In at least one embodiment, one or more cooling distribution units (CDUs) 125 are used between the primary cooling circuit 1007 and a second or secondary cooling circuit 1009 to facilitate heat extraction from the second or secondary cooling circuit 1009 to the primary cooling circuit 1007. In at least one embodiment, the secondary cooling circuit 1009 can access various plumbing lines extending into the server tray as needed. In at least one embodiment, the circuits 1007 and 1009 are illustrated as line drawings, but a person skilled in the art would recognize that one or more plumbing features can be used. In at least one embodiment, flexible polyvinyl chloride (PVC) tubing, along with associated installation, can be used to move the fluid along each of the circuits 1007 and 1009.In at least one embodiment, one or more coolant pumps can be used to maintain pressure differentials within the loops 1007, 1009 in order to enable the movement of the coolant according to temperature sensors at different locations, including in the room, in one or more racks 106 and / or in server boxes or server trays within the racks 106.

[0011] In at least one embodiment, the coolant in the primary cooling circuit 1007 and in the secondary cooling circuit 1009 can be at least water and an additive, for example, glycol or propylene glycol. In operation, in at least one embodiment, each of the primary and secondary cooling circuits has its own coolant. In at least one embodiment, the coolant in the secondary cooling circuits can be proprietary to meet the requirements of the components in the server tray or racks 106. In at least one embodiment, the one or more CDUs 125 are capable of sophisticatedly controlling the coolants in the circuits 1007 and 1009 independently or simultaneously. In at least one embodiment, the CDU can be configured to control the flow rate so that the coolant(s) are adequately distributed to extract heat generated within the racks 106.In at least one embodiment, more flexible tubing material 1015 is provided by the secondary cooling circuit 1009 to enter each server tray and to provide coolant to the electrical and / or computing components.

[0012] In at least one embodiment, the electrical and / or computing components are used interchangeably to refer to the heat-generating components that benefit from the present data center cooling system. In at least one embodiment, the tubing material 1019, which forms part of the secondary cooling circuit 1009, can be referred to as room manifolds. Separately, in at least one embodiment, the tubing material 1017, which extends from the tubing material 1019, can also be part of the secondary cooling circuit 108, but can be referred to as row manifolds. In at least one embodiment, the tubing material 1015 enters the racks as part of the secondary cooling circuit 108, but can be referred to as rack cooling manifolds. In at least one embodiment, the row manifolds 1017 extend along a row in the data center infrastructure 100 to all racks.In at least one embodiment, the piping of the secondary cooling circuit 1009, including the manifolds 1019, 1017, and 1015, can be improved by at least one embodiment of the present disclosure. In at least one embodiment, a chiller 1021 can be provided in the primary cooling circuit within the data center infrastructure 100 to assist cooling upstream of the cooling tower 1005. In at least one embodiment, a person skilled in the art, to the extent that additional circuits exist in the primary control circuit, would recognize upon reading the present disclosure that the additional circuits provide cooling outside the rack and outside the secondary cooling circuit; and can be taken together with the primary cooling circuit for the purposes of this disclosure.

[0013] In at least one embodiment, heat generated within the server trays of the racks 106 during operation can be transferred via flexible tubing of the rack distributor 1015 of the second cooling circuit 1009 to a coolant exiting the racks 106. In at least one embodiment, the second coolant (in the secondary cooling circuit 1009) moves from one or more CDUs 125 to the racks 106 for cooling. In at least one embodiment, the second coolant flows from one or more CDUs 125 from one side of the room distributor via tubing 1019, through the row distributor 1017, to one side of the rack 106 and via tubing 1015 through one side of the server tray. In at least one embodiment, used or recirculated second coolant (or exiting second coolant carrying heat from the computing components) exits from another side of the server (such as...).It enters the left side of the rack and exits the right side of the rack for the server tray (after being routed through the server tray or components on the server tray). In at least one embodiment, the used second coolant exiting the server tray or rack 106 exits from another side (such as the outlet side) of the hose 1015 and moves to a parallel outlet side of the manifold 1017. In at least one embodiment, the used second coolant moves from the manifold 1017 in a parallel section of the space distributor 1019, which runs in the opposite direction to the incoming second coolant (which may also be the renewed second coolant), and to one or more CDUs 125.

[0014] In at least one embodiment, the consumed second coolant exchanges its heat with a primary coolant in the primary cooling circuit 1007 via one or more CDUs 125. In at least one embodiment, the consumed second coolant can be renewed (e.g., cooled relative to the temperature in the consumed second coolant stage) and be ready to be recirculated (cycled) through the second cooling circuit 1009 to the computing components. In at least one embodiment, various flow and temperature control features in the one or more CDUs 125 enable control of the heat exchanged by the consumed second coolant or of the flow of the second coolant into and out of the one or more CDUs 125.In at least one embodiment, the one or more CDUs 125 may also be able to control a flow of the primary coolant in the primary cooling circuit 1007.

[0015] In at least one embodiment, pressure sensors can be placed at the server level within each rack 106 along the coolant lines near the coolant inlet and / or outlet of a data server, so that a coolant inlet pressure and a coolant outlet pressure, referred to as the "differential pressure," can be measured. In at least one embodiment, the differential pressure drop between the coolant inlet and the coolant outlet can be monitored so that the flow rate for each piece of computer hardware on rack 106 can be dynamically adjusted. In at least one embodiment, the data center infrastructure 100 can include a power delivery system comprising a number of different units throughout the server rack 106.

[0016] In at least one embodiment, one or more control valves can be placed along coolant lines for the primary cooling circuit 1007 or the secondary cooling circuit 1009 in the data center infrastructure 100 to control the coolant flow in the cooling system. In at least one embodiment, for example, control valves can be installed on the line distributors 1017, rack-level distributors 1015, or at the server level for each server on rack 106. In at least one embodiment, each control valve can control the coolant flow to and from various components of a data center. In at least one embodiment, control valves can be automatically reconfigured to decrease or increase the coolant flow in response to a signal. In at least one embodiment, a signal for reconfiguring control valves can be transmitted from a control unit via a wireless or wired connection.

[0017] In at least one embodiment, one or more air-to-liquid heat exchangers can be placed and operated above racks 106 in a data center 100. In at least one embodiment, the one or more heat exchangers (such as a fan coil unit (FCU)) can have fans to draw heated airflow from the computing equipment in the server racks 106 into the FCU, whereby the heated air can in turn be cooled by the primary cooling circuit. In at least one embodiment, the FCU can then release cooled airflow that is transported back to the server rack 106. In at least one embodiment, heated air from a "hot aisle" is transferred by fans to the FCU, and air cooled by the FCU is transferred by fans to a "cold aisle," which can consist of cool server racks 106.In at least one embodiment, one or more FCUs can be placed that are associated with CDU 125.

[0018] In at least one embodiment, more than one type of cooling system can be used in the data center cooling system 100. For example, one type of cooling system is immersion cooling, which involves immersing data center components in a thermally conductive but electrically insulating liquid. As another example, a different cooling system is direct-to-chip (D2C) cooling, which involves circulating a cold coolant or refrigerant to hardware within a server through coolant lines.

[0019] In at least one embodiment, data center infrastructure 100 can support both liquid cooling (as described above) and air cooling. For example, in at least one embodiment, the data center infrastructure 100 can include various cooling equipment, such as, among others, CDUs, rack fans, and air conditioning units (ACUs), to cool data center hardware (e.g., GPUs and CPUs in servers on racks 106) collectively. In at least one embodiment, the various cooling equipment can be controlled and monitored by physically informed real-time simulations integrated with live sensor data. In at least one embodiment, the various cooling equipment can be controlled and monitored by a physically informed machine learning system, as described herein.

[0020] In at least one embodiment, designing and / or building the data center infrastructure required 100 conventionally time-consuming experimental tests. In at least one embodiment, this experimental testing of cooling hardware can be costly and requires expensive laboratory facilities and manual labor. In at least one embodiment, any change to the design of the data center infrastructure can require 100 new tests, which can cause even more manual labor and laboratory work.

[0021] Fig. 2A-2D illustrate exemplary block diagrams that depict modular neural networks trained to simulate various data center hardware in order to collectively simulate a data center system running in Fig. 1 is described, according to the embodiments described herein. In at least one embodiment, such as in Fig. As shown in Figure 2A, a digitally simulated data center cooling system can be built to replicate the operation of a physical data center cooling system similar to 100. Fig. 1 to simulate. In at least one embodiment, the digitally simulated data center cooling system can facilitate real-time monitoring of performance and control.

[0022] In at least one embodiment, the digitally simulated data center cooling system 200 can have a modular structure comprising a digitally simulated cooling tower module 205, a chiller module 206, a CDU module 225, a data center room module 230, which further comprises one or more rack modules 235, one or more server modules 236, one or more piping modules 241 and / or an airflow simulation module 233 and / or the like.

[0023] In at least one embodiment, each module can be a neural network trained to simulate the operation of a specific type of cooling hardware. For example, cooling tower module 205 can be trained to operate a physical cooling tower (e.g., 1005 in [location]). Fig. 1) to simulate. As another example, chiller module 206 can be trained to simulate a physical chiller (e.g., 1021 in Fig. 1) to simulate. As another example, CDU module 225 can be trained, a physical CDU (e.g., 125 in Fig. 1) to simulate. As another example, Rackmodule 235 can be trained, a physical server rack (e.g., 106 in Fig. 1) to simulate. As another example, Server Module 236 can be trained to simulate a physical server placed on a physical rack. As another example, Pipeline Module 241 can be trained to simulate physical pipework such as room distributors, manifolds, and / or rack distributors (e.g., 1019, 1017, and 1015 in Fig. 1) to simulate. As another example, Airflow Module 233 can be trained to simulate a hot or cold airflow along the aisles in a physical data center room. As another example, Air Cooling Module 237 can be trained to simulate the operation of a computer room air purifier (CRAH) that uses chiller water to cool a hot airflow, and / or a computer room air conditioner (CRAC) that uses a refrigerant to cool the hot airflow.

[0024] In at least one embodiment, for example, each neural network simulating a single type of cooling hardware can, upon inference, receive input data representing the characteristics of a physical cooling input to the cooling hardware, one or more control parameters of the single type of cooling hardware, and / or other data center conditions outside the specific type of cooling hardware, and, based on this, generate output data that reveals characteristics of the physical cooling output of the cooling hardware. For example, a CDU module 225 can receive an input describing characteristics of a return coolant (e.g., temperature, flow rate, and / or the like), control parameters (e.g., heat exchanger control settings, valve settings, pump settings, etc.), any external conditions outside the CDU module 225 (e.g.,Data center conditions (such as the number of racks, airflow conditions, and / or the like) are used to predict an output including characteristics of a supply coolant (e.g., flow rate, temperature, and / or the like). As another example, a rack module 235 can receive an input that includes at least characteristics of a supply coolant (from the CDU module 225), control parameters, and / or other external data center conditions to generate predicted characteristics of a return coolant from the rack, which is then returned to the CDU module 225. In at least one embodiment, where the CDU module 225 and the rack module 235 can be interconnected in a data center cooling system architecture, real-time communication between modules, as described below, is possible. Fig. 3A-3B described, facilitate the integration of different neural networks to jointly simulate data center cooling system 200.

[0025] In at least one embodiment, for example, the neural networks can be trained to jointly simulate estimated operating cost and / or performance metrics, such as a total cost of ownership (TCO) and a performance metric such as power utilization effectiveness (PUE) and / or the like, corresponding to the specific data center design. In at least one embodiment, TCO and / or performance metrics associated with the data center design can also be obtained from the simulation results 216 of the data center simulation platform 200.

[0026] Total Cost of Ownership (TCO) can refer to the comprehensive costs of operating and maintaining a data center over a period of time, including all direct and indirect costs associated with the data center, from initial capital investment to ongoing operating expenses. In at least one embodiment, TCO can include capital expenditures (CapEx) and operating expenses (OpEx). OpEx can include, for example, the costs of operating and maintaining the data center, such as, but not limited to, the costs associated with the power required to operate all cooling hardware modules in the data center design. OpEx can be updated in real time as the data center's condition changes, enabling data-driven decision-making. For example, the data center's location can influence the choice of heat dissipation systems (chillers vs. radiators).dry coolers) and their operating conditions, CDUs and ACUs, which in turn affects the overall cost of cooling.

[0027] As another example, CapEx may include costs to construct the data center simulated by the Data Center Simulation Platform 200, including the costs of the number of racks, technology cooling system, plant piping, aisle enclosure, servers, CDUs, chillers and plant costs and / or the like.

[0028] In at least one embodiment, PUE is a further metric to show the ratio of the total power delivered to the data center cooling system to the total power used by IT systems (total data center power / total IT power). For example, PUE can be calculated based on a simulation of the Data Center Simulation Platform 200 by monitoring the power consumption of the cooling systems and IT modules of the Data Center Simulation Platform 200 in real time.

[0029] In at least one embodiment, Data Center Simulation Platform 200 can additionally generate an output to assist users in selecting the heat dissipation system for information technology (IT) requirements, based at least partially on the climatic conditions of a specific location. In at least one embodiment, Data Center Simulation Platform 200 can receive suggested operating parameters (flow rate, temperature, and cooling load) and a city location to generate an output that analyzes historical extreme weather conditions, e.g., from the last 50 years. Data Center Simulation Platform 200 can then generate possible solutions with the goal of identifying the most energy-efficient scenario for data center designs. In at least one embodiment, a CapEx and OpEx comparison of different solutions can be provided, e.g.,via a user interface, which enables data center design engineers to make informed decisions for each specific design.

[0030] In at least one embodiment, performance metrics for evaluating data center design may also include tokens per kilowatt, tokens per dollar and / or the like.

[0031] In this way, Data Center Simulation Platform 200 can facilitate rapid design adjustments and efficient real-time monitoring of data center infrastructure without the costly physical manual process of design and testing. In at least one embodiment, Data Center Simulation Platform 200 can provide long-term financial implications for the construction, operation, and maintenance of a data center by supplying performance metrics such as TCO and PUE, thus facilitating decision-making during the design and construction phases. Decisions regarding data center location, changes in IT load on data center computer hardware, and / or similar factors can be made based on TCO and return-on-investment data.

[0032] In at least one embodiment, the various modules can form a digitally simulated data center cooling system 200, which together simulate the operation of the data center cooling system, which is considered a simulated version of the data center infrastructure 100 in Fig. 1. For example, the digitally simulated data center cooling system 200 can accept input of various control parameters for different modules, such as CDU settings, valve configuration settings, rack fan settings, and / or the like, each of which is then entered into the corresponding cooling hardware module. The digitally simulated data center cooling system 200 can then generate predicted thermal states, such as coolant temperature, coolant flow rate at various points in the piping, server temperature, airflow dynamics, and / or the like, based on the input control parameters.

[0033] In at least one embodiment, one or more neural networks, each simulating a single type of cooling hardware, can be selected to design and / or build a data center cooling system 200. In at least one embodiment, a user can select the desired modular neural networks or specify design requirements (for automatic module creation) to form the data center cooling system 200 via a graphical user interface (GUI). This allows for the creation of data transmission pipelines (see, for example, [reference to relevant example]). Fig. 3A-3B) between the modular neural networks to facilitate an overall simulation of the data center cooling system 200, which features the selected modular neural networks.

[0034] In at least one embodiment, to build a data center cooling system 200 from modular neural networks, a user can program and / or configure computer code programs that configure and / or integrate the modular neural networks to simulate them together as an integrated cooling system 200. In at least one embodiment, instead of the user manually selecting, designing, and / or programming the desired simulation cooling system, a neural network such as a Large Language Model (LLM) can be used to automatically design, select, and / or generate computer code programs to integrate one or more desired modular neural networks into the digitally simulated data center cooling system 200. Embodiments of using an LLM to design a digitally simulated data center cooling system are described in concurrently pending and jointly assigned U.S. Patent Application No.18 / 949,706, Attorney File No. 42940.83US01, which was filed on the same day.

[0035] In at least one embodiment, the digitally simulated data center cooling system 200 can update predicted simulation results of the data center cooling system in real time in response to any changes in the control parameters. In at least one embodiment, the digitally simulated data center cooling system 200 can be updated by removing, adding, or replacing one or more cooling hardware modules that reflect a change in the data center design, and then generating predicted simulation results in real time.

[0036] In this way, digitally simulated data center cooling system 200 can facilitate rapid design adjustments and efficient real-time monitoring of a data center infrastructure, while optimizing performance without costly physical testing.

[0037] In at least one embodiment, as in Fig. As shown in Figure 2B, a data center room module 230 can be constructed by combining server module 236, rack module 235, hot / cold airflow aisle simulation module 233, piping module 241, electrical module 242, airflow simulation module 245, coolant flow simulation module 246, air cooling module 237, and / or the like. In at least one embodiment, each of these modules can be trained on training data representing data center conditions, such as the number of racks, heat loads, external conditions, and / or the like, and corresponding thermal states of the coolant and / or physical results of the airflow in aisles. For example, data center room module 230 can perform real-time inference of the physical results based on an input of the current data center conditions.

[0038] In at least one embodiment, each of the server module 236, rack module 235, hot / cold airflow simulation module 233, piping module 241, electrical module 242, and air cooling module 237, airflow simulation 245, and coolant flow simulation 246 can, upon inference, generate an output that reveals one or more thermal states and / or other characteristics of one or more cooling outputs associated with the individual cooling hardware. For example, a rack module 235 can receive an input that reveals at least characteristics of a supply coolant (from the CDU module 225), control parameters (such as rack fan speed, airflows, etc.), and / or other external data center conditions to generate predicted characteristics of a return coolant from the rack, which is then returned to the CDU module 225.As another example, each server module 236 can receive an input that specifies at least characteristics of a server-level supply coolant, a heat load and / or other external data center conditions to generate predicted server-level return coolant characteristics.

[0039] In at least one embodiment, one or more of the server module 236, rack module 235, hot / cold airflow aisle simulation module 233, piping module 241, electrical module 242, and air cooling module 237, airflow simulation 245, and coolant flow simulation 246 can perform joint inference, since the modules can operate in conjunction with each other. For example, a rack module 235 can include one or more server modules 236, and inference of the rack module 235 can be performed in conjunction with at least the airflow aisle simulation module 233, the coolant flow simulation module 246, and the airflow simulation 245, while the latter generates a prediction of airflow dynamic characteristics that shows how airflow passes through a physical rack.In at least one embodiment, additional details of the air cooling module 237, the hot and cold aisle module 233, and the airflow simulation module 245, which performs physically informed simulation of airflow dynamics, can be found in the concurrently pending and jointly assigned US patent application No. 18 / 949,430, attorney file number 42940.81US01, which was filed on the same day.

[0040] In at least one embodiment, as in Fig. As shown in Figure 2C, a CDU module 225 can be constructed by combining a pump module 251, a heat exchanger module 252, a module of valves and filters 253, a control module 254, a failure and failover simulation module 255, and / or the like. In at least one embodiment, each of these modules 251-255 can be trained on training data representing thermal performance data available to the respective module, so that the CDU module 225 is aware of all its capabilities. For example, performance data for training can include control data (such as, but not limited to, pump control parameters, heat exchanger fan speed, valve setting, and / or the like), curves of thermal states (e.g., coolant and / or airflow temperature, airflow velocity, pressure over time, etc.), failure and failover simulation (e.g., CDU failure, etc.), and / or the like.For example, the CDU module 225 can perform real-time inference of simulated physical results, such as, but not limited to, thermal states (e.g., temperature and flow rate of supply coolant, etc.) of a simulated CDU, based on an input of one or more of control parameters (e.g., valve setting, fan speed, and / or the like), thermal states of return coolant (e.g., temperature and flow rate of a return coolant), data center conditions (such as number of racks, heat loads, external conditions), and / or the like.

[0041] In at least one embodiment, control module 254 can control and / or coordinate the operation of the CDU module 225, as further described in Fig. 4A described.

[0042] In at least one embodiment, CDU module 225 can update predicted simulation results of the CDU in real time in response to any changes in the control parameters. In at least one embodiment, CDU module 225 can be updated by removing, adding, or replacing one or more cooling hardware modules 251-255 that reflect a change in the CDU design, and then generating predicted simulation results in real time. In this way, CDU 225 can facilitate rapid design adjustments and efficient real-time monitoring of CDU designs, while optimizing performance without costly physical testing.

[0043] In at least one embodiment, as in Fig. As shown in 2D, a heat rejection module 410 can be constructed by combining a cooling tower module 206, a chiller module 205, and / or a dry cooler module 207. In at least one embodiment, each of the cooling tower module 206, chiller module 205, and / or dry cooler module 207 can be trained on training data representing thermal performance data available to the respective module, so that the heat rejection module 410 is aware of all its performance capabilities. For example, performance data for training from the cooling tower module 206 can include control data (such as, but not limited to, water pump control parameters, cooling tower fan speed, valve setting, and / or the like), curves of thermal states (e.g., cooling fluid temperature, airflow velocity, pressure, etc.), and / or the like.

[0044] In at least one embodiment, the heat dissipation module 410, combining the cooling tower module 206, chiller module 205, and / or dry cooler module 207, can be further trained to generate a thermal response in response to a specific site's climatic condition, such as historical weather conditions from the past 50 years. For example, in at least one embodiment, the heat dissipation module 410 can generate updated thermal states based on an input that includes suggested operating parameters (flow rate, temperature, and cooling load) and an urban location that reflects a climatic condition. In this way, the data center simulation platform 200, which integrates the heat dissipation module 410, can generate an output to assist users in selecting the heat dissipation system (e.g., chiller vs. dry cooler) for information technology (IT) requirements, based at least partially on the climatic conditions of a specific site.The data center simulation platform can then generate 200 possible solutions with the goal of identifying the most energy-efficient scenario for data center designs. In at least one implementation, a CapEx and OpEx comparison of different solutions can be provided, for example, via a user interface, enabling data center design engineers to make informed decisions for each specific design.

[0045] Fig. Figures 3A-3B illustrate exemplary data transmission between neural module networks corresponding to different types of data center modules, to simulate the process described in Fig. to facilitate the data center system described in section 1 according to the embodiments described herein. In at least one embodiment, such as in Fig. As shown in Figure 3A, all modules can be individually trained to simulate a single type of cooling hardware. In at least one embodiment, one or more modules can be combined and trained together to simulate a cooling component of a data center. For example, chiller module 206 and cooling tower module 205 can be trained together to simulate an external heat exchanger circuit of a data center. As another example, modules 251–255 can be trained together to simulate a CDU module 225. As yet another example, modules 233, 235, 236, and 242 can be trained together to simulate a data center room module 230.

[0046] In at least one embodiment, one or more modules can be selected and integrated based on a data center design and / or data center cooling hardware. In at least one embodiment, a data pipeline 310 or 320 can be created to manage the transfer of data between modules for real-time inference of performance and results.

[0047] In at least one embodiment, data transmission pipeline 310 can transmit real-time inference results from the cooling tower module 205 and the chiller module 206, such as, but not limited to, a predicted primary coolant supply flow rate, primary coolant temperature, and / or the like, to an input in the CDU module 225. In at least one embodiment, data transmission pipeline 310 can transmit real-time inference results from the CDU module 225, such as, but not limited to, a predicted primary coolant return flow rate, primary coolant temperature, and / or the like, to an input in the chiller module 206.

[0048] In at least one embodiment, data transmission pipeline 311 can transmit real-time inference results from the cooling tower module 205 and the chiller module 206, such as, but not limited to, a predicted primary coolant supply flow rate, primary coolant temperature, and / or the like, to an input to the air cooling module 237. In at least one embodiment, the air cooling module 237 may not receive data transmission 311, for example, when the air cooling module 237 is simulating a CRAC that includes a compressor / condenser module to use a refrigerant to cool hot air streams.

[0049] In at least one embodiment, data transmission pipeline 320 can transmit real-time inference results from CDU module 225, such as, but not limited to, a predicted secondary coolant supply flow rate, secondary supply coolant temperature, and / or the like, to an input to rack module 235. In at least one embodiment, data transmission pipeline 320 can transmit real-time inference results from data center room module 230, which includes modules 233, 235, 236, and 242, such as, but not limited to, a predicted secondary coolant return flow rate, secondary return coolant temperature, and / or the like, to an input to CDU module 225.

[0050] In at least one embodiment, data transmission pipeline 321 can transmit real-time inference results from the air cooling module 237, such as, but not limited to, the temperature of the cold airflow, airflow velocity, pressure, and / or the like, to an input for the rack module 235. In at least one embodiment, data transmission pipeline 321 can transmit real-time inference results from the data center room module 230, which includes modules 233, 235, 236, and 242, such as, but not limited to, the hot airflow rate, airflow temperature, airflow dynamics, and / or the like, to an input for the air cooling module 237.

[0051] In at least one embodiment, such as in Fig. As shown in Figure 3B, real-time simulation of the CDU module 225 and the data center room module 230 can be performed by inferring from the combined trained modules 205, 206, 225, and 230 and transferring data between components for transient simulation. In at least one embodiment, the CDU module 225 and the data center room module 230 can exchange real-time data via 405a-b. In at least one embodiment, the performance of the CDU module 225 can be monitored in real time, for example, by observing variations in data center parameters of the physical CDU 125 (at 403).

[0052] In this way, real-time simulation of the CDU module 225 and data center room module 230 allows for the rapid implementation of design changes to the CDU or the data center (e.g., adding, replacing, or removing a module), and the results are monitored by updating the real-time simulations. In at least one embodiment, real-time simulations can be provided to a user through a user interface, enabling the user to decide whether to retain a design change or modify it further.

[0053] Fig. Figures 4A-4B illustrate exemplary control modules corresponding to modular neural networks, which correspond to different types of data center modules, according to embodiments described herein. In at least one embodiment, such as in Fig. As shown in Figure 4A, the transient and real-time responses of the digitally simulated modules can enable the design and tuning of the control system. For example, control module 254a can be trained to transfer data from the combined chiller / tower module 410 (featuring modules 205-206) to the CDU module 225 and / or air cooling module 237, while providing feedback control.

[0054] In at least one embodiment, control module 254a can receive an input 402 containing data specifying a primary coolant temperature, a primary coolant flow rate, and / or the like, and convert such an input 402 via control parameters of the control elements 405 into an output 406 (e.g., adjusted primary coolant temperature, adjusted primary coolant flow rate, and / or the like), which is passed to the CDU module 225 and / or the air cooling module 237.

[0055] In at least one embodiment, based on data transmitted to the CDU module 225 and / or air cooling module 237 from the combined chiller / tower module 410 (comprising modules 205-206), control module 254a can receive feedback 407 from the CDU module 225 and / or other data center modules. For example, feedback 407 can include the temperature and flow rate of a secondary supply coolant, the temperature and flow rate of a secondary return coolant, server cooling capacity, and / or the like. In at least one embodiment, feedback 407 can be fed back to control elements 405 to generate control signals to adjust the valves and pumps for CDU module 225 and / or air cooling module 237 to ensure that cooling requirements are met.

[0056] In at least one embodiment, a user can adjust control module 254a, for example, by updating or adjusting one or more control parameters and / or control mechanisms via a user interface, and observe the response of the digitally simulated data center 200 in real time, such as updated thermal states of the digitally simulated data center 200. For example, conventional physical CDUs have manual tuning control systems for engineers to manually adjust the control system and change the parameters based on a target cooling requirement. In at least one embodiment, control module 254a can allow a user to adjust control parameters while observing any performance change caused by control parameters without physical work or hardware modification.

[0057] In at least one embodiment, such as in Fig. As shown in Figure 4B, control module 254b can be trained to transmit data from CDU module 225 (e.g., temperature and flow rate of the secondary supply coolant and / or the like) and / or air cooling module 237 (e.g., airflow dynamics, temperature, etc.) to the data center room module 230 while providing feedback control. In at least one embodiment, control module 254b can receive an input 412 containing data specifying a secondary coolant temperature, a secondary coolant flow rate, airflow dynamics, air temperature, and / or the like, and convert such input 412, via control parameters of the controls 415, into an output 416 (e.g., adjusted secondary coolant temperature, adjusted secondary coolant flow rate, and / or the like) that is passed to the data center room module 230.In at least one embodiment, based on data transmitted to the data center room module 230 from the CDU module 225 and / or air cooling module 237, control module 254b can receive feedback 417 from the data center room module 230 and / or other data center modules. For example, feedback 417 can include the temperature and flow rate of a secondary return coolant, server cooling capacity, hot air dynamics, temperature, and / or the like. In at least one embodiment, feedback 417 can be fed back to control elements 405b to generate control signals to adjust the valves and pumps for data center room 230 to ensure that cooling requirements are met.

[0058] In at least one embodiment, a self-tuning control system can be trained by training data 419, which includes at least feedback 417 and / or feedback 407 and / or other feedback from another cooling hardware module and / or physical simulation results of the digitally simulated data center, to automatically generate control parameters for the CDU module 225 and / or air cooling module 237 and / or various modules within the data center room module 230. In at least one embodiment, for a single design of the data center system, the self-tuning control system can generate predicted control parameters for each cooling hardware component, depending on the individual design. In at least one embodiment, if a change is made to the design of the data center system (e.g.,(through a user who modifies, adds, removes, or replaces cooling hardware), the self-tuning control system generates updated control parameters for the updated cooling hardware, depending on the changed design of the data center system.

[0059] Fig. Figure 5 illustrates an exemplary logic flow diagram of a method for using one or more modular neural networks to jointly simulate the operation of a data center cooling system, according to at least one embodiment. One or more of the processes of Method 500 may be implemented, at least partially, in the form of executable code stored on non-transitory, tangible, machine-readable media which, when executed by one or more processors, can cause the one or more processors to perform one or more of the processes. In some embodiments, Method 500 corresponds to the operation of the digital data center simulation module 630 in Fig. 6A.

[0060] As illustrated, Procedure 500 includes a number of listed steps, but aspects of Procedure 500 may include additional steps before, after, and between the listed steps. In some aspects, one or more of the listed steps may be omitted or performed in a different order.

[0061] In at least one embodiment, in step 502, one or more modular neural networks (e.g., modules 205, 206, 225, 235, 236, 233, 241 and / or the like, which are in Fig. (shown in 2A-4B) are trained to simulate different types of data center cooling hardware, each using training data of data center conditions and / or performance data. For example, one or more types of data center hardware may include one or more of: a chiller (e.g., 206 in Fig. 2A), a cooling tower (e.g. 205 in Fig. 2A), a CDU (e.g. 225 in Fig. 2A), a server (e.g. 236 in Fig. 2A), a server rack (e.g. 235 in Fig. 2A), a data center corridor (e.g. 233 in Fig. 2A) and an electrical power module (e.g. 242 in Fig. 2B). For example, data center conditions may include one or more of: a number of server racks, a thermal load for each server rack, and cooling capacity and configuration data specific to a particular type of data input hardware. As another example, performance data may include one or more control parameters associated with one or more types of data center hardware, failure and failover associated with one or more types of data center hardware, thermal conditions (e.g., coolant flow rate, temperature, airflow velocity, pressure, and / or the like) associated with one or more types of data center hardware, and / or the like.

[0062] In at least one embodiment, the one or more neural networks comprise a first neural network (e.g., 225 in Fig. 2C) which is trained to simulate the CDU, and wherein the first neural network further comprises one or more sub-neural networks, each of which is trained using a dataset of CDU control data, CDU failures, and CDU failovers to simulate a respective subcomponent of the CDU. The respective subcomponent of the CDU comprises one or more of: a CDU pump (e.g., 251 in Fig. 2C), a heat exchanger (e.g. 252 in Fig. 2C), one or more valves and filters (e.g. 253 in Fig. 2C) and a CDU control module (e.g. 254 in Fig. 2C).

[0063] In at least one embodiment, in step 502, the one or more neural networks corresponding to the one or more types of data center hardware can be integrated based at least partially on a data center cooling system architecture. For example, in at least one embodiment, a user can select or modify the one or more neural networks corresponding to the one or more types of data center hardware via a graphical user interface. In at least one embodiment, a user can connect and integrate the selected one or more neural networks based on a desired data center cooling system architecture via a user interface.

[0064] As another example, in at least one embodiment, another neural network can determine and integrate the one or more neural networks based on an input of desired data center cooling system characteristics. In at least one embodiment, another neural network can generate one or more code programs that integrate the one or more neural networks and perform data transmissions (e.g., 310, 320 in Fig. 3A, 405a-405b in Fig. 3B) establish a connection between one or more neural networks.

[0065] In at least one embodiment, in step 503, one or more modular neural networks can jointly simulate the operation of the data center cooling system, including real-time data transmission (e.g., 310, 320 in Fig. 3A, 405a-405b in Fig. 3B) between the one or more integrated neural networks.

[0066] In at least one embodiment, at step 504, simulation results of the data center cooling system can be updated in real time in response to a change in the data center cooling system. For example, the change might involve user input adding, removing, modifying, or replacing cooling hardware. A selection of one or more neural networks is updated in real time in response to a change in data center hardware within the data center cooling system.

[0067] As another example, the modification can include one or more control parameters associated with one or more types of data center cooling hardware. In at least one embodiment, at least one neural network of the one or more neural networks is caused to generate one or more control parameters for at least one type of data center hardware based at least partially on one or more simulation results of the data center cooling system.

[0068] Fig. Figure 6A is a simplified diagram illustrating a computing device implementing a digitally simulated data center cooling system located in Fig. 1-5, according to an embodiment described herein. As described in Fig. As shown in Figure 6, computing device 612 includes a processor 618 coupled to memory 620. Operation of computing device 612 is controlled by processor 618. Although computing device 612 is shown with only one processor 618, it is understood that processor 618 can represent one or more central processing units, multi-core processors, microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), graphics processing units (GPUs), and / or the like in computing device 612. Computing device 612 can be implemented as a standalone subsystem, as a circuit board added to a computing device, and / or as a virtual machine.

[0069] Memory 620 can be used to store software executed by Computing Device 612 and / or one or more data structures used during operation of Computing Device 612. Memory 620 can include one or more types of machine-readable media. Some common forms of machine-readable media may include floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punched cards, paper tape, any other physical medium with hole patterns, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other medium for which a processor or computer is designed to read.

[0070] The processor 618 and / or memory 620 can be arranged in any suitable physical configuration. In some embodiments, the processor 618 and / or memory 620 can be implemented on the same board, in the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), and / or the like. In some embodiments, the processor 618 and / or memory 620 can include distributed, virtualized, and / or containerized computing resources. Consistent with such embodiments, the processor 618 and / or memory 620 can be located in one or more data centers and / or cloud computing facilities.

[0071] In another embodiment, processor 618 can have multiple microprocessors and / or memory 620 can have multiple registers and / or other memory elements, so that processor 618 and / or memory 620 can be arranged in the form of a hardware-based neural network.

[0072] In some examples, memory 620 can contain non-transient, tangible, machine-readable media containing executable code which, when executed by one or more processors (e.g., processor 610), can cause the one or more processors to perform the procedures described in more detail herein. For example, as shown, memory 620 contains instructions for the digital data center simulation module 630, which can be used to implement and / or emulate the systems and models and / or to implement any of the procedures described herein. The digitally simulated data center simulation module 630 can receive input 640, such as input training data (e.g., data center conditions, performance data, etc.), via the data interface 615 and produce an output 650, which may be simulated thermal states, etc., in response to input 640.

[0073] The data interface 615 can include a communication interface, a user interface (such as a voice input interface, a graphical user interface, and / or the like). For example, the computing device 600 can receive input 640 (such as a training data set) from a networked database via a communication interface. Or the computing device 600 can receive input 640, such as user-configured control parameters, from a user via the user interface.

[0074] In some embodiments, the Digital Data Center Simulation Module 630 is configured to digitally simulate a physical data center cooling system, as in Fig. 1 described. The digital data center simulation module 630 can further comprise several neural network submodules 631-632, as described in Fig. 2A -5 described, which together simulate, monitor and / or control a data center cooling system in real time.

[0075] Some examples of computing devices, such as Computing Device 600, can include non-transitory, tangible, machine-readable media containing executable code which, when executed by one or more processors (e.g., Processor 610), can cause the one or more processors to perform the processes of the procedure. Some common forms of machine-readable media that can contain the processes of the procedure are, for example, floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punched cards, paper tape, any other physical medium with hole patterns, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other medium for which a processor or computer is designed to read. Servers and data centers

[0076] The following figures represent, without limitation, exemplary network server and data center-based systems that can be used to implement at least one embodiment.

[0077] Fig. Figure 6B illustrates a distributed system 600 according to at least one embodiment. In at least one embodiment, the distributed system 600 includes one or more client computing devices 602, 604, 606, and 608 configured to run and operate a client application, such as a web browser, a proprietary client, and / or variations thereof, over one or more networks 610. In at least one embodiment, the server 612 can be communicatively coupled with remote client computing devices 602, 604, 606, and 608 over network 610.

[0078] In at least one embodiment, Server 612 can be configured to run one or more services or software applications, such as services and applications that can manage session activity for single sign-on (SSO) access across multiple data centers. In at least one embodiment, Server 612 can also provide other services, or software applications can include non-virtual and virtual environments. In at least one embodiment, these services can be offered as web-based or cloud services, or under a software-as-a-service (SaaS) model, to users of client computing devices 602, 604, 606, and / or 608. In at least one embodiment, users operating client computing devices 602, 604, 606, and / or 608 can, in turn, use one or more client applications to interact with Server 612 to use services provided by these devices.

[0079] In at least one embodiment, software components 618, 620, and 622 of distributed system 600 are implemented on server 612. In at least one embodiment, one or more components of distributed system 600 and / or services provided by these components may also be implemented by one or more of the client computing devices 602, 604, 606, and / or 608. In at least one embodiment, users operating client computing devices may then use one or more client applications to access services provided by these components. In at least one embodiment, these components may be implemented in hardware, firmware, software, or combinations thereof. It is understood that various different system configurations are possible that may differ from distributed system 600. Fig. The embodiment shown in Figure 6 is therefore an example of a distributed system for implementing an embodiment system and is not intended to be restrictive.

[0080] In at least one embodiment, client computing devices 602, 604, 606, and / or 608 can include various types of computing systems. In at least one embodiment, a client computing device can include portable handheld devices (e.g., an iPhone®, a mobile phone, an iPad®, a computer tablet, a personal digital assistant (PDA)) or wearable devices (e.g., a head-mounted display of Google Glass®), running software such as Microsoft Windows Mobile®, and / or a variety of mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 60, Palm OS, and / or variations thereof. In at least one embodiment, devices can support various applications, such as various internet-related apps, email, instant messaging (SMS) applications, and can use various other communication protocols.In at least one embodiment, client computing devices may also include general-purpose personal computers, including, for example, personal computers and / or laptop computers running various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems. In at least one embodiment, client computing devices may be workstation computers running any one of a variety of commercially available UNIX® or UNIX-like operating systems, including, without limitation, a variety of GNU / Linux operating systems, such as Google Chrome OS. In at least one embodiment, client computing devices may also include electronic devices, such as a thin client computer, an internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device capable of communicating over a network or networks.Although distributed system 600 in . Fig. Although Figure 6 shows four client computing devices, any number of client computing devices can be supported. Other devices, such as devices with sensors, etc., can interact with Server 612.

[0081] In at least one embodiment, network(s) 610 in the distributed system 600 can be any type of network capable of supporting data communications using any of a variety of available protocols, including, without limitation, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Packet Exchange), AppleTalk, and / or variations thereof. In at least one embodiment, the network(s) 610 can be a local area network (LAN), networks based on Ethernet, Token Ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g.,a network operating under any of the Institute of Electrical and Electronics (IEEE) 802.11 protocol suite, Bluetooth® and / or any other wireless protocol) and / or any combination of these and / or other networks.

[0082] In at least one embodiment, Server 612 can consist of one or more general-purpose computers, specialized server computers (including, for example, PC (Personal Computer) servers, UNIX® servers, mid-range servers, mainframes, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. In at least one embodiment, Server 612 can include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization. In at least one embodiment, one or more flexible pools of logical storage devices can be virtualized to maintain virtual storage for a server. In at least one embodiment, virtual networks can be controlled by Server 612 using software-defined networking.In at least one embodiment, Server 612 can be designed to run one or more services or software applications.

[0083] In at least one embodiment, Server 612 can run any operating system as well as any commercially available server operating system. In at least one embodiment, Server 612 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP (Hypertext Transport Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Java® servers, database servers, and / or variations thereof. In at least one embodiment, exemplary database servers include, without limitation, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variations thereof.

[0084] In at least one embodiment, Server 612 may include one or more applications for analyzing and consolidating data feeds and / or event updates received by users of client computing devices 602, 604, 606, and 608. In at least one embodiment, data feeds and / or event updates may include, among others, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third-party information sources and continuous data streams that report real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and / or variations thereof.In at least one embodiment, Server 612 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 602, 604, 606 and 608.

[0085] In at least one embodiment, distributed system 600 can also include one or more databases 614 and 616. In at least one embodiment, the databases can provide a mechanism for storing information, such as user interaction information, usage pattern information, adaptation rules information, and other information. In at least one embodiment, databases 614 and 616 can reside in a plurality of locations. In at least one embodiment, one or more of the databases 614 and 616 can reside locally on (and / or be resident in) server 612 on a non-transitory storage medium. In at least one embodiment, databases 614 and 616 can be located remotely from server 612 and communicate with server 612 via a network-based or dedicated connection.In at least one embodiment, databases 614 and 616 can reside in a storage area network (SAN). In at least one embodiment, all necessary files for performing functions attributed to server 612 can be stored locally on server 612 and / or remotely, as appropriate. In at least one embodiment, databases 614 and 616 can include relational databases, such as databases designed to store, update, and retrieve data in response to SQL-formatted commands.

[0086] In at least one embodiment, System 600 can include a data center that can be digitally simulated by one or more neural networks, as in at least one embodiment described in Fig. 1-5 is described.

[0087] Fig. Figure 7 illustrates an exemplary data center 700 according to at least one embodiment. In at least one embodiment, data center 700 includes, without limitation, a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0088] In at least one embodiment, as in Fig. As shown in Figure 7, data center infrastructure layer 710 can include a resource orchestrator 712, clustered compute resources 714, and node compute resources (“Node CRs”) 716(1)-216(N), where “N” is a positive integer. In at least one embodiment, Node CRs 716(1)-216(N) can include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (“FPGAs”), graphics processing units, etc.), storage devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk storage), network input / output devices (“NW I / O”), network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node CRs from node CRs 716(1)-216(N) can be a server with one or more of the above-mentioned computing resources.

[0089] In at least one embodiment, clustered compute resources 714 can include separate groupings of node CRs located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of node CRs within clustered compute resources 714 can include clustered compute, network, short-term memory, or long-term storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs, including CPUs or processors, can be grouped within one or more racks to provide compute resources to support one or more workloads.In at least one embodiment, one or more racks can also include any number of power modules, cooling modules and network switches, in any combination.

[0090] In at least one embodiment, Resource Orchestrator 712 can configure or otherwise control one or more nodes of CRs 716(1)-216(N) and / or grouped compute resources 714. In at least one embodiment, Resource Orchestrator 712 can include a software design infrastructure (“SDI”) management entity for data center 700. In at least one embodiment, Resource Orchestrator 712 can include hardware, software, or a combination thereof.

[0091] In at least one embodiment, as in Fig. As shown in Figure 7, framework layer 720 includes, without limitation, a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, framework layer 720 may include a framework to support software 752 of software layer 730 and / or one or more applications 742 of application layer 740. In at least one embodiment, software 752 or application(s) 742 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 720 may be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter “Spark”), which may use distributed file system 738 for large-scale data processing (e.g., “Big Data”).In at least one embodiment, the job scheduler 732 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 700. In at least one embodiment, the configuration manager 734 can be able to configure different layers, such as the software layer 730 and the framework layer 720, including Spark and the distributed file system 738, to support large-scale data processing. In at least one embodiment, the resource manager 736 can be able to manage clustered or grouped compute resources mapped to or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, clustered or grouped compute resources can include grouped compute resources 714 at the data center infrastructure layer 710.In at least one embodiment, resource manager 736 can coordinate with resource orchestrator 712 to manage these mapped or allocated computing resources.

[0092] In at least one embodiment, software 752, which is included in software layer 730, may include software used by at least portions of nodes CRs 716(1)-216(N), grouped computing resources 714, and / or distributed file systems 738 of frame layer 720. One or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.

[0093] In at least one embodiment, the application(s) 742 included in application layer 740 may include one or more types of applications that utilize at least proportions of node CRs 716(1)-216(N), grouped compute resources 714, and / or distributed file system 738 of frame layer 720. At least one or more types of applications may include, without limitation, CUDA applications, 5G network applications, artificial intelligence applications, data center applications, and / or variations thereof.

[0094] In at least one embodiment, each of the configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible way. In at least one embodiment, self-modifying actions can relieve a data center 700 operator of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of a data center.

[0095] In at least one embodiment, Data Center 700 can adopt a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0096] Fig. Figure 8 illustrates a client-server network 804 formed by a plurality of network server computers 802 interconnected according to at least one embodiment. In at least one embodiment, each network server computer 802 stores data that can be accessed by other network server computers 802 and client computers 806 and networks 808 connecting to a wide area network 804. In at least one embodiment, the configuration of a client-server network 804 can change over time as client computers 806 and one or more networks 808 connect to and disconnect from a network 804, and as one or more trunk-line server computers 802 are added to or removed from a network 804.In at least one embodiment, when a client computer 806 and a network 808 are connected to network server computers 802, the client-server network includes such client computer 806 and network 808. In at least one embodiment, the term computer includes any device or machine capable of receiving data, applying prescribed processes to data, and producing results of processes.

[0097] In at least one embodiment, a client-server network 804 stores information that is accessible to network server computers 802, remote networks 808, and client computers 806. In at least one embodiment, network server computers 802 are formed by mainframes, minicomputers, and / or microcomputers, each with one or more processors. In at least one embodiment, server computers 802 are interconnected by wired and / or wireless transmission media, such as conductive wire, fiber optic cable, and / or microwave transmission media, satellite transmission media, or other conductive, optical, or electromagnetic wave transmission media. In at least one embodiment, client computers 806 access a network server computer 802 by means of a similar wired or wireless transmission medium.In at least one embodiment, a client computer 806 can connect to a client-server network 804 using a modem and a standard telephone communication network. In at least one embodiment, alternative carrier systems, such as cable and satellite communication systems, can also be used to connect to the client-server network 804. In at least one embodiment, other private or time-shared carrier systems can be used. In at least one embodiment, network 804 is a global information network, such as the Internet. In at least one embodiment, network 804 is a private intranet that uses protocols similar to the Internet, but with additional security measures and restricted access controls. In at least one embodiment, network 804 is a private or semi-private network that uses proprietary communication protocols.

[0098] In at least one embodiment, client computer 806 is any end-user computer and may also be a mainframe, minicomputer, or microcomputer with one or more microprocessors. In at least one embodiment, server computer 802 may temporarily function as a client computer accessing another server computer 802. In at least one embodiment, remote network 808 may be a local area network, a network added to a wide area network by an independent internet service provider (ISP), or any other group of computers interconnected by wired or wireless transmission media in a configuration that is either fixed or changes over time. In at least one embodiment, client computers 806 may connect to and access a network 804 independently or through a remote network 808.

[0099] In at least one embodiment, various components of Network 804 can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0100] Fig. Figure 9 illustrates a computer network 900 that connects one or more computing machines according to at least one embodiment. In at least one embodiment, network 900 can be any type of electronically connected group of computers, including, for example, the following networks: Internet, intranet, local area networks (LANs), wide area networks (WANs), or a connected combination of these network types. In at least one embodiment, connectivity within a network 908 can be a remote modem, Ethernet (IEEE 802.3), Token Ring (IEEE 802.5), Fiber Distributed Datalink Interface (FDDI), Asynchronous Transfer Mode (ATM), or any other communication protocol. In at least one embodiment, computing devices connected to a network can be desktop computers, servers, portable computers, handheld computers, set-top boxes, personal digital assistants (PDAs), terminals, or any other desired type or configuration.In at least one embodiment, depending on their functionality, networked devices can vary greatly in terms of processing power, internal memory, and other performance aspects. In at least one embodiment, communications within a network and to or from computing devices connected to a network can be either wired or wireless. In at least one embodiment, Network 908 can include at least part of the worldwide public Internet, which generally connects a large number of users according to a client-server model as defined by a Transmission Control Protocol / Internet Protocol (TCP / IP) specification. In at least one embodiment, client-server networking is a dominant model for communication between two computers. In at least one embodiment, a client computer ("client") issues one or more commands to a server computer ("server").In at least one embodiment, a server fulfills client commands by accessing available network resources and returning information to a client according to the client commands. In at least one embodiment, client computer systems and network resources residing on network servers are assigned a network address for identification during communication between network elements. In at least one embodiment, communication from other network-connected systems to servers includes the network address of a relevant server / network resource as part of the communication, so that a suitable target of a data / request is identified as the receiver.In at least one embodiment, where a network 908 comprises the global Internet, a network address is an IP address in a TCP / IP format that can route at least some data to an email account, a website, or another Internet tool residing on a server. In at least one embodiment, information and services residing on network servers can be available to a web browser on a client computer via a domain name (e.g., www.site.com) that maps to an IP address of a network server.

[0101] In at least one embodiment, a plurality of clients 902, 904, and 906 are connected to a network 908 via respective communication links. In at least one embodiment, each of these clients can access a network 908 via any desired form of communication, such as a dial-up modem connection, a cable connection, a digital subscriber line (DSL), a wireless or satellite connection, or any other form of communication. In at least one embodiment, each client can communicate using any machine compatible with a network 908, such as a personal computer (PC), a workstation, a dedicated terminal, a personal data assistant (PDA), or other similar equipment. In at least one embodiment, clients 902, 904, and 906 may or may not be located in the same geographic area.

[0102] In at least one embodiment, a plurality of servers 910, 912, and 914 are connected to a network 908 to serve clients communicating with the network 908. In at least one embodiment, each server is typically a high-performance computer or device that manages network resources and responds to client commands. In at least one embodiment, servers include computer-readable data storage media, such as hard disk drives and RAM, which store program commands and data. In at least one embodiment, the servers 910, 912, and 914 run application programs that respond to client commands. In at least one embodiment, server 910 can run a web server application to respond to client requests for HTML pages and can also run a mail server application to receive and route electronic mail.In at least one embodiment, other application programs, such as an FTP server or a media server for streaming audio / video data to clients, can also run on a Server 910. In at least one embodiment, different servers can be dedicated to perform different tasks. In at least one embodiment, Server 910 can be a dedicated web server that manages website resources for different users, whereas a Server 912 can be dedicated to provide electronic mail (email) management. In at least one embodiment, other servers can be dedicated to media (audio, video, etc.), File Transfer Protocol (FTP), or a combination of any two or more services typically available or provided over a network.In at least one embodiment, each server can be located at a position that is the same or different from that of other servers. In at least one embodiment, there can be multiple servers that perform mirrored tasks for users, thereby relieving congestion or minimizing traffic directed to and from a single server. In at least one embodiment, servers 910, 912, and 914 are under the control of a web hosting provider in a business of maintaining and delivering third-party content over a network 908.

[0103] In at least one embodiment, web hosting providers deliver services to two different types of clients. In at least one embodiment, one type, which can be called a browser, requests content from servers 910, 912, 914, such as web pages, email messages, video clips, etc. In at least one embodiment, a second type, which can be called a user, engages a web hosting provider to maintain a network resource, such as a website, and make it available to browsers. In at least one embodiment, users agree with a web hosting provider to make storage space, processing capacity, and communication bandwidth available for their desired network resource according to a set of server resources that a user wishes to use.

[0104] In at least one embodiment, for a web hosting provider to provide services to both of these clients, application programs that manage a server-hosted network resource must be properly configured. In at least one embodiment, the program configuration process involves defining a set of parameters that at least partially control an application program's response to browser requests and that also at least partially define a server that is available to a particular user.

[0105] In one embodiment, an intranet server 916 communicates with a network 908 via a communication link. In at least one embodiment, the intranet server 916 communicates with a server manager 918. In at least one embodiment, the server manager 918 has a database of application program configuration parameters that are used in servers 910, 912, and 914. In at least one embodiment, users modify a database 920 via an intranet 916, and a server manager 918 interacts with servers 910, 912, and 914 to modify application program parameters so that they match the contents of a database. In at least one embodiment, a user logs on to an intranet server 916 by connecting to the intranet 916 via the client 902 and entering authentication information, such as a username and password.

[0106] In at least one embodiment, when a user wishes to log on to a new service or modify an existing service, an intranet server 916 authenticates the user and provides the user with an interactive screen display / control panel that allows the user to access configuration parameters for a specific application program. In at least one embodiment, a number of modifiable text boxes are presented to the user, describing aspects of a configuration of the user's website or other network resource. In at least one embodiment, when a user wishes to increase storage space reserved on a server for their website, a field is provided in which the user specifies the desired storage space.In at least one embodiment, an intranet server 916 updates a database 920 in response to receiving this information. In at least one embodiment, a server manager 918 forwards this information to a suitable server, and a new parameter is used during application program operation. In at least one embodiment, an intranet server 916 is configured to provide users with access to configuration parameters of hosted network resources (e.g., web pages, email, FTP sites, media sites, etc.) for which a user has a contract with a web hosting service provider.

[0107] In at least one embodiment, various components of Network 900 can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0108] Fig. Figure 10A illustrates a networked computer system 1000A according to at least one embodiment. In at least one embodiment, the networked computer system 1000A comprises a plurality of nodes or personal computers (“PCs”) 1002, 1018, 1020. In at least one embodiment, a personal computer or node 1002 comprises a processor 1014, memory 1016, video camera 1004, microphone 1006, mouse 1008, speaker 1010, and monitor 1012. In at least one embodiment, nodes 1002, 1018, 1020 can each run one or more desktop servers of an internal network within a given company or can be servers of a general network that is not limited to a specific environment. In at least one embodiment, there is one server per PC node of a network, such that each PC node of a network represents a single network server with a single network URL address.In at least one embodiment, each server is by default a standard web page for that server's users, which may itself contain embedded URLs that point to further subpages of that user on that server or to other servers or pages on other servers in a network.

[0109] In at least one embodiment, nodes 1002, 1018, 1020, and other nodes of a network are interconnected via medium 1022. In at least one embodiment, medium 1022 can be a communication channel, such as an Integrated Services Digital Network (“ISDN”). In at least one embodiment, different nodes of a networked computer system can be connected through a variety of communication media, including local area networks (“LANs”), simple telephone lines (“POTS”), sometimes referred to as public switched telephone networks (“PSTN”), and / or variations thereof. In at least one embodiment, different nodes of a network can also be computer system users who are interconnected via a network such as the Internet.In at least one embodiment, each server in a network (running from a single node of a network at a given instance) has a unique address or identification within a network, which may be specified in the form of a URL.

[0110] In at least one embodiment, a plurality of multipoint conference units (“MCUs”) can be used to transmit data to and from different nodes or “endpoints” of a conference system. In at least one embodiment, nodes and / or MCUs can be connected via an ISDN link or through a local area network (“LAN”), in addition to various other communication media, such as nodes connected via the internet. In at least one embodiment, nodes of a conference system can generally be directly connected to a communication medium, such as a LAN or via an MCU, and a conference system can include other nodes or elements, such as routers, servers, and / or variations thereof.

[0111] In at least one embodiment, processor 1014 is a general-purpose programmable processor. In at least one embodiment, processors of nodes of networked computer system 1000A can also be special-purpose video processors. In at least one embodiment, various peripheral devices and components of a node, such as those of node 1002, can differ from those of other nodes. In at least one embodiment, nodes 1018 and 1020 can be configured identically to or differently from node 1002. In at least one embodiment, a node can be implemented on any suitable computer system in addition to PC systems.

[0112] Fig. Figure 10B illustrates a networked computer system 1000B according to at least one embodiment. In at least one embodiment, system 1000B illustrates a network, such as LAN 1024, which can be used to connect a variety of nodes that can communicate with each other. In at least one embodiment, a variety of nodes, such as PC nodes 1026, 1028, and 1030, are connected to LAN 1024. In at least one embodiment, a node can also be connected to the LAN via a network server or other means. In at least one embodiment, system 1000B includes other types of nodes or elements, for example, including routers, servers, and nodes.

[0113] Fig. Figure 10C illustrates a networked computer system 1000C according to at least one embodiment. In at least one embodiment, system 1000C illustrates a WWW system with communications over a backbone communication network (“Network”), such as the Internet, which can be used to connect a variety of nodes of a Network. In at least one embodiment, WWW is a set of protocols that operate over the Internet and enables a graphical interface system to work on it to access information over the Internet. In at least one embodiment, a variety of nodes, such as PCs 1040, 1042, and 1044, are connected to Network 1032 in WWW. In at least one embodiment, a node is connected via an interface to other nodes of WWW via a WWW HTTP server, such as servers 1034 and 1036.In at least one embodiment, PC 1044 can be a PC that forms a node of network 1032 and runs its own server 1036, even though PC 1044 and server 1036 are in . Fig. 10C are illustrated separately for illustrative purposes.

[0114] In at least one embodiment, WWW is a distributed type of application characterized by WWW HTTP, WWW's protocol, which runs on the Internet Transfer Control Protocol / Internet Protocol (“TCP / IP”). In at least one embodiment, WWW can thus be characterized by a set of protocols (i.e., HTTP) running on the Internet as its “backbone”.

[0115] In at least one embodiment, a web browser is an application running on a node of a network which, in WWW-compatible network systems, allows users of a single server or node to view such information and thus allows a user to browse graphical and text-based files linked together using hypertext links embedded in documents or files available from servers in a network that understand HTTP. In at least one embodiment, when a given web page from a first server associated with a first node is retrieved by a user using another server on a network such as the Internet, a retrieved document may have various hypertext links embedded within it, and a local copy of a page is created locally for the retrieving user.In at least one embodiment, when a user clicks on a hypertext link, locally stored information relating to a selected hypertext link is typically sufficient to allow a user machine to open a connection across the Internet to a server indicated by a hypertext link.

[0116] In at least one embodiment, more than one user can be connected to each HTTP server, for example, via a LAN, such as LAN 1038, as illustrated with respect to WWW-HTTP-Server 1034. In at least one embodiment, System 1000C can also include other types of nodes or elements. In at least one embodiment, a WWW-HTTP-Server is an application running on a machine, such as a PC. In at least one embodiment, each user can be considered to have a unique "server," as illustrated with respect to PC 1044. In at least one embodiment, a server can be considered to be a server, such as WWW-HTTP-Server 1034, that provides network access for a LAN or a plurality of nodes or a plurality of LANs.In at least one embodiment, there are multiple users, each with a desktop PC or network node, with each desktop PC potentially hosting a server for one of those users. In at least one embodiment, each server is associated with a single network address or URL, which, when accessed, provides a default web page for that user. In at least one embodiment, a web page can contain further links (embedded URLs) that point to further subpages for that user on that server, or to other servers in a network, or to pages on other servers in a network.

[0117] In at least one embodiment, various components of networked computer system 1000A-C can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5. Cloud computing and services

[0118] The following figures represent, without limitation, exemplary cloud-based systems that can be used to implement at least one embodiment.

[0119] In at least one embodiment, cloud computing is a style of computing in which dynamically scalable and often virtualized resources are provided as a service over the internet. In at least one embodiment, users do not need to have knowledge of, expertise in, or control over the technology infrastructure that can be described as "in the cloud" that supports them. In at least one embodiment, cloud computing includes infrastructure as a service, platform as a service, software as a service, and other variations that share a common theme of dependence on the internet to meet users' computing needs. In at least one embodiment, a typical cloud deployment, such as in a private cloud (e.g., a corporate network) or a data center (DC) in a public cloud (e.g., a community cloud), may be described as a cloud service.The cloud (Internet), consists of thousands of servers (or alternatively VMs), hundreds of Ethernet, Fibre Channel, or Fibre Channel over Ethernet (FCoE) ports, circuitry, and storage infrastructure, etc. In at least one embodiment, the cloud can also consist of network service infrastructure such as IPsec VPN hubs, firewalls, load balancers, wide area network (WAN) optimizers, etc. In at least one embodiment, remote participants can access cloud applications and services by connecting via a VPN tunnel, such as an IPsec VPN tunnel.

[0120] In at least one embodiment, cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction.

[0121] In at least one embodiment, cloud computing is characterized by on-demand self-service, where a consumer can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed, without requiring human interaction with each service provider. In at least one embodiment, cloud computing is characterized by broad network access, where capabilities are available over a network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).In at least one embodiment, cloud computing is characterized by resource pooling, where a provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with various physical and virtual resources being dynamically allocated and reallocated according to consumer demand. In at least one embodiment, a sense of location independence arises from the fact that a customer generally has no control over or knowledge of the precise location of provided resources but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center). In at least one embodiment, examples of resources include long-term storage, processing, short-term memory, network bandwidth, and virtual machines.In at least one embodiment, cloud computing is characterized by rapid elasticity, where capabilities can be provisioned and released quickly and elastically, in some cases automatically, to scale rapidly. In at least one embodiment, capabilities appear to be available for procurement to a consumer, often in unlimited quantities, and can be acquired at any time. In at least one embodiment, cloud computing is characterized by metered service, where cloud systems automatically control and optimize resource usage by exploiting a metering capability at a certain level of abstraction appropriate for a particular type of service (e.g., storage, processing, bandwidth, and active user accounts).In at least one embodiment, resource usage can be monitored, controlled and reported, thereby providing transparency for both a provider and a consumer of a used service.

[0122] In at least one embodiment, cloud computing can be associated with various services. In at least one embodiment, cloud software-as-a-service (SaaS) can be defined as a service where one capability provided to a consumer is the ability to use a provider's applications running on a cloud infrastructure. In at least one embodiment, applications are accessible from various client devices via a thin client interface, such as a web browser (e.g., web-based email). In at least one embodiment, the consumer does not manage or control any underlying cloud infrastructure, including network, servers, operating systems, storage, or even individual application capabilities, with a possible exception of limited user-specific application configuration settings.

[0123] In at least one embodiment, Cloud Platform-as-a-Service (PaaS) can refer to a service where a capability provided to a consumer is to deploy consumer-created or acquired applications on cloud infrastructure, created using programming languages ​​and tools supported by a provider. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but has control over deployed applications and potentially application host environment configurations.

[0124] In at least one embodiment, cloud infrastructure-as-a-service (IaaS) can refer to a service where a capability provided to a consumer consists of providing processing, storage, networking, and other basic computing resources, whereby a consumer is able to deploy and run any software, which may include operating systems and applications. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure but has control over operating systems, storage, deployed applications, and possibly limited control over selected network components (e.g., host firewalls).

[0125] In at least one embodiment, cloud computing can be used in various ways. In at least one embodiment, a private cloud can refer to a cloud infrastructure operated exclusively for one organization. In at least one embodiment, a private cloud can be managed by an organization or a third party and can exist on-premises or off-premises. In at least one embodiment, a community cloud can refer to a cloud infrastructure shared by multiple organizations and supporting a specific community with common interests (e.g., mission, security requirements, policy, and compliance considerations). In at least one embodiment, a community cloud can be managed by organizations or a third party and can exist on-premises.Cloud computing environments can exist on-premises or off-premises. In at least one embodiment, a public cloud can refer to cloud infrastructure made available to the general public or a large industry group and owned by an organization that provides cloud services. In at least one embodiment, a hybrid cloud can refer to cloud infrastructure that is a composition of two or more clouds (private, community, or public) that remain distinct entities but are linked by standardized or proprietary technology that enables data and application portability (e.g., cloud bursting for load balancing between clouds). In at least one embodiment, a cloud computing environment is service-oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability.

[0126] Fig. Figure 11 illustrates one or more components of a system environment 1100 in which services can be offered as third-party network services, according to at least one embodiment. In at least one embodiment, a third-party network can be referred to as a cloud, a cloud network, a cloud computing network, and / or variations thereof. In at least one embodiment, the system environment 1100 includes one or more client computing devices 1104, 1106, and 1108 that can be used by users to interact with a third-party network infrastructure system 1102 that provides third-party network services, which can be referred to as cloud computing services. In at least one embodiment, the third-party network infrastructure system 1102 can comprise one or more computers and / or servers.

[0127] It is understood that the third-party network infrastructure system 1102, which is in Fig. Figure 11 shows that it may include components other than those depicted. Furthermore, it shows Fig. 11 an embodiment of a third-party network infrastructure system. In at least one embodiment, the third-party network infrastructure system 1102 may comprise more or fewer components than in Fig. Figure 11 shows that it can combine two or more components or can include a different configuration or arrangement of components.

[0128] In at least one embodiment, client computing devices 1104, 1106, and 1108 can be configured to run a client application, such as a web browser, a proprietary client application, or another application that can be used by a user of a client computing device to interact with third-party network infrastructure system 1102 in order to use services provided by third-party network infrastructure system 1102. Although an exemplary system environment 1100 with three client computing devices is shown, any number of client computing devices can be supported. In at least one embodiment, other devices, such as devices with sensors, etc., can interact with third-party network infrastructure system 1102. In at least one embodiment, a network or networks 1110 can facilitate communication and data exchange between client computing devices 1104, 1106, and 1108 and third-party network infrastructure system 1102.

[0129] In at least one embodiment, services provided by the third-party network infrastructure system 1102 can include a host of services that are made available on-demand to users of the third-party network infrastructure system 1102. In at least one embodiment, various services can also be offered, including, without limitation, online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database management and processing, managed technical support services, and / or variations thereof. In at least one embodiment, services provided by the third-party network infrastructure system 1102 can dynamically scale to meet the needs of its users.

[0130] In at least one embodiment, a specific instantiation of a service provided by third-party network infrastructure system 1102 can be referred to as a "service instance." In at least one embodiment, any service provided over a communications network, such as the Internet, by a third-party network service provider's system is generally referred to as a "third-party network service." In at least one embodiment, in a public third-party network environment, the servers and systems that comprise a third-party network service provider's system are distinct from a consumer's own servers and systems on-site. In at least one embodiment, a third-party network service provider's system can host an application, and a user can order and use an application on-demand over a communications network, such as the Internet.

[0131] In at least one embodiment, a service in a third-party computer network infrastructure may include protected computer network access to storage, a hosted database, a hosted web server, a software application, or another service provided by a third-party network vendor to a user. In at least one embodiment, a service may include password-protected access to remote storage in a third-party network via the internet. In at least one embodiment, a service may include a web service-based hosted relational database and a scripting language middleware engine for private use by a networked developer. In at least one embodiment, a service may include access to an email software application hosted on a third-party network vendor's website.

[0132] In at least one embodiment, the third-party network infrastructure system 1102 can include a range of application, middleware, and database service offerings that are delivered to a customer in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. In at least one embodiment, the third-party network infrastructure system 1102 can also provide "big data"-related computing and analytics services. In at least one embodiment, the term "big data" is used generally to refer to extremely large datasets that can be stored and manipulated by analysts and researchers to visualize large volumes of data, detect trends, and / or otherwise interact with data. In at least one embodiment, big data and related applications can be hosted and / or manipulated by an infrastructure system at many levels and at different scales.In at least one embodiment, tens, hundreds, or thousands of processors linked in parallel can act on such data to present it or to simulate external forces on the data or what it represents. In at least one embodiment, these datasets can involve structured data, such as data organized in a database or otherwise according to a structured model, and / or unstructured data (e.g., emails, images, data blobs (large binary objects), web pages, complex event processing).In at least one embodiment, by exploiting a capability of an embodiment to relatively quickly focus more (or fewer) computing resources on a target, a third-party network infrastructure system can be made more available to perform tasks on large datasets based on demand from a company, government agency, research organization, private individual, group of like-minded individuals or organizations, or other entity.

[0133] In at least one embodiment, the third-party network infrastructure system 1102 can be designed to automatically provision, manage, and track a customer's subscription to services offered by the third-party network infrastructure system 1102. In at least one embodiment, the third-party network infrastructure system 1102 can provide third-party network services through various deployment models. In at least one embodiment, services can be provided under a public third-party network model in which the third-party network infrastructure system 1102 is owned by an organization that sells third-party network services, and services are made available to the general public or various industrial enterprises.In at least one embodiment, services can be provided under a private third-party network model, in which the third-party network infrastructure system 1102 is operated exclusively for a single organization and can provide services to one or more entities within that organization. In at least one embodiment, third-party network services can also be provided under a community third-party network model, in which the third-party network infrastructure system 1102 and the services provided by the third-party network infrastructure system 1102 are shared by multiple organizations in a related community. In at least one embodiment, third-party network services can also be provided under a hybrid third-party network model, which is a combination of two or more different models.

[0134] In at least one embodiment, services provided by the third-party network infrastructure system 1102 may include one or more services provided under the Software-as-a-Service (SaaS) category, Platform-as-a-Service (PaaS) category, Infrastructure-as-a-Service (IaaS) category, or other categories of services, including hybrid services. In at least one embodiment, a customer may order one or more services provided by the third-party network infrastructure system 1102 via a subscription order. In at least one embodiment, the third-party network infrastructure system 1102 then performs processing to deliver services in a customer's subscription order.

[0135] In at least one embodiment, services provided by third-party network infrastructure system 1102 can include, without limitation, application services, platform services, and infrastructure services. In at least one embodiment, application services can be provided by a third-party network infrastructure system via a SaaS platform. In at least one embodiment, the SaaS platform can be configured to provide third-party network services that fall under a SaaS category. In at least one embodiment, the SaaS platform can provide capabilities to build and deliver a suite of on-demand applications on an integrated development and deployment platform. In at least one embodiment, the SaaS platform can manage and control the underlying software and infrastructure for delivering SaaS services.In at least one embodiment, customers can use applications running on a third-party network infrastructure system by utilizing services provided by a SaaS platform. In at least one embodiment, customers can acquire application services without having to purchase separate licenses and support. In at least one embodiment, various different SaaS services can be provided. In at least one embodiment, examples include, without limitation, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations.

[0136] In at least one embodiment, platform services can be provided by a third-party network infrastructure system 1102 via a PaaS platform. In at least one embodiment, the PaaS platform can be configured to provide third-party network services that fall under a PaaS category. In at least one embodiment, examples of platform services can, without limitation, include services that enable organizations to consolidate existing applications on a common, shared architecture, as well as enabling the development of new applications that utilize shared services provided by a platform. In at least one embodiment, the PaaS platform can manage and control the underlying software and infrastructure for providing PaaS services.In at least one embodiment, customers can acquire PaaS services provided by third-party network infrastructure system 1102 without having to purchase separate licenses and support.

[0137] In at least one embodiment, customers can use programming languages ​​and tools supported by a third-party network infrastructure system and control the deployed services by using services provided by a PaaS platform. In at least one embodiment, platform services provided by a third-party network infrastructure system can include third-party database network services, third-party middleware network services, and third-party network services. In at least one embodiment, third-party database network services can support shared service deployment models that allow organizations to pool database resources and offer customers a database as a service in the form of a third-party database network.In at least one embodiment, third-party middleware network services can provide a platform for customers to develop and deploy various business applications, and third-party network services can provide a platform for customers to deploy applications in a third-party network infrastructure system.

[0138] In at least one embodiment, various different infrastructure services can be provided by an IaaS platform within a third-party network infrastructure system. In at least one embodiment, infrastructure services facilitate the management and control of underlying computing resources, such as storage, networks, and other fundamental computing resources, for customers using services provided by a SaaS platform and a PaaS platform.

[0139] In at least one embodiment, the third-party network infrastructure system 1102 can also include infrastructure resources 1130 for providing resources used to deliver various services to customers of a third-party network infrastructure system. In at least one embodiment, the infrastructure resources 1130 can include pre-integrated and optimized combinations of hardware, such as server, storage, and network resources for running services provided by a PaaS platform and a SaaS platform, and other resources.

[0140] In at least one embodiment, resources in third-party network infrastructure system 1102 can be shared by multiple users and dynamically reassigned as needed. In at least one embodiment, resources can be assigned to users in different time zones. In at least one embodiment, third-party network infrastructure system 1102 can allow a first set of users in a first time zone to use resources of a third-party network infrastructure system for a specified number of hours, and then allow the same resources to be reassigned to another set of users located in a different time zone, thereby maximizing resource utilization.

[0141] In at least one embodiment, a number of internal shared services 1132 can be provided, which are shared by various components or modules of the third-party network infrastructure system 1102 to enable the third-party network infrastructure system 1102 to provide services. In at least one embodiment, these internal shared services can include, without limitation, a security and identity service, an integration service, an enterprise repository service, an enterprise manager service, a virus scanning and whitelisting service, a high availability, backup and recovery service, a service to enable third-party network support, an email service, a notification service, a file transfer service, and / or variations thereof.

[0142] In at least one embodiment, the third-party network infrastructure system 1102 can provide comprehensive management of third-party network services (e.g., SaaS, PaaS, and IaaS services) within the third-party network infrastructure system. In at least one embodiment, the third-party network management functionality can include capabilities for provisioning, managing, and tracking a customer subscription received by the third-party network infrastructure system 1102 and / or variations thereof.

[0143] In at least one embodiment, as in Fig. As shown in Figure 11, third-party network management functionality can be provided by one or more modules, such as an order management module 1120, an order orchestration module 1122, an order delivery module 1124, an order management and monitoring module 1126, and an identity management module 1128. In at least one embodiment, these modules can include or be provided using one or more computers and / or servers, which may be general-purpose computers, specialized server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0144] In at least one embodiment, in step 1134, a customer using a client device, such as client computing devices 1104, 1106, or 1108, can interact with third-party network infrastructure system 1102 by requesting one or more services provided by third-party network infrastructure system 1102 and placing an order for a subscription to one or more services offered by third-party network infrastructure system 1102. In at least one embodiment, a customer can access a third-party network user interface (UI), such as third-party network UI 1112, third-party network UI 1114, and / or third-party network UI 1116, and place a subscription order through these UIs.In at least one embodiment, order information received by the third-party network infrastructure system 1102 in response to a customer placing an order may include information that identifies a customer and one or more services offered by the third-party network infrastructure system 1102 that the customer wishes to subscribe to.

[0145] In at least one embodiment, in step 1136, order information received from a customer can be stored in an order database 1118. In at least one embodiment, if this is a new order, a new entry for the order can be created. In at least one embodiment, order database 1118 can be one of several databases operated by a third-party network infrastructure system 1102 and operated in conjunction with other system elements.

[0146] In at least one embodiment, at step 1138, order information can be forwarded to an order management module 1120, which can be configured to perform billing and accounting functions in relation to an order, such as verifying an order and, upon verification, posting an order.

[0147] In at least one embodiment, at step 1140, information regarding an order can be communicated to an order orchestration module 1122, which is configured to orchestrate the provisioning of services and resources for an order placed by a customer. In at least one embodiment, order orchestration module 1122 can use services from order provisioning module 1124 for provisioning. In at least one embodiment, order orchestration module 1122 enables the management of business processes associated with each order and applies business logic to determine whether an order should proceed with provisioning.

[0148] In at least one embodiment, at step 1142, after receiving an order for a new subscription, order orchestration module 1122 sends a request to order provisioning module 1124 to allocate and configure resources needed to fulfill a subscription order. In at least one embodiment, order provisioning module 1124 enables the allocation of resources for services ordered by a customer. In at least one embodiment, order provisioning module 1124 provides an abstraction layer between third-party network services provided by third-party network infrastructure system 1100 and a physical implementation layer used to provision resources for delivering requested services.In at least one embodiment, this allows order orchestration module 1122 to be isolated from implementation details, such as whether services and resources are actually provided in real time or in advance and only allocated / assigned on demand.

[0149] In at least one embodiment, at step 1144, once services and resources are provisioned, a notification can be sent to subscribing customers indicating that a requested service is now ready for use. In at least one embodiment, information (e.g., a link) can be sent to a customer that enables the customer to begin using the requested services.

[0150] In at least one embodiment, in step 1146, a customer subscription order can be managed and tracked by an order management and monitoring module 1126. In at least one embodiment, the order management and monitoring module 1126 can be configured to collect usage statistics regarding customer use of subscribed services. In at least one embodiment, statistics can be collected for the amount of memory used, the amount of data transferred, the number of users, and the amount of system up time and system down time, and / or variations thereof.

[0151] In at least one embodiment, the third-party network infrastructure system 1100 can include an identity management module 1128 configured to provide identity services, such as access management and authorization services, within the third-party network infrastructure system 1100. In at least one embodiment, the identity management module 1128 can manage information about customers who wish to use services provided by the third-party network infrastructure system 1102. In at least one embodiment, such information can include information that authenticates the identities of such customers and information that describes which actions these customers are authorized to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.).In at least one embodiment, identity management module 1128 can also include the management of descriptive information about each customer and how and by whom this descriptive information can be accessed and modified.

[0152] In at least one embodiment, various components of the network environment 1100 can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0153] Fig. Figure 12 illustrates a cloud computing environment 1202 according to at least one embodiment. In at least one embodiment, the cloud computing environment 1202 comprises one or more computer systems / servers 1204 with which computing devices such as a personal digital assistant (PDA) or mobile phone 1206A, desktop computer 1206B, laptop computer 1206C, and / or automotive computer system 1206N communicate. In at least one embodiment, this allows infrastructure, platforms, and / or software to be offered as services of the cloud computing environment 1202, thus avoiding the need for each client to maintain such resources separately. It is understood that in Fig. The 12 types of computing devices shown in 1206A-N are intended to be illustrative only, and the Cloud Computing Environment 1202 can communicate with any type of computer-based device over any type of network and / or network / addressable connection (e.g., using a web browser).

[0154] In at least one embodiment, a Computer System / Server 1204, which may be referred to as a cloud computing node, is compatible with numerous other general-purpose or special-purpose computing environments or configurations. In at least one embodiment, examples of computing systems, environments, and / or configurations suitable for use with Computer System / Server 1204 include personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments, including any of the above systems or devices, and / or variations thereof.

[0155] In at least one embodiment, Computer System / Server 1204 can be described in a general context of computer system executable instructions, such as program modules, that are executed by a computer system. In at least one embodiment, program modules include routines, programs, objects, components, logic, data structures, and so on, which perform specific tasks or implement individual abstract data types. In at least one embodiment, Exemplary Computer System / Server 1204 can be used in distributed cloud computing environments in which tasks are performed by remote processing devices linked by a communication network. In at least one embodiment, in a distributed cloud computing environment, program modules can reside in both local and remote computer system storage media, including storage devices.

[0156] In at least one embodiment, various components of the cloud computing environment 1202 can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0157] Fig. 13 illustrates a set of functional abstraction layers used by cloud computing environment 1202 ( Fig. 12) shall be provided, according to at least one embodiment. It should be understood in advance that in Fig. The 13 components, layers and functions shown are for illustrative purposes only and components, layers and functions may vary.

[0158] In at least one embodiment, the hardware and software layer 1302 comprises hardware and software components. Examples of hardware components in at least one embodiment include mainframes, various RISC (Reduced Instruction Set Computer) architecture-based servers, various computing systems, supercomputers, storage devices, networks, network components, and / or variations thereof. Examples of software components in at least one embodiment include network application server software, various application server software, various database software, and / or variations thereof.

[0159] In at least one embodiment, virtualization layer 1304 provides an abstraction layer from which the following exemplary virtual entities can be provided: virtual servers, virtual storage, virtual networks, including virtual private networks, virtual applications, virtual clients and / or variations thereof.

[0160] In at least one embodiment, the management layer 1306 provides various functions. In at least one embodiment, resource provisioning provides dynamic procurement of compute resources and other resources needed to perform tasks within a cloud computing environment. In at least one embodiment, metering provides usage tracking while resources are used within a cloud computing environment and billing for the consumption of these resources. In at least one embodiment, resources can have application software licenses. In at least one embodiment, security provides identity verification for users and tasks, as well as protection for data and other resources. In at least one embodiment, the user interface provides access to a cloud computing environment for both users and system administrators.In at least one embodiment, service-level management provides cloud computing resource allocation and management to meet required service levels. In at least one embodiment, service-level agreement (SLA) management provides pre-planning and procurement of cloud computing resources for which a future requirement is anticipated according to an SLA.

[0161] In at least one embodiment, workload layer 1308 provides functionality that requires a cloud computing environment. Examples of workloads and functions that can be provided by this layer in at least one embodiment include: mapping and navigation, software development and management, educational services, data analysis and processing, transaction processing, and service delivery. Supercomputing

[0162] The following figures represent, without limitation, exemplary supercomputer-based systems that can be used to implement at least one embodiment.

[0163] In at least one embodiment, a supercomputer can refer to a hardware system exhibiting substantial parallelism and comprising at least one chip, wherein chips in a system are interconnected by a network and placed in hierarchically organized enclosures. In at least one embodiment, a large hardware system filling a machine room, with multiple racks, each containing multiple boards / rack modules, each containing multiple chips, all interconnected by a scalable network, is a single example of a supercomputer. In at least one embodiment, a single rack of such a large hardware system is another example of a supercomputer.In at least one embodiment, a single chip exhibiting substantial parallelism and containing multiple hardware components can equally be considered a supercomputer, since if feature sizes can decrease, the amount of hardware that can be incorporated into a single chip can also increase.

[0164] Fig. Figure 14 illustrates a chip-level supercomputer according to at least one embodiment. In at least one embodiment, within an FPGA or ASIC chip, main computation is performed within finite state machines (904) called thread units. In at least one embodiment, task and synchronization networks (902) connect finite state machines and are used to dispatch threads and execute operations in the correct order. In at least one embodiment, a multi-level partitioned on-chip cache hierarchy (908, 1412) is accessed using memory networks (906, 1410). In at least one embodiment, off-chip memory is accessed using memory controllers (916) and an off-chip memory network (914).In at least one embodiment, I / O controller (918) is used for cross-chip communication when a design does not fit into a single logic chip.

[0165] Fig. Figure 15 illustrates a supercomputer at the rack module level according to at least one embodiment. In at least one embodiment, within a rack module, there are multiple FPGA or ASIC chips (1002) connected to one or more DRAM units (1004) that constitute the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its adjacent FPGA / ASIC chip using wide buses on a board with differential high-speed signaling (1006). In at least one embodiment, each FPGA / ASIC chip is also connected to at least one serial high-speed communication cable.

[0166] Fig. Figure 16 illustrates a rack-level supercomputer according to at least one embodiment. Fig. Figure 17 illustrates a supercomputer at the overall system level according to at least one embodiment. In at least one embodiment, with reference to Fig. 16 and Fig. 17 High-speed serial optical or copper cables (1102, 1702) are used between rack modules in a rack and across racks in an entire system to implement a scalable, potentially incomplete hypercube network. In at least one embodiment, one of the FPGA / ASIC chips of an accelerator is connected to a host system via a PCI Express connection (1204). In at least one embodiment, the host system includes a host microprocessor (1208) on which a software portion of an application runs, and memory consisting of one or more host memory DRAM units (1206) that are kept coherent with memory on an accelerator. In at least one embodiment, the host system can be a separate module on one of the racks or can be integrated into one of the modules of a supercomputer.In at least one embodiment, Cube-Connected Cycles topology provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, a small group of FPGA / ASIC chips on a rack module can act as a single hypercube node, thus increasing the total number of external links for each group compared to a single chip. In at least one embodiment, a group comprises chips A, B, C, and D on a rack module with internal wide differential buses that connect A, B, C, and D in a torus organization. In at least one embodiment, there are 17 serial communication cables connecting a rack module to the outside world. In at least one embodiment, chip A on a rack module is connected to serial communication cables 0, 1, and 2. In at least one embodiment, chip B is connected to cables 3, 4, and 5.In at least one embodiment, chip C is connected to 6, 7, 8. In at least one embodiment, chip D is connected to 9, 15, 16. In at least one embodiment, an entire group {A, B, C, D} constituting a rack module can form a hypercube node within a supercomputer system with up to 212 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, for chip A to send a message on connection 4 of the group {A, B, C, D}, a message must first be routed to chip B via an on-board differential wide bus connection. In at least one embodiment, a message arriving at connection 4 (i.e., arriving at B) in a group {A, B, C, D} and destined for chip A must also first be routed to a correct destination chip (A) internally within the group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes can also be implemented.

[0167] In at least one embodiment, supercomputers located in Fig. Figures 14-17 show that the computer is housed in a data center which can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5. Artificial intelligence

[0168] The following figures represent, without limitation, exemplary artificial intelligence-based systems that can be used to implement at least one embodiment.

[0169] Fig. Figure 18A illustrates inference and / or training logic 1815, which is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1815 are given below in conjunction with Fig. 18A and / or 18B provided.

[0170] In at least one embodiment, inference and / or training logic 1815 can, without limitation, include code and / or data storage 1801 for storing forward and / or output weights and / or input / output data and / or other parameters for configuring neurons or layers of a neural network that are trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, training logic 1815 can include or be coupled to code and / or data storage 1801 to store graph code or other software for controlling timing and / or sequencing, into which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, code and / or data storage 1801 stores weight parameters and / or input / output data of each layer of a neural network, trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 1801 may be contained in another on-chip or off-chip data storage, including L1, L2, or L3 cache or system memory of a processor.

[0171] In at least one embodiment, any portion of the code and / or data memory 1801 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data memory 1801 can be cache memory, dynamic directly addressable memory (“DRAM”), static directly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether code and / or code and / or data storage 1801 is, for example, inside or outside a processor, or comprises DRAM, SRAM, Flash or another type of memory, may depend on available on-chip versus off-chip memory, latency requirements of training and / or inference functions performed, batch size of data used in inferencing and / or training a neural network, or a combination of these factors.

[0172] In at least one embodiment, inference and / or training logic 1815 can, without limitation, include code and / or data storage 1805 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network that are trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 1805 stores weight parameters and / or input / output data of each layer of a neural network, trained or used in conjunction with one or more embodiments, during backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, training logic 1815 may include or be coupled to code and / or data storage 1805 to store graph code or other software for controlling timing and / or sequencing, into which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).

[0173] In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, any portion of code and / or data memory 1805 can be contained in another on-chip or off-chip data memory, including L1, L2, or L3 cache or system memory of a processor. In at least one embodiment, any portion of code and / or data memory 1805 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data memory 1805 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether code and / or data storage 1805 is, for example, inside or outside a processor, or whether it comprises DRAM, SRAM, Flash or another type of memory, may depend on available on-chip versus off-chip memory, latency requirements of training and / or inference functions performed, batch size of data used in inferencing and / or training a neural network, or a combination of these factors.

[0174] In at least one embodiment, code and / or data memory 1801 and code and / or data memory 1805 can be separate memory structures. In at least one embodiment, code and / or data memory 1801 and code and / or data memory 1805 can be a combined memory structure. In at least one embodiment, code and / or data memory 1801 and code and / or data memory 1805 can be partially combined and partially separate. In at least one embodiment, any portion of code and / or data memory 1801 and code and / or data memory 1805 can be contained in another on-chip or off-chip data memory, including L1, L2, or L3 cache or system memory of a processor.

[0175] In at least one embodiment, inference and / or training logic 1815 may, without limitation, include one or more arithmetic logic unit(s) (“ALU(s)”) 1810, including integer and / or floating-point units, to perform logical and / or mathematical operations that are at least partially based on or indicated by training and / or inference code (e.g., graph code), wherein a result thereof may produce activations (e.g., output values ​​of layers or neurons within a neural network) that are stored in an activation memory 1820, which are functions of input / output and / or weight parameter data that are stored in code and / or data memory 1801 and / or code and / or data memory 1805.In at least one embodiment, activations stored in activation memory 1820 are generated according to linear algebraic and / or matrix-based mathematics performed by ALU(s) 1810 in response to the execution of instructions or other code, wherein weight values ​​stored in code and / or data memory 1805 and / or data memory 1801 are used as operands together with other values, such as bias values, gradient information, pulse values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data memory 1805 or code and / or data memory 1801 or another on- or off-chip memory.

[0176] In at least one embodiment, ALU(s) 1810 are contained within one or more processors or other hardware logic devices or circuits, whereas in another embodiment, ALU(s) 1810 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, ALU(s) 1810 may be contained within the execution units of a processor or otherwise within a bank of ALUs that can be accessed by the execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, code and / or data memory 1801, code and / or data memory 1805, and activation memory 1820 can share a processor or other hardware logic device or circuit, whereas in another embodiment, they can be located in different processors or other hardware logic devices or circuits, or a combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation memory 1820 can be contained in another on-chip or off-chip data memory, including L1, L2, or L3 cache or system memory of a processor.Furthermore, inference and / or training code can be stored with other code that is accessible to a processor or other hardware logic or circuitry and can be retrieved and / or processed using retrieval, decoding, scheduling, execution, shutdown, and / or other logic circuitry of a processor.

[0177] In at least one embodiment, the activation memory 1820 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the activation memory 1820 can be located wholly or partially inside or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation memory 1820 is located inside or external to a processor, or whether it comprises DRAM, SRAM, flash memory, or another type of memory, can depend on available on-chip versus off-chip memory, the latency requirements of training and / or inference functions performed, the batch size of data used in inferencing and / or training a neural network, or a combination of these factors.

[0178] In at least one embodiment, inference and / or training logic 1815, which is in Fig. 18A is illustrated, in conjunction with an application-specific integrated circuit (“ASIC”), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, inference and / or training logic 1815, which is illustrated in Fig. Figure 18A illustrates how they can be used in conjunction with hardware of a central processing unit (“CPU”), hardware of a graphics processing unit (“GPU”) or other hardware, such as field-programmable gate arrays (“FPGAs”).

[0179] Fig. Figure 18B illustrates inference and / or training logic 1815 according to at least one embodiment. In at least one embodiment, inference and / or training logic 1815 can, without limitation, include hardware logic in which computing resources are dedicated or otherwise used exclusively in connection with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, inference and / or training logic 1815, which is described in Fig. 18B is illustrated, in conjunction with an application-specific integrated circuit (ASIC), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, inference and / or training logic 1815, which is illustrated in Fig. Figure 18B illustrates the use of inference and / or training logic 1815 in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware, such as field-programmable gate arrays (FPGAs). In at least one embodiment, inference and / or training logic 1815 includes, without limitation, code and / or data memory 1801 and code and / or data memory 1805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, pulse values, and / or other parameter or hyperparameter information. In at least one embodiment, which is illustrated in Figure 18B, the inference and / or training logic 1815 includes code and / or data memory 1801 and code and / or data memory 1805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, pulse values, and / or other parameter or hyperparameter information. Fig. As illustrated in Figure 18B, each of code and / or data memory 1801 and code and / or data memory 1805 is associated with a dedicated computing resource, such as computing hardware 1802 and computing hardware 1806, respectively. In at least one embodiment, each of computing hardware 1802 and computing hardware 1806 has one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in code and / or data memory 1801 and code and / or data memory 1805, respectively, the result of which is stored in activation memory 1820.

[0180] In at least one embodiment, each of the code and / or data storage units 1801 and 1805 and the corresponding computing hardware 1802 and 1806, respectively, corresponds to different layers of a neural network, such that the resulting activation from one memory / computing pair 1801 / 1302 of code and / or data storage unit 1801 and computing hardware 1802 is provided as an input for the next memory / computing pair 1805 / 1306 of code and / or data storage unit 1805 and computing hardware 1806, in order to reflect a conceptual organization of a neural network. In at least one embodiment, each of the memory / computing pairs 1801 / 1302 and 1805 / 1306 can correspond to more than one neural network layer. In at least one embodiment, additional memory / computing pairs (not shown) may be included after or in parallel to the memory / computing pairs 1801 / 1302 and 1805 / 1306 in inference and / or training logic 1815.

[0181] Fig. Figure 19 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 1906 is trained using a set of training data 1902. In at least one embodiment, the training frame 1904 is a PyTorch frame, whereas in other embodiments, the training frame 1904 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frame. In at least one embodiment, the training frame 1904 trains an untrained neural network 1906 and allows it to be trained using processing resources described herein to generate a trained neural network 1908. In at least one embodiment, weights can be selected randomly or by pretraining using a deep belief network.In at least one embodiment, training can be carried out in a supervised, partially supervised, or unsupervised manner.

[0182] In at least one embodiment, an untrained neural network 1906 is trained using supervised learning, wherein the set of training data 1902 contains an input paired with a desired output for that input, or wherein the set of training data 1902 contains an input paired with a known output, and an output of the neural network 1906 is manually graded. In at least one embodiment, the untrained neural network 1906 is trained in a supervised manner and processes inputs from the set of training data 1902 and compares resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are then propagated back by the untrained neural network 1906. In at least one embodiment, the training frame 1904 adjusts weights that control the untrained neural network 1906.In at least one embodiment, training frame 1904 includes tools to monitor how well an untrained neural network 1906 converges to a model, such as a trained neural network 1908, that is capable of producing correct answers, as in result 1914, based on input data such as a new dataset 1912. In at least one embodiment, training frame 1904 repeatedly trains the untrained neural network 1906 while adjusting weights to refine an output from the untrained neural network 1906 using a loss function and a fitting algorithm, such as stochastic gradient descent. In at least one embodiment, training frame 1904 trains the untrained neural network 1906 until the untrained neural network 1906 achieves a desired accuracy.In at least one embodiment, the trained neural network can then be used to implement any number of machine learning operations.

[0183] In at least one embodiment, an untrained neural network 1906 is trained using unsupervised learning, wherein the untrained neural network 1906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised training set 1902 contains input data without any associated output data or ground-truth data. In at least one embodiment, the untrained neural network 1906 can learn groupings within the training set 1902 and can determine how individual inputs relate to the training set 1902. In at least one embodiment, unsupervised training can be used to generate a self-organizing mapping in the trained neural network 1908, which is capable of performing operations useful for reducing the dimensionality of the new dataset 1912.In at least one embodiment, unsupervised training can also be used to perform anomaly detection, allowing the identification of data points in new data set 1912 that deviate from normal patterns of new data set 1912.

[0184] In at least one embodiment, semi-supervised learning can be used, which is a technique in which the set of training data 1902 contains a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 1904 can be used to perform incremental learning, such as through transferred learning techniques. In at least one embodiment, incremental learning allows the trained neural network 1908 to adapt to new data 1912 without forgetting knowledge that was fed into the trained neural network 1908 during initial training.

[0185] In at least one embodiment, neural networks that are in Fig. Figures 18-19 are shown to be used to predict thermal conditions, as in Fig. 4-5 discussed. 5G networks

[0186] The following figures represent, without limitation, exemplary 5G network-based systems that can be used to implement at least one embodiment.

[0187] Fig. Figure 20 illustrates an architecture of a System 2000 network according to at least one embodiment. In at least one embodiment, System 2000 is shown to include a User Device (UE) 2002 and a UE 2004. In at least one embodiment, UEs 2002 and 2004 are illustrated as smartphones (e.g., handheld mobile touchscreen computing devices that can connect to one or more cellular networks), but they can also include any mobile or non-mobile computing device, such as Personal Data Assistants (PDAs), pagers, laptop computers, desktop computers, wireless handsets, or any computing device that includes a wireless communication interface.

[0188] In at least one embodiment, each of the UEs 2002 and 2004 can include an Internet of Things (IoT) UE, which may have a network access layer designed for low-power IoT applications using short-lived UE connections. In at least one embodiment, an IoT UE can use technologies such as machine-to-machine (M2M) or machine-type communication (MTC) to exchange data with an MTC server or device over a public terrestrial mobile network (PLMN), proximity-based service (ProSe) or device-to-device (D2D) communication, sensor networks, or IoT networks. In at least one embodiment, an M2M or MTC data exchange can be a machine-initiated data exchange.In at least one embodiment, an IoT network describes interconnected IoT UEs that may include uniquely identifiable embedded computing devices (within the internet infrastructure) with short-lived connections. In at least one embodiment, IoT UEs may execute background applications (e.g., keep-alive messages, status updates, etc.) to facilitate connections within an IoT network.

[0189] In at least one embodiment, UEs 2002 and 2004 can be configured to connect to a radio access network (RAN) 2016, for example, to communicate. In at least one embodiment, RAN 2016 can be, for example, an Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN), a NextGen RAN (NG RAN), or another type of RAN. In at least one embodiment, UEs 2002 and 2004 each use connections 2012 and 2014, respectively, each of which has a physical communication interface or layer.In at least one embodiment, connections 2012 and 2014 are illustrated as an air interface to enable communicative coupling and can be consistent with cellular communication protocols such as a Global System for Mobile Communications (GSM) protocol, a Code Duplex Multi-Access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a Fifth Generation (5G) protocol, a New Radio (NR) protocol and variations thereof.

[0190] In at least one embodiment, UEs 2002 and 2004 can also directly exchange communication data via a ProSe interface 2006. Alternatively, in at least one embodiment, the ProSe interface 2006 can be described as a sidelink interface comprising one or more logical channels, including, but not limited to, a physical sidelink control channel (PSCCH), a physical shared sidelink channel (PSSCH), a physical sidelink discovery channel (PSDCH), and a physical sidelink broadcast channel (PSBCH).

[0191] In at least one embodiment, it is shown that UE 2004 is configured to access access point (AP) 2010 via connection 2008. In at least one embodiment, connection 2008 can have a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, where AP 2010 would have a Wireless Fidelity (WiFi®) router. In at least one embodiment, it is shown that AP 2010 is connected to the internet without connecting to a core network of a wireless system.

[0192] In at least one embodiment, RAN 2016 can include one or more access nodes enabling connections 2012 and 2014. In at least one embodiment, these access nodes (ANs) can be designated as base stations (BSs), NodeBs, evolved NodeBs (eNBs), next-generation NodeBs (gNBs), RAN nodes, and so forth, and can include ground stations (e.g., terrestrial access points) or satellite stations that provide coverage within a geographic area (e.g., a cell). In at least one embodiment, RAN 2016 can include one or more RAN nodes for providing macrocells, e.g., Macro-RAN Node 2018, and one or more RAN nodes for providing femtocells or picocells (e.g., cells with smaller coverage areas, smaller user capacity, or higher bandwidth compared to macrocells). B. Low Power (LP) RAN Node 2020.

[0193] In at least one embodiment, each of the RAN 2018 and 2020 nodes can terminate an air interface protocol and can be a first contact point for UEs 2002 and 2004. In at least one embodiment, each of the RAN 2018 and 2020 nodes can perform various logical functions for RAN 2016, including, but not limited to, radio network control (RNC) functions such as radio support management, dynamic uplink and downlink radio resource management, and data packet scheduling and mobility management.

[0194] In at least one embodiment, UEs 2002 and 2004 can be configured to communicate with each other or with any of the RAN nodes 2018 and 2020 over a multi-carrier communication channel using orthogonal frequency-division multiplexing (OFDM) communication signals, according to various communication techniques, such as, but not limited to, orthogonal frequency-division multiple access (OFDMA) communication techniques (e.g., for downlink communications) or single-carrier frequency-division multiplexing (SC-FDMA) communication techniques (e.g., for uplink and ProSe or sidelink communications) and / or variations thereof. In at least one embodiment, OFDM signals can have a plurality of orthogonal subcarriers.

[0195] In at least one embodiment, a downlink resource grid can be used for downlink transmissions from each of the RAN nodes 2018 and 2020 to UEs 2002 and 2004, while uplink transmissions can use similar techniques. In at least one embodiment, a grid can be a time-frequency grid, called a resource grid or time-frequency resource grid, which is a physical resource in a downlink in each slot. In at least one embodiment, such a time-frequency plane representation is common practice for OFDM systems, making it intuitive for radio resource allocation. In at least one embodiment, each column and each row of a resource grid corresponds to an OFDM symbol and an OFDM subcarrier, respectively. In at least one embodiment, a duration of a resource grid in a time domain corresponds to a slot in a radio frame.In at least one embodiment, the smallest time-frequency unit in a resource grid is referred to as a resource element. In at least one embodiment, each resource grid has a number of resource blocks that describe a mapping of certain physical channels to resource elements. In at least one embodiment, each resource block has a collection of resource elements. In at least one embodiment, this can represent a smallest set of resources that can currently be allocated in a frequency domain. In at least one embodiment, there are several different physical downlink channels that are communicated using such resource blocks.

[0196] In at least one embodiment, a physical shared downlink channel (PDSCH) can carry user data and higher-layer signaling to UEs 2002 and 2004. In at least one embodiment, a physical downlink control channel (PDCCH) can carry information about a transport format and resource allocations relating, among other things, to the PDSCH channel. In at least one embodiment, it can also inform UEs 2002 and 2004 about a transport format, resource allocation, and HARQ (Hybrid Automatic Repeat Request) information relating to a shared uplink channel.In at least one embodiment, downlink planning (allocating control and shared channel resource blocks to UE 2002 within a cell) can typically be performed at each of the RAN nodes 2018 and 2020, based on channel quality information fed back from each of the UEs 2002 and 2004. In at least one embodiment, downlink resource allocation information can be sent on a PDCCH that is used (e.g., allocated) for each of the UEs 2002 and 2004.

[0197] In at least one embodiment, a PDCCH can use control channel elements (CCEs) to communicate control information. In at least one embodiment, before mapping to resource elements, complex-valued PDCCH symbols can first be organized into quadruples, which can then be permuted using a subblock nester for rate matching. In at least one embodiment, each PDCCH can be transmitted using one or more of these CCEs, each CCE being able to correspond to nine sets of four physical resource elements known as resource element groups (REGs). In at least one embodiment, four quadrature phase shift keying (QPSK) symbols can be mapped to each REG. In at least one embodiment, PDCCH can be transmitted using one or more CCEs, depending on the size of a downlink control information (DCI) and a channel condition.In at least one embodiment, there can be four or more different PDCCH formats defined in LTE with different numbers of CCEs (e.g., aggregation level, L = 1, 2, 4 or 8).

[0198] In at least one embodiment, an extended physical downlink control channel (EPDCCH) utilizing PDSCH resources can be used for transmitting control information. In at least one embodiment, EPDCCH can be transmitted using one or more extended control channel elements (ECCEs). In at least one embodiment, each ECCE can correspond to nine sets of four physical resource elements known as extended resource element groups (EREGs). In at least one embodiment, an ECCE can, in some situations, comprise a different number of EREGs.

[0199] In at least one embodiment, RAN 2016 is communicatively coupled to a core network (CN) 2038 via an S1 interface 2022. In at least one embodiment, CN 2038 can be an Evolved Packet Core (EPC) network, a NextGen Packet Core (NPC) network, or another type of CN. In at least one embodiment, the S1 interface 2022 is divided into two parts: an S1-U interface 2026, which carries traffic data between RAN nodes 2018 and 2020 and the Serving Gateway (S-GW) 2030, and an S1 Mobility Management Entity (MME) interface 2024, which is a signaling interface between RAN nodes 2018 and 2020 and the MME(s) 2028.

[0200] In at least one embodiment, CN 2038 includes MME(s) 2028, S-GW 2030, Packet Data Network (PDN) Gateway (P-GW) 2034, and a Home Subscriber Server (HSS) 2032. In at least one embodiment, MME(s) 2028 may function similarly to a control plane of Legacy Serving General Packet Radio Service (GPRS) Support Nodes (SGSN). In at least one embodiment, MME 2028(s) or MME(s) 2028 may manage mobility aspects of access, such as gateway selection and tracking area list management. In at least one embodiment, HSS 2032 may include a database for network users, including subscription-related information, to support the handling of communication sessions by network entities. In at least one embodiment, CN 2038 can include one or more HSSs 2032, depending on the number of mobile subscribers, the capacity of an equipment, the organization of a network, etc.In at least one embodiment, HSS 2032 can provide support for routing / roaming, authentication, authorization, naming / addressing resolution, location dependencies, etc.

[0201] In at least one embodiment, S-GW 2030 can terminate an S1 interface 2022 to RAN 2016 and route data packets between RAN 2016 and CN 2038. In at least one embodiment, S-GW 2030 can be a local mobility anchor point for inter-RAN node handoffs and can also provide an anchor for inter-3GPP mobility. In at least one embodiment, other responsibilities can include lawful interception, charging, and certain policy enforcement.

[0202] In at least one embodiment, P-GW 2034 can terminate an SGi interface to a PDN. In at least one embodiment, P-GW 2034 can route data packets between CN 2038 and external networks, such as a network containing an application server 2040 (alternatively referred to as an application function (AF)), via an Internet Protocol (IP) interface 2042. In at least one embodiment, the application server 2040 can be an element that provides applications that use IP support resources with a core network (e.g., a UMTS Packet Services (PS) domain, LTE PS data services, etc.). In at least one embodiment, P-GW 2034 is shown to be communicatively coupled to an application server 2040 via an IP communication interface 2042. In at least one embodiment, the application server 2040 can also be configured to provide one or more communication services (e.g.,to support Voice-over-Internet Protocol (VoIP) sessions, PTT sessions, group communication sessions, social networking services, etc.) for UEs 2002 and 2004 via CN 2038.

[0203] In at least one embodiment, P-GW 2034 can further be a policy enforcement and charge data collection node. In at least one embodiment, the policy and charge enforcement function (PCRF) 2036 is a policy and charge control element of CN 2038. In at least one embodiment, in a non-roaming scenario, a single PCRF can be present in a Home Public Land Mobile Network (HPLMN) associated with an Internet Protocol Connectivity Access Network (IP-CAN) session of the UE. In at least one embodiment, in a local traffic outbreak roaming scenario, two PCRFs can be present, associated with an IP-CAN session of the UE: a Home PCRF (H-PCRF) within an HPLMN and a Visited PCRF (V-PCRF) within a Visited Public Land Mobile Network (VPLMN). In at least one embodiment, PCRF 2036 can be communicatively coupled with application server 2040 via P-GW 2034.In at least one embodiment, Application Server 2040 can signal PCRF 2036 to indicate a new service flow and select an appropriate Quality of Service (QoS) and charge parameter. In at least one embodiment, PCRF 2036 can provide this rule to a Policy and Charge Enforcement Function (PCEF) (not shown) with an appropriate Traffic Flow Template (TFT) and QoS Class Identifier (QCI), which initiates QoS and charge as specified by Application Server 2040.

[0204] In at least one embodiment, various components of System 2000 can be built on a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0205] Fig. Figure 21 illustrates an architecture of a System 2100 of a network according to at least one embodiment. In at least one embodiment, it is shown that System 2100 includes a UE 2102, a 5G access node or RAN node (shown as (R)A node 2108), a user-level function (shown as UPF 2104), a data network (DN 2106), which may be, for example, operator services, internet access, or third-party services, and a 5G core network (5GC) (shown as CN 2110).

[0206] In at least one embodiment, CN 2110 includes an authentication server function (AUSF 2114); a core access and mobility management function (AMF 2112); a session management function (SMF 2118); a network exposure function (NEF 2116); a policy control function (PCF 2122); a network function (NF) repository function (NRF 2120); a unified data management function (UDM 2124); and an application function (AF 2126). In at least one embodiment, CN 2110 may also include other elements not shown, such as a structured data storage network function (SDSF), an unstructured data storage network function (UDSF), and variations thereof.

[0207] In at least one embodiment, UPF 2104 can act as an anchor point for intra-RAT and inter-RAT mobility, an external PDU session federation point to DN 2106, and a branching point to support multi-homed PDU sessions. In at least one embodiment, UPF 2104 can also perform packet routing and forwarding, packet inspection, enforcement of user-level policy rules, and lawful packet interception (UP collection). Traffic usage reporting, performing user-level QoS handling (e.g., packet filtering, gating, UL / DL rate enforcement), performing uplink traffic verification (e.g., SDF-to-QoS flow mapping), transport-level packet tagging in uplink and downlink, and downlink packet buffering and downlink data notification triggering.In at least one embodiment, UPF 2104 can include an uplink classifier to assist in routing traffic flows to a data network. In at least one embodiment, DN 2106 can represent various network operator services, internet access, or third-party services.

[0208] In at least one embodiment, AUSF 2114 can store data for authenticating UE 2102 and handle authentication-related functionality. In at least one embodiment, AUSF 2114 can facilitate a common authentication framework for various access types.

[0209] In at least one embodiment, AMF 2112 can be responsible for registration management (e.g., for registering UE 2102, etc.), connection management, reachability management, mobility management, and lawful interception of AMF-related events, as well as access authentication and authorization. In at least one embodiment, AMF 2112 can provide transport for SM messages to SMF 2118 and act as a transparent proxy for routing SM messages. In at least one embodiment, AMF 2112 can also provide transport for short message (SMS) messages between UE 2102 and an SMS Function (SMSF) (not shown). Fig. 21) provide. In at least one embodiment, AMF 2112 can act as a security anchor function (SEA) that may involve interaction with AUSF 2114 and UE 2102 and receiving an intermediate key established as a result of the authentication process of UE 2102. In at least one embodiment where USIM-based authentication is used, AMF 2112 can retrieve security material from AUSF 2114. In at least one embodiment, AMF 2112 can also include a security context management (SCM) function that receives a key from SEA, which it uses to derive access network-specific keys. In at least one embodiment, AMF 2112 can further be a RAN-CP interface termination point (N2 reference point), a NAS (NI) signaling termination point, and perform NAS encryption and integrity protection.

[0210] In at least one embodiment, the AMF 2112 can also support NAS signaling with a UE 2102 via an N3 Interworking Function (IWF) interface. In at least one embodiment, N3IWF can be used to provide access to untrusted entities. In at least one embodiment, N3IWF can be a termination point for N2 and N3 interfaces for both the control plane and user plane, and can thus handle N2 signaling from SMF and AMF for PDU sessions and QoS, encapsulate / decapsulate packets for IPSec and N3 tunneling, mark N3 user-plane packets in uplinking, and enforce QoS according to N3 packet marking, taking into account QoS requirements associated with such marking received via N2.In at least one embodiment, N3IWF can also relay uplink and downlink control plane NAS(NI) signaling between UE 2102 and AMF 2112, and relay uplink and downlink user plane packets between UE 2102 and UPF 2104. In at least one embodiment, N3IWF also provides mechanisms for establishing an IPsec tunnel with UE 2102.

[0211] In at least one embodiment, SMF 2118 can be responsible for session management (e.g., session setup, modification, and release, including tunnel maintenance between UPF and AN nodes); UE IP address assignment and management (including optional authorization); UP function selection and control; configuring traffic routing at UPF to direct traffic to the appropriate destination; terminating interfaces to policy control functions; controlling the policy enforcement and QoS portion; lawful interception (for SM events and interface to the LI system); terminating SM portions of NAS messages; downward link data notification; initiating AN-specific SM information sent via AMF over N2 to AN; and determining the SSC mode of a session.In at least one embodiment, SMF 2118 can include the following roaming functionality: handling local enforcement for applying QoS-SLAB (VPLMN); charging data collection and charging interface (VPLMN); legitimate interception (in VPLMN for SM events and interface to LI system); support for interaction with external DN for transporting signaling for PDU session authorization / authentication by external DN.

[0212] In at least one embodiment, NEF 2116 can provide means for the secure exposure of services and capabilities provided by 3GPP network functions to third parties, internal exposure / re-exposure, application functions (e.g., AF 2126), edge computing or fog computing systems, etc. In at least one embodiment, NEF 2116 can authenticate, authorize, and / or throttle AFs. In at least one embodiment, NEF 2116 can also translate information exchanged with AF 2126 and information exchanged with internal network functions. In at least one embodiment, NEF 2116 can translate between an AF service identifier and internal 5GC information. In at least one embodiment, NEF 2116 can also receive information from other network functions (NFs) based on exposed capabilities of other network functions.In at least one embodiment, this information can be stored in NEF 2116 as structured data or in a data storage device NF using a standardized interface. In at least one embodiment, the stored information can then be re-exposed by NEF 2116 to other NFs and AFs and / or used for other purposes such as analytics.

[0213] In at least one embodiment, NRF 2120 can support service discovery functions, receive NF discovery requests from NF instances, and provide information from discovered NF instances to NF instances. In at least one embodiment, NRF 2120 also maintains information about available NF instances and their supported services.

[0214] In at least one embodiment, PCF 2122 can provide policy rules for control plane function(s) to enforce them and can also support a unified policy framework to govern network behavior. In at least one embodiment, PCF 2122 can also implement a front end (FE) to access subscription information relevant for policy decisions in a UDR of UDM 2124.

[0215] In at least one embodiment, UDM 2124 can handle subscription-related information to support network entity handling of communication sessions and can store subscription data from UE 2102. In at least one embodiment, UDM 2124 can comprise two parts: an application front end (FE) and a user data repository (UDR). In at least one embodiment, UDM can include a UDM FE responsible for credential processing, location management, subscription management, and so on. In at least one embodiment, multiple different front ends can serve the same user in different transactions. In at least one embodiment, the UDM FE accesses subscription information stored in a UDR and performs authentication credential processing; user identification handling; access authorization; registration / mobility management; and subscription management.In at least one embodiment, UDR can interact with PCF 2122. In at least one embodiment, UDM 2124 can also support SMS management, with an SMS FE implementing similar application logic as previously discussed.

[0216] In at least one embodiment, AF 2126 can provide application influence on traffic routing, access to a network capability exposure (NCE), and interact with a policy framework for policy control. In at least one embodiment, NCE can be a mechanism that allows a 5GC and AF 2126 to share information about NEF 2116, which can be used for edge computing implementations. In at least one embodiment, network operators and third-party services can be hosted near UE 2102 access point of connection to achieve efficient service delivery through reduced end-to-end latency and load on a transport network. In at least one embodiment, for edge computing implementations, 5GC can select a UPF 2104 near UE 2102 and perform traffic routing from UPF 2104 to DN 2106 via the N6 interface.In at least one embodiment, this can be based on UE subscription data, UE position, and information provided by AF 2126. In at least one embodiment, AF 2126 can influence UPF (re)selection and traffic routing. In at least one embodiment, based on operator deployment, when AF 2126 is considered a trusted entity, a network operator can allow AF 2126 to interact directly with relevant NFs.

[0217] In at least one embodiment, CN 2110 can include an SMSF that can be responsible for SMS subscription checks and verification, and can forward SM messages to / from UE 2102 to / from other entities, such as an SMS-GMSC / IWMSC / SMS router. In at least one embodiment, SMS can also interact with AMF 2112 and UDM 2124 for a notification procedure indicating that UE 2102 is available for SMS transmission (e.g., setting a UE unreachable flag and notifying UDM 2124 when UE 2102 is available for SMS).

[0218] In at least one embodiment, System 2100 can include the following service-based interfaces: Namf: service-based interface exhibited by AMF; Nsmf: service-based interface exhibited by SMF; Nnef: service-based interface exhibited by NEF; Npcf: service-based interface exhibited by PCF; Nudm: service-based interface exhibited by UDM; Naf: service-based interface exhibited by AF; Nnrf: service-based interface exhibited by NRF; and Nausf: service-based interface exhibited by AUSF.

[0219] In at least one embodiment, System 2100 may include the following reference points: N1: reference point between UE and AMF; N2: reference point between (R)AN and AMF; N3: reference point between (R)AN and UPF; N4: reference point between SMF and UPF; and N6: reference point between UPF and a data network. In at least one embodiment, there may be many more reference points and / or service-based interfaces between an NF service and NFs; however, these interfaces and reference points have been omitted for clarity. In at least one embodiment, an NS reference point may be between a PCF and AF; an N7 reference point may be between PCF and SMF; an N11 reference point may be between AMF and SMF; and so on. In at least one embodiment, CN 2110 may include an Nx interface, which is an inter-CN interface between MME and AMF 2112 to enable interworking between CN 2110 and CN 7216.

[0220] In at least one embodiment, System 2100 can include multiple RAN nodes (such as (R)AN nodes 2108), wherein an Xn interface is defined between two or more (R)AN nodes 2108 (e.g., gNBs) connecting to 5GC 410, between an (R)AN node 2108 (e.g., gNB) connecting to CN 2110 and an eNB (e.g., a macro-RAN node), and / or between two eNBs connecting to CN 2110.

[0221] In at least one embodiment, the Xn interface can include an Xn user-level (Xn-U) interface and an Xn control-level (Xn-C) interface. In at least one embodiment, Xn-U can provide non-guaranteed delivery of user-level PDUs and support / provide data forwarding and flow control functionality. In at least one embodiment, Xn-C can provide management and fault handling functionality, functionality for managing an Xn-C interface, and mobility support for UE 2102 in a connected mode (e.g., CM-CONNECTED), including functionality for managing UE mobility for connected mode between one or more (R)AN nodes 2108.In at least one embodiment, mobility support can include context transfer from an old (source) serving(R)AN node 2108 to a new (destination) serving (R)AN node 2108; and control of user-level tunnels between old (source) serving (R)AN node 2108 and new (destination) serving (R)AN node 2108.

[0222] In at least one embodiment, an Xn-U protocol stack can include a transport network layer built on top of an Internet Protocol (IP) transport layer and a GTP-U layer on top of a UDP and / or IP layer(s) to carry user-level PDUs. In at least one embodiment, an Xn-C protocol stack can include an application-layer signaling protocol (referred to as the Xn application protocol (Xn-AP)) and a transport network layer built on top of an SCTP layer. In at least one embodiment, the SCTP layer can reside on top of an IP layer. In at least one embodiment, the SCTP layer provides guaranteed delivery of application-layer messages. In at least one embodiment, point-to-point transmission is used in a transport IP layer to deliver signaling PDUs.In at least one embodiment, the Xn-U protocol stack and / or an Xn-C protocol stack can be the same as or similar to a user-level and / or control-level protocol stack or protocol stacks shown and described herein.

[0223] In at least one embodiment, various components of System 2100 can be powered by a data center that can be digitally simulated by one or more neural networks, as described in at least one embodiment in Fig. 1-5.

[0224] Fig. Figure 22 illustrates a control plane protocol stack according to at least one embodiment. In at least one embodiment, a control plane 2200 is shown as a communications protocol stack between UE 1502 (or alternatively UE 1504), RAN 1516, and MME(s) 1528.

[0225] In at least one embodiment, PHY layer 2202 can transmit or receive information used by MAC layer 2204 over one or more air interfaces. In at least one embodiment, PHY layer 2202 can further perform link matching or adaptive modulation and coding (AMC), power control, cell search (e.g., for initial synchronization and handover purposes), and other measurements used by higher layers, such as an RRC layer 2210. In at least one embodiment, PHY layer 2202 can further perform fault detection on transport channels, forward error correction (FEC) encoding / decoding of transport channels, modulation / demodulation of physical channels, interleaving, rate matching, mapping to physical channels, and multiple-input multiple-output (MIMO) antenna processing.

[0226] In at least one embodiment, MAC layer 2204 can perform mapping between logical channels and transport channels, multiplex MAC service data units (SDUs) from one or more logical channels onto transport blocks (TBs) to be delivered to PHY via transport channels, demultiplex MAC-SDUs onto one or more logical channels of transport blocks (TBs) to be delivered by PHY via transport channels, multiplex MAC-SDUs onto TBs, scheduling information reporting, error correction through hybrid automatic retry request (HARD), and logical channel prioritization.

[0227] In at least one embodiment, RLC layer 2206 can operate in a variety of modes, including: Transparent Mode (TM), Unacknowledged Mode (UM), and Acknowledged Mode (AM). In at least one embodiment, RLC layer 2206 can perform transmission of upper-layer protocol data units (PDUs), error correction by automatic retry request (ARQ) for AM data transmissions, and concatenation, segmentation, and recomposition of RLC SDUs for UM and AM data transmissions. In at least one embodiment, RLC layer 2206 can also perform resegmentation of RLC data PDUs for AM data transmissions, reorder RLC data PDUs for UM and AM data transmissions, detect duplicate data for UM and AM data transmissions, discard RLC SDUs for UM and AM data transmissions, detect protocol errors for AM data transmissions, and perform RLC re-establishment.

[0228] In at least one embodiment, PDCP layer 2208 can perform header compression and decompression of IP data, maintain PDCP sequence numbers (SNs), perform in-sequence delivery of upper-layer PDUs on new lower-layer setups, perform duplicates of lower-layer SDUs on new lower-layer setups for radio support mapped to RLC AM, encrypt and decrypt control-plane data, perform integrity protection and integrity verification of control-plane data, control timer-based data discarding, and perform security operations (e.g., encryption, decryption, integrity protection, integrity verification, etc.).

[0229] In at least one embodiment, key services and functions of an RRC layer 2210 may include broadcasting system information (e.g., contained in Master Information Blocks (MIBs) or System Information Blocks (SIBs) with respect to a Non-Access Stratum (NAS)), broadcasting system information with respect to an Access Stratum (AS), paging, establishing, maintaining, and releasing an RRC link between a UE and E-UTRAN (e.g., RRC link paging, RRC link establishment, RRC link modification, and RRC link release), establishing, configuring, maintaining, and releasing point-to-point radio supports, security functions including key management, Inter-Radio Access Technology (RAT) mobility, and meter configuration for UE meter reporting.In at least one embodiment, the MIBs and SIBs can have one or more information elements (IEs), each of which can have individual data fields or data structures.

[0230] In at least one embodiment, UE 1502 and RAN 1516 can use a Uu interface (e.g., an LTE Uu interface) to exchange control plane data via a protocol stack comprising PHY layer 2202, MAC layer 2204, RLC layer 2206, PDCP layer 2208, and RRC layer 2210.

[0231] In at least one embodiment, non-access stratum (NAS) protocols (NAS protocols 2212) form a highest stratum of a control plane between UE 1502 and MME(s) 1528. In at least one embodiment, NAS protocols 2212 support mobility of UE 1502 and session management procedures to establish and maintain IP connectivity between UE 1502 and P-GW 1534.

[0232] In at least one embodiment, the Si application protocol (S1-AP) layer (Si-AP layer 2222) can support Si interface functions and include elementary procedures (EPs). In at least one embodiment, an EP is an interaction unit between RAN 1516 and CN 1538. In at least one embodiment, S1-AP layer services can have two groups: UE-associated services and non-UE-associated services. In at least one embodiment, these applications perform functions including, but not limited to: E-UTRAN radio access support (E-RAB) management, UE capability indication, mobility, NAS signaling transport, RAN information management (RIM), and configuration transmission.

[0233] In at least one embodiment, the Stream Control Transmission Protocol (SCTP) layer (alternatively referred to as a Stream Control Transmission Protocol / Internet Protocol (SCTP / IP) layer) (SCTP layer 2220) can ensure reliable delivery of signaling messages between RAN 1516 and MME(s) 1528, partly based on an IP protocol supported by an IP layer 2218. In at least one embodiment, L2 layer 2216 and an L1 layer 2214 can refer to communication links (e.g., wired or wireless) used by a RAN node and MME to exchange information.

[0234] In at least one embodiment, RAN 1516 and MME(s) 1528 can use an S1-MME interface to exchange control plane data via a protocol stack comprising an L1 layer 2214, L2 layer 2216, IP layer 2218, SCTP layer 2220 and Si-AP layer 2222.

[0235] In at least one embodiment, control level 2200 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0236] Fig. Figure 23 illustrates a user-level protocol stack according to at least one embodiment. In at least one embodiment, a user level 2300 is shown as a communication protocol stack between a UE 1502, RAN 1516, S-GW 1530, and P-GW 1534. In at least one embodiment, user level 2300 can use the same protocol layers as control level 1700. For example, in at least one embodiment, UE 1502 and RAN 1516 can use a Uu interface (e.g., an LTE Uu interface) to exchange user-level data over a protocol stack comprising PHY layer 1702, MAC layer 1704, RLC layer 1706, and PDCP layer 1708.

[0237] In at least one embodiment, the General Packet Radio Service (GPRS) Tunneling Protocol can be used for a User-Level Layer (GTP-U) (GTP-U layer 2304) to carry user data within a GPRS core network and between a radio access network and a core network. In at least one embodiment, the transported user data can be, for example, packets in any of the IPv4, IPv6, or PPP formats. In at least one embodiment, the UDP and IP Security Layer (UDP / IP layer 2302) can provide checksums for data integrity, port numbers for addressing various functions at a source and destination, and encryption and authentication on selected data flows.In at least one embodiment, RAN 1516 and S-GW 1530 can use an S1-U interface to exchange user-level data over a protocol stack comprising L1 layer 1714, L2 layer 1716, UDP / IP layer 2302, and GTP-U layer 2304. In at least one embodiment, S-GW 1530 and P-GW 1534 can use an S5 / S8a interface to exchange user-level data over a protocol stack comprising L1 layer 1714, L2 layer 1716, UDP / IP layer 2302, and GTP-U layer 2304. In at least one embodiment, as above with respect to... Fig. 17 discussed, NAS protocols support UE 1502 mobility and session management procedures to establish and maintain IP connectivity between UE 1502 and P-GW 1534.

[0238] In at least one embodiment, user level 2300 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0239] Fig. Figure 24 illustrates components 2400 of a core network according to at least one embodiment. In at least one embodiment, components of CN 1538 can be implemented in a physical node or separate physical nodes, including components for reading and executing instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium). In at least one embodiment, network function virtualization (NFV) is used to virtualize any or all of the network node functions described above via executable instructions stored in one or more computer-readable storage media (described in more detail below). In at least one embodiment, a logical instantiation of CN 1538 can be designated as a network slice 2402 (e.g., it is shown that network slice 2402 includes HSS 1532, MME(s) 1528, and S-GW 1530).In at least one embodiment, a logical instantiation of a portion of CN 1538 can be referred to as network sub-slice 2404 (e.g., it is shown that network sub-slice 2404 includes P-GW 1534 and PCRF 1536).

[0240] In at least one embodiment, NFV architectures and infrastructures can be used to virtualize one or more network functions, alternatively performed by proprietary hardware, onto physical resources comprising a combination of industry-standard server hardware, storage hardware, or switches. In at least one embodiment, NFV systems can be used to execute virtual or reconfigurable implementations of one or more EPC components / functions.

[0241] In at least one embodiment, various components 2400 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0242] Fig. Figure 25 is a block diagram illustrating components according to at least one embodiment of a System 2500 to support network function virtualization (NFV). In at least one embodiment, System 2500 is illustrated to include a virtualized infrastructure manager (shown as VIM 2502), a network function virtualization infrastructure (shown as NFVI 2504), a VNF manager (shown as VNFM 2506), virtualized network functions (shown as VNF 2508), an element manager (shown as EM 2510), an NFV orchestrator (shown as NFVO 2512), and a network manager (shown as NM 2514).

[0243] In at least one embodiment, VIM 2502 manages resources of NFVI 2504. In at least one embodiment, NFVI 2504 may include physical or virtual resources and applications (including hypervisors) used to run System 2500. In at least one embodiment, VIM 2502 may manage a virtual resource lifecycle with NFVI 2504 (e.g., creating, maintaining, and terminating virtual machines (VMs) associated with one or more physical resources), track VM instances, monitor the performance, failure, and security of VM instances and associated physical resources, and expose VM instances and associated physical resources to other management systems.

[0244] In at least one embodiment, VNFM 2506 can manage VNF 2508. In at least one embodiment, VNF 2508 can be used to execute EPC components / functions. In at least one embodiment, VNFM 2506 can manage a lifecycle of VNF 2508 and track the performance, faults, and safety of virtual aspects of VNF 2508. In at least one embodiment, EM 2510 can track the performance, faults, and safety of functional aspects of VNF 2508. In at least one embodiment, tracking data from VNFM 2506 and EM 2510 can include, for example, performance measurement (PM) data used by VIM 2502 or NFVI 2504. In at least one embodiment, both VNFM 2506 and EM 2510 can scale a set of VNFs up / down from System 2500.

[0245] In at least one embodiment, NFVO 2512 can coordinate, authorize, release, and engage resources from NFVI 2504 to provide a requested service (e.g., to execute an EPC function, component, or slice). In at least one embodiment, NM 2514 can provide a package of end-user functions responsible for managing a network, which may include network elements with VNFs, non-virtualized network functions, or both (management of VNFs may be handled via EM 2510).

[0246] In at least one embodiment, various components of System 2500 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described. Computer-based systems

[0247] The following figures represent, without limitation, exemplary computer-based systems that can be used to implement at least one embodiment.

[0248] Fig. Figure 26 illustrates a processing system 2600 according to at least one embodiment. In at least one embodiment, the processing system 2600 includes one or more processors 2602 and one or more graphics processors 2608 and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 2602 or processor cores 2607. In at least one embodiment, the processing system 2600 is a processing platform incorporated within an integrated system-on-a-chip (“SoC”) circuit for use in mobile, handheld, or embedded devices.

[0249] In at least one embodiment, the processing system 2600 may include or be incorporated within a server-based gaming platform, a gaming console, a media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, the processing system 2600 is a mobile phone, smartphone, tablet computer, or mobile internet device. In at least one embodiment, the processing system 2600 may also include, couple with, or be integrated within a portable device such as a wearable smartwatch, smart eyewear, augmented reality, or virtual reality device. In at least one embodiment, the processing system 2600 is a television or set-top box device with one or more processors 2602 and a graphical interface generated by one or more graphics processors 2608.

[0250] In at least one embodiment, one or more processors 2602 each include one or more processor cores 2607 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2607 is configured to process a specific instruction set 2609. In at least one embodiment, instruction set 2609 can facilitate Complex Instruction Set Computing (“CISC”), Reduced Instruction Set Computing (“RISC”), or Computing via a Very Long Instruction Word (“VLIW”). In at least one embodiment, processor cores 2607 can each process a different instruction set 2609, which may include instructions to facilitate the emulation of other instruction sets. In at least one embodiment, processor core 2607 can also include other processing devices, such as a digital signal processor (“DSP”).

[0251] In at least one embodiment, processor 2602 includes cache memory (“cache”) 2604. In at least one embodiment, processor 2602 can include a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared by different components of processor 2602. In at least one embodiment, processor 2602 also uses an external cache (e.g., a level 3 cache (“L3”) or a last-level cache (“LLC”)) (not shown) that can be shared by processor cores 2607 using known cache coherence techniques. In at least one embodiment, register file 2606 is additionally included in processor 2602, which can contain different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and an instruction pointer register).In at least one embodiment, register file 2606 may contain general-purpose registers or other registers.

[0252] In at least one embodiment, a processor 2602 or multiple processors 2602 are coupled to an interface bus 2610 or multiple interface buses 2610 to transmit communication signals, such as address, data, or control signals, between the processor 2602 and other components in the processing system 2600. In at least one embodiment, the interface bus 2610 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 2610 is not limited to a DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express (PCIe)), memory buses, or other types of interface buses. In at least one embodiment, the processor(s) 2602 include an integrated memory controller 2616 and a platform controller hub 2630.In at least one embodiment, the storage controller 2616 facilitates communication between a storage device and other components of the processing system 2600, while the platform controller hub (“PCH”) 2630 provides connections to input / output (“I / O”) devices via a local I / O bus.

[0253] In at least one embodiment, the storage device 2620 can be a dynamic random-access memory (“DRAM”) device, a static random-access memory (“SRAM”) device, a flash memory device, a phase-change memory device, or another storage device with suitable performance to serve as processor memory. In at least one embodiment, the storage device 2620 can operate as system memory for the processing system 2600 to store data 2622 and instructions 2621 for use when one or more processors 2602 execute an application or process. In at least one embodiment, the storage controller 2616 also couples with an optional external graphics processor 2612, which can communicate with one or more graphics processors 2608 in processors 2602 to perform graphics and media operations. In at least one embodiment, a display device 2611 can be connected to processor(s) 2602.In at least one embodiment, display device 2611 can include one or more internal display devices, such as in a mobile electronic device or a laptop, or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, display device 2611 can include a head-mounted display (“HMD”), such as a stereoscopic display device for use in virtual reality (“VR”) or augmented reality (“AR”) applications.

[0254] In at least one embodiment, the platform controller hub 2630 enables peripheral devices to connect to the storage device 2620 and the processor 2602 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2646, a network controller 2634, a firmware interface 2628, a wireless transceiver 2626, touch sensors 2625, and a data storage device 2624 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2624 can connect via a storage interface (e.g., SATA) or via a peripheral bus, such as PCI or PCIe. In at least one embodiment, the touch sensors 2625 can include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, wireless transceiver 2626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (“LTE”) transceiver. In at least one embodiment, firmware interface 2628 enables communication with system firmware and can, for example, be a unified extensible firmware interface (“UEFI”). In at least one embodiment, network controller 2634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled with interface bus 2610. In at least one embodiment, audio controller 2646 is a multi-channel, high-resolution audio controller. In at least one embodiment, processing system 2600 includes an optional legacy I / O controller 2640 for coupling legacy (e.g.,Personal System 2 (“PS / 2”) devices with processing system 2600. In at least one embodiment, platform controller hub 2630 can also be connected to one or more Universal Serial Bus (“USB”) controllers 2642 that connect input devices, such as keyboard and mouse 2643 combinations, a camera 2644 or other USB input devices.

[0255] In at least one embodiment, an instance of memory controller 2616 and platform controller hub 2630 can be integrated into a discrete external graphics processor, such as the external graphics processor 2612. In at least one embodiment, the platform controller hub 2630 and / or memory controller 2616 can be external to one or more processors 2602. For example, in at least one embodiment, processing system 2600 can include an external memory controller 2616 and platform controller hub 2630, which can be configured as a memory controller hub and peripheral controller hub within a system chipset that communicates with processor(s) 2602.

[0256] In at least one embodiment, various components of processing system 2600 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0257] Fig. Figure 27 illustrates a computer system 2700 according to at least one embodiment. In at least one embodiment, computer system 2700 can be a system with interconnected devices and components, a system-on-a-chip (SOC), or a combination thereof. In at least one embodiment, computer system 2700 comprises a processor 2702, which may include execution units for executing an instruction. In at least one embodiment, computer system 2700 can, without limitation, include a component, such as a processor 2702, to employ execution units, including logic, for performing algorithms for processing data.In at least one embodiment, Computer System 2700 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) may also be used. In at least one embodiment, Computer System 2700 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0258] In at least one embodiment, Computer System 2700 can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (DSP), a system-on-a-chip (SoC), network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system capable of executing one or more instructions.

[0259] In at least one embodiment, Computer System 2700 may, without limitation, include a Processor 2702, which may, without limitation, include one or more Execution Units 2708 that may be configured to execute a Compute Unified Device Architecture (“CUDA”) (CUDA® is developed by NVIDIA Corporation of Santa Clara, CA) program. In at least one embodiment, a CUDA program is at least a portion of a software application written in a CUDA programming language. In at least one embodiment, Computer System 2700 is a single-processor desktop or server system. In at least one embodiment, Computer System 2700 may be a multiprocessor system.In at least one embodiment, processor 2702 can, without limitation, include a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other processing device, such as a digital signal processor. In at least one embodiment, processor 2702 can be coupled to a processor bus 2710, which can transmit data signals between processor 2702 and other components in the computer system 2700.

[0260] In at least one embodiment, processor 2702 can include, without restriction, an internal level 1 cache memory (“L1”) 2704. In at least one embodiment, processor 2702 can include a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory can reside outside of processor 2702. In at least one embodiment, processor 2702 can also include a combination of both internal and external caches. In at least one embodiment, a register file 2706 can store different types of data in different registers, including, but not limited to, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0261] In at least one embodiment, execution unit 2708, containing, among other things, logic for performing integer and floating-point operations, also resides in processor 2702. Processor 2702 may also include a microcode ("ucode") read-only memory ("ROM") that stores microcode for certain macro instructions. In at least one embodiment, execution unit 2708 may include logic for handling a packed instruction set 2709. In at least one embodiment, by including a packed instruction set 2709 in an instruction set of a general-purpose processor 2702, together with associated circuitry for executing instructions, operations used by many multimedia applications can be performed using packed data in a general-purpose processor 2702.In at least one embodiment, many multimedia applications can be accelerated and run more efficiently by using the full width of a processor's data bus to perform operations on packed data, which can eliminate the need to transfer smaller data units across a processor's data bus to perform one or more operations per data element.

[0262] In at least one embodiment, the execution unit 2708 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 2700 can include a memory 2720 without restriction. In at least one embodiment, the memory 2720 can be implemented as a DRAM device, an SRAM device, a flash memory device, or another type of memory device. The memory 2720 can store instructions 2719 and / or data 2721, represented by data signals that can be executed by the processor 2702.

[0263] In at least one embodiment, a system logic chip can be coupled with a processor bus 2710 and memory 2720. In at least one embodiment, a system logic chip can, without restriction, include a memory controller hub (“MCH”) 2716, and processor 2702 can communicate with MCH 2716 via processor bus 2710. In at least one embodiment, MCH 2716 can provide a high-bandwidth memory path 2718 to memory 2720 for instruction and data storage and for storing graphics instructions, data, and textures. In at least one embodiment, MCH 2716 can route data signals between processor 2702, memory 2720, and other components in the computer system 2700 and bridge data signals between processor bus 2710, memory 2720, and a system I / O 2722. In at least one embodiment, the system logic chip can provide a graphics port for coupling with a graphics controller.In at least one embodiment, MCH 2716 can be coupled to memory 2720 via high-bandwidth memory path 2718, and graphics / video card 2712 can be coupled to MCH 2716 via an Accelerated Graphics Port (“AGP”) assembly 2714.

[0264] In at least one embodiment, computer system 2700 can use system I / O 2722, which is a proprietary hub interface bus, to couple MCH 2716 with I / O controller (“ICH”) 2730. In at least one embodiment, ICH 2730 can provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus can, without limitation, include a high-speed I / O bus for connecting peripheral devices to memory 2720, a chipset, and processor 2702. Examples may include, but are not limited to, an audio controller 2729, a firmware hub (“flash BIOS”) 2728, a wireless transceiver 2726, a data storage device 2724, a legacy I / O controller 2723 which includes a user input interface 2725 and a keyboard interface, a serial expansion port 2727, such as a USB, and a network controller 2734.Data storage device 2724 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0265] In at least one embodiment, illustrated Fig. 27 a system comprising interconnected hardware devices or “chips”. In at least one embodiment, Fig. 27 illustrates an exemplary SoC. In at least one embodiment, in Fig. Figure 27 illustrates devices that can be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more System 2700 components are interconnected using Compute Express Link (“CXL”) interconnects.

[0266] In at least one embodiment, computer system 2700 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0267] Fig. Figure 28 illustrates a System 2800 according to at least one embodiment. In at least one embodiment, System 2800 is an electronic device that requires a Processor 2810. In at least one embodiment, System 2800 can be, for example, and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device.

[0268] In at least one embodiment, System 2800 can include, without limitation, Processor 2810, which is communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, Processor 2810 is coupled using a bus or interface, such as an I 2 C-Bus, a system management bus (“SMBus”), a low-pin-count bus (“LPC”), a serial peripheral interface (“SPI”), a high-resolution audio bus (“HDA”), an advanced serial technology interface bus (“SATA”), a USB bus (versions 1, 2, 3), or a universal asynchronous receiver / transmitter bus (“UART”). In at least one embodiment, illustrated Fig. 28 a system comprising interconnected hardware devices or “chips”. In at least one embodiment, Fig. 28 illustrates an exemplary SoC. In at least one embodiment, in Fig. 28 illustrated devices are interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 28 interconnected using CXL assemblies.

[0269] In at least one embodiment, Fig. 28 include a display 2824, a touchscreen 2825, a touchpad 2830, a near field communication (“NFC”) 2845, a sensor hub 2840, thermal sensors 2839, an Express chipset (“EC”) 2835, a trusted platform module (“TPM”) 2838, BIOS / firmware / flash memory (“BIOS, FW Flash”) 2822, a DSP 2860, a solid state disk (“SSD”) or a hard disk drive (“HDD”) 2820, a wireless local area network (“WLAN”) 2850, a Bluetooth unit 2852, a wireless wide area network (“WWAN”) 2856, a global positioning system (“GPS”) 2855, a camera (“USB 3.0 camera”) 2854, such as a USB 3.0 camera, or a low power double data rate The LPDDR3 memory unit (2815), for example, is implemented in the LPDDR3 standard. These components can each be implemented in any suitable way.

[0270] In at least one embodiment, other components can be communicatively coupled to processor 2810 via the components discussed above. In at least one embodiment, an accelerometer 2841, an ambient light sensor (“ALS”) 2842, a compass 2843, and a gyroscope 2844 can be communicatively coupled to sensor hub 2840. In at least one embodiment, a thermal sensor 2839, a fan 2837, a keyboard 2836, and a touchpad 2830 can be communicatively coupled to EC 2835. In at least one embodiment, a loudspeaker 2863, headphones 2864, and a microphone (“mic”) 2865 can be communicatively coupled to an audio unit (“audio codec and class d amplifier”) 2862, which in turn can be communicatively coupled to DSP 2860. In at least one embodiment, audio unit 2862 can, for example and without limitation, include an audio encoder / decoder (“codec”) and a class D amplifier.In at least one embodiment, a SIM card (“SIM”) 2857 can be communicatively coupled with a WWAN unit 2856. In at least one embodiment, components such as a WLAN unit 2850 and a Bluetooth unit 2852, as well as a WWAN unit 2856, can be implemented in a Next Generation Form Factor (“NGFF”).

[0271] In at least one embodiment, computer system 2800 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0272] Fig. Figure 29 illustrates an exemplary integrated circuit 2900 according to at least one embodiment. In at least one embodiment, the exemplary integrated circuit 2900 is a SoC that can be fabricated using one or more IP cores. In at least one embodiment, the integrated circuit 2900 includes an application processor 2905 or more application processors 2905 (e.g., CPUs), at least one graphics processor 2910, and may additionally include an image processor 2915 and / or a video processor 2920, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 2900 includes peripheral or bus logic, including a USB controller 2925, a UART controller 2930, an SPI / SDIO controller 2935, and an I 2 S / I 2The integrated circuit 2900 includes a C-controller 2940. In at least one embodiment, the integrated circuit 2900 can include a display device 2945 coupled to one or more High-Definition Multimedia Interface (“HDMI”) controllers 2950 and Mobile Industry Processor Interface (“MIPI”) display interfaces 2955. In at least one embodiment, memory can be provided by a flash memory subsystem 2960, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface can be provided via a memory controller 2965 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 2970.

[0273] In at least one embodiment, the integrated circuit 2900 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0274] Fig. Figure 30 illustrates a computing system 3000 according to at least one embodiment. In at least one embodiment, the computing system 3000 includes a processing subsystem 3001 with one or more processors 3002 and a system memory 3004, which communicates via a compound path that may include a memory hub 3005. In at least one embodiment, the memory hub 3005 may be a separate component within a chipset component or may be integrated within one or more processors 3002. In at least one embodiment, the memory hub 3005 is coupled to an I / O subsystem 3011 via a communication link 3006. In at least one embodiment, the I / O subsystem 3011 includes an I / O hub 3007, which may enable the computing system 3000 to receive inputs from one or more input devices 3008.In at least one embodiment, the I / O hub 3007 enables a display controller, which may be included in one or more processors 3002, to provide outputs to one or more display devices 3010A. In at least one embodiment, one or more display devices 3010A coupled to the I / O hub 3007 may include a local, internal, or embedded display device.

[0275] In at least one embodiment, processing subsystem 3001 includes one or more parallel processors 3012 coupled to memory hub 3005 via a bus or other communication interface 3013. In at least one embodiment, the communication interface 3013 can be one of any number of standards-based communication interface technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 3012 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a multi-integrated core processor.In at least one embodiment, a parallel processor 3012 or several parallel processors 3012 form a graphics processing subsystem that can output pixels to a display device 3010A or several display devices 3010A coupled via an I / O hub 3007. In at least one embodiment, a parallel processor 3012 or several parallel processors 3012 can also include a display controller and a display interface (not shown) to enable a direct connection to a display device 3010B or several display devices 3010B.

[0276] In at least one embodiment, a system storage unit 3014 can be connected to an I / O hub 3007 to provide a storage mechanism for a computing system 3000. In at least one embodiment, an I / O switch 3016 can be used to provide an interface mechanism to enable connections between the I / O hub 3007 and other components, such as a network adapter 3018 and / or a wireless network adapter 3019, which can be integrated into a platform, and various other devices that can be added via one or more add-in devices 3020. In at least one embodiment, the network adapter 3018 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 3019 can include one or more wireless radios from a Wi-Fi, Bluetooth, NFC, or other networking device.

[0277] In at least one embodiment, the computing system 3000 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video recording devices and / or variations thereof, which may also be connected to the I / O hub 3007. In at least one embodiment, communication paths connecting various components in Fig. 30 interconnected, implemented using any suitable protocols, such as PCI-based protocols (e.g. PCIe) or other bus or point-to-point communication interfaces and / or protocol(s), such as NVLink high-speed federation or federation protocols.

[0278] In at least one embodiment, a parallel processor 3012 or multiple parallel processors 3012 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constitute a graphics processing unit (“GPU”). In at least one embodiment, a parallel processor 3012 or multiple parallel processors 3012 include circuitry optimized for general-purpose processing. In at least one embodiment, components of computer system 3000 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, a parallel processor 3012 or multiple parallel processors 3012, memory hub 3005, processor 3002 or processors 3002, and I / O hub 3007 can be integrated into an integrated SoC circuit.In at least one embodiment, components of Computer System 3000 can be integrated into a single package to form a system-in-package (“SIP”) configuration. In at least one embodiment, at least a proportion of components of Computer System 3000 can be integrated into a multi-chip module (“MCM”), which can be interconnected with other multi-chip modules to form a modular computer system. In at least one embodiment, I / O subsystem 3011 and display devices 3010B from Computer System 3000 are omitted.

[0279] In at least one embodiment, computer system 3000 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described. Processing systems

[0280] The following figures represent, without limitation, exemplary processing systems that can be used to implement at least one embodiment.

[0281] Fig. Figure 31 illustrates an accelerated processing unit (“APU”) 3100 according to at least one embodiment. In at least one embodiment, APU 3100 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, APU 3100 can be configured to execute an application program, such as a CUDA program. In at least one embodiment, APU 3100 includes, without limitation, a core complex 3110, a graphics complex 3140, Fabric 3160, I / O interfaces 3170, a memory controller 3180, a display controller 3192, and a multimedia engine 3194. In at least one embodiment, APU 3100 can include, without limitation, any number of core complexes 3110, any number of graphics complexes 3140, any number of display controllers 3192, and any number of multimedia engines 3194 in any combination.For explanatory purposes, multiple instances of similar objects are referred to herein by reference numbers that identify an object, and bracket numbers that identify an instance where required.

[0282] In at least one embodiment, core complex 3110 is a CPU, graphics complex 3140 is a GPU, and APU 3100 is a processing unit that integrates both the 3110 and 3140 on a single chip without restriction. In at least one embodiment, some tasks can be assigned to core complex 3110 and other tasks can be assigned to graphics complex 3140. In at least one embodiment, core complex 3110 is configured to execute main control software associated with APU 3100, such as an operating system. In at least one embodiment, core complex 3110 is a main processor of APU 3100 that controls and coordinates operations of other processors. In at least one embodiment, core complex 3110 issues instructions that control an operation of graphics complex 3140.In at least one embodiment, core complex 3110 can be configured to execute host executable code derived from CUDA source code, and graphics complex 3140 can be configured to execute device executable code derived from CUDA source code.

[0283] In at least one embodiment, core complex 3110 includes, without limitation, cores 3120(1)-2620(4) and an L3 cache 3130. In at least one embodiment, core complex 3110 can include, without limitation, any number of cores 3120 and any number and type of caches in any combination. In at least one embodiment, cores 3120 are configured to execute instructions of a single instruction set architecture (“ISA”). In at least one embodiment, each core 3120 is a CPU core.

[0284] In at least one embodiment, without limitation, each core 3120 includes, among other things, a retrieval / decoding unit 3122, an integer execution engine 3124, a floating-point execution engine 3126, and an L2 cache 3128. In at least one embodiment, the retrieval / decoding unit 3122 retrieves instructions, decodes such instructions, generates microoperations, and sends separate microinstructions to the integer execution engine 3124 and the floating-point execution engine 3126. In at least one embodiment, the retrieval / decoding unit 3122 can simultaneously send one microinstruction to the integer execution engine 3124 and another microinstruction to the floating-point execution engine 3126. In at least one embodiment, without limitation, the integer execution engine 3124 performs, among other things, integer and memory operations. In at least one embodiment, the floating-point engine 3126 performs, without limitation, floating-point and vector operations, among other things.In at least one embodiment, the retrieval / decoding unit 3122 sends micro-instructions to a single execution engine, which replaces both the integer execution engine 3124 and the floating-point execution engine 3126.

[0285] In at least one embodiment, each core 3120(i), where i is an integer representing a single instance of core 3120, can access L2 cache 3128(i), contained in core 3120(i). In at least one embodiment, each core 3120, contained in core complex 3110(j), where j is an integer representing a single instance of core complex 3110, is connected to other cores 3120, contained in core complex 3110(j), via L3 cache 3130(j), contained in core complex 3110(j). In at least one embodiment, cores 3120, contained in core complex 3110(j), where j is an integer representing a single instance of core complex 3110, can access all L3 caches 3130(j), contained in core complex 3110(j). In at least one embodiment, L3-Cache 3130 can contain any number of slices without restriction.

[0286] In at least one embodiment, graphics complex 3140 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 3140 is configured to execute graphics pipeline operations, such as drawing commands, pixel operations, geometric calculations, and other operations associated with rendering an image on a display. In at least one embodiment, graphics complex 3140 is configured to perform operations that are not related to graphics. In at least one embodiment, graphics complex 3140 is configured to perform both graphics-related and non-graphics-related operations.

[0287] In at least one embodiment, the graphics complex 3140 includes, without limitation, any number of processing units 3150 and an L2 cache 3142. In at least one embodiment, the processing units 3150 share the L2 cache 3142. In at least one embodiment, the L2 cache 3142 is partitioned. In at least one embodiment, the graphics complex 3140 includes, without limitation, any number of processing units 3150 and any number (including zero) and type of caches. In at least one embodiment, the graphics complex 3140 includes, without limitation, any amount of dedicated graphics hardware.

[0288] In at least one embodiment, computing units 3150 include, without limitation, any number of SIMD units 3152 and a shared memory 3154. In at least one embodiment, each SIMD unit 3152 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, computing units 3150 can execute any number of thread blocks, but each thread block is executed on a single computing unit 3150. In at least one embodiment, a thread block includes, without limitation, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 3152 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g.,16 threads), wherein each thread in a warp belongs to a single thread block and is configured to process a different data set based on a single set of instructions. In at least one embodiment, prediction can be used to disable one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a worker is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block can synchronize with each other and communicate over shared memory 3154.

[0289] In at least one embodiment, Fabric 3160 is a system assembly that facilitates data and control transmissions across the core complex 3110, graphics complex 3140, I / O interfaces 3170, memory controller 3180, display controller 3192, and multimedia engine 3194. In at least one embodiment, APU 3100 can, without limitation, include any number and type of system assembly in addition to or instead of Fabric 3160, facilitating data and control transmissions across any number and type of directly or indirectly linked components, which may be located inside or outside APU 3100. In at least one embodiment, I / O interfaces 3170 represent any number and type of I / O interfaces (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.).In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 3170. In at least one embodiment, peripheral devices coupled to I / O interfaces 3170 can include, without limitation, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, and so on.

[0290] In at least one embodiment, the AMD92 display controller displays images on one or more display devices, such as a liquid crystal display (“LCD”) device. In at least one embodiment, the multimedia engine 3194 includes, without limitation, any quantity and type of switching technology related to multimedia, such as a video decoder, a video encoder, an image signal processor, etc. In at least one embodiment, the memory controller 3180 facilitates data transfers between the APU 3100 and a unified system memory 3190. In at least one embodiment, the core complex 3110 and the graphics complex 3140 jointly utilize the unified system memory 3190.

[0291] In at least one embodiment, the APU 3100 implements a memory subsystem that includes, without restriction, any number and type of memory controllers 3180 and memory devices (e.g., shared memory 3154) that can be dedicated to a component or shared by multiple components. In at least one embodiment, the APU 3100 implements a cache subsystem that includes, without restriction, one or more cache memories (e.g., L2 caches 2728, L3 cache 3130, and L2 cache 3142) that can each be private or shared by any number of components (e.g., cores 3120, core complex 3110, SIMD units 3152, arithmetic units 3150, and graphics complex 3140).

[0292] In at least one embodiment, the APU 3100 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0293] Fig. Figure 32 illustrates a CPU 3200 according to at least one embodiment. In at least one embodiment, the CPU 3200 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, the CPU 3200 can be configured to execute an application program. In at least one embodiment, the CPU 3200 is configured to execute main control software, such as an operating system. In at least one embodiment, the CPU 3200 issues instructions that control an operation of an external GPU (not shown). In at least one embodiment, the CPU 3200 can be configured to execute host executable code derived from CUDA source code, and an external GPU can be configured to execute device executable code derived from such CUDA source code.In at least one embodiment, CPU 3200 includes without limitation any number of core complexes 3210, fabric 3260, I / O interfaces 3270 and memory controller 3280.

[0294] In at least one embodiment, core complex 3210 includes, without limitation, cores 3220(1)-2720(4) and an L3 cache 3230. In at least one embodiment, core complex 3210 can include, without limitation, any number of cores 3220 and any number and type of caches in any combination. In at least one embodiment, cores 3220 are configured to execute instructions of a single ISA. In at least one embodiment, each core 3220 is a CPU core.

[0295] In at least one embodiment, each core 3220 includes, without limitation, a retrieval / decoding unit 3222, an integer execution engine 3224, a floating-point execution engine 3226, and an L2 cache 3228. In at least one embodiment, the retrieval / decoding unit 3222 retrieves instructions, decodes such instructions, generates microoperations, and sends separate microinstructions to the integer execution engine 3224 and the floating-point execution engine 3226. In at least one embodiment, the retrieval / decoding unit 3222 can simultaneously send one microinstruction to the integer execution engine 3224 and another microinstruction to the floating-point execution engine 3226. In at least one embodiment, the integer execution engine 3224 performs, among other things, integer and memory operations. In at least one embodiment, the floating-point machine 3226 performs, among other things, floating-point and vector operations.In at least one embodiment, the retrieval / decoding unit 3222 sends micro-instructions to a single execution machine, which replaces both the integer execution machine 3224 and the floating-point execution machine 3226.

[0296] In at least one embodiment, each kernel 3220(i), where i is an integer representing a single instance of kernel 3220, can access L2 cache 3228(i) contained in kernel 3220(i). In at least one embodiment, each kernel 3220, contained in kernel complex 3210(j), where j is an integer representing a single instance of kernel complex 3210, is connected to other kernels 3220 in kernel complex 3210(j) via L3 cache 3230(j), contained in kernel complex 3210(j). In at least one embodiment, kernels 3220, contained in kernel complex 3210(j), where j is an integer representing a single instance of kernel complex 3210, can access all L3 caches 3230(j) contained in kernel complex 3210(j). In at least one embodiment, L3-Cache 3230 can contain any number of slices without restriction.

[0297] In at least one embodiment, Fabric 3260 is a system assembly that facilitates data and control transmissions via core complexes 3210(1)-2710(N) (where N is an integer greater than zero), I / O interfaces 3270, and memory controllers 3280. In at least one embodiment, CPU 3200 can, without limitation, include any number and type of system assembly in addition to or instead of Fabric 3260, facilitating data and control transmissions via any number and type of directly or indirectly linked components, which may be located inside or outside CPU 3200. In at least one embodiment, I / O interfaces 3270 are representative of any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.).In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 3270. In at least one embodiment, peripheral devices coupled to I / O interfaces 3270 can include, without limitation, displays, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, and so on.

[0298] In at least one embodiment, memory controllers 3280 facilitate data transfers between the CPU 3200 and a system memory 3290. In at least one embodiment, the core complex 3210 and the graphics complex 3240 share system memory 3290. In at least one embodiment, the CPU 3200 implements a memory subsystem that includes, without limitation, any number and type of memory controllers 3280 and memory devices, which can be dedicated to a component or shared by multiple components. In at least one embodiment, the CPU 3200 implements a cache subsystem that includes, without limitation, one or more cache memories (e.g., L2 caches 3228 and L3 caches 3230), each of which can be private or shared by any number of components (e.g., cores 3220 and core complexes 3210).

[0299] In at least one embodiment, the CPU 3200 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0300] Fig. Figure 33 illustrates an exemplary accelerator integration slice 3390 according to at least one embodiment. As used herein, a “slice” comprises a specified portion of the processing resources of an accelerator integration circuit. In at least one embodiment, an accelerator integration circuit provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines included in a graphics acceleration module. Graphics processing engines may each comprise a separate GPU. Alternatively, graphics processing engines may comprise different types of graphics processing engines within a single GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, a graphics acceleration module may be a GPU containing multiple graphics processing engines.In at least one embodiment, graphics processing engines can be individual GPUs integrated on a common package, line card or chip.

[0301] An application-effective address space 3382 within system memory 3314 stores process elements 3383. In one embodiment, process elements 3383 are stored in response to GPU calls 3381 from applications 3380 running on processor 3307. A process element 3383 contains process state for the corresponding application 3380. A work descriptor (“WD”) 3384 contained in process element 3383 can be a single job requested by an application or it can contain a pointer to a queue of jobs. In at least one embodiment, WD 3384 is a pointer to a job request queue in application-effective address space 3382.

[0302] The Graphics Acceleration Module 3346 and / or individual graphics processing engines can be shared by all or a subset of processes in a system. In at least one embodiment, an infrastructure for setting process state and sending WD 3384 to the Graphics Acceleration Module 3346 to start a job in a virtualized environment can be included.

[0303] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns either the 3346 Graphics Acceleration Module or a custom graphics processing engine. Because the 3346 Graphics Acceleration Module is owned by a single process, a hypervisor initializes an accelerator integration circuit for an owned partition, and an operating system initializes an accelerator integration circuit for an owned process when the 3346 Graphics Acceleration Module is allocated.

[0304] During operation, a WD retrieval unit 3391 in accelerator integration slice 3390 retrieves the next WD 3384, which contains a display of work to be performed by one or more graphics processing engines of graphics acceleration module 3346. Data from WD 3384 can be stored in registers 3345 and used by a memory management unit (MMU) 3339, interrupt management circuit 3347, and / or context management circuit 3348, as illustrated. For example, one embodiment of MMU 3339 incorporates segment / page walk circuitry for accessing segment / page tables 3386 within OS virtual address space 3385. The interrupt management circuit 3347 can process interrupt events (INTs) 3392 received from graphics acceleration module 3346. When performing graphics operations, an effective address 3393, generated by a graphics processing engine, is translated into a real address by MMU 3339.

[0305] In one embodiment, an identical set of registers 3345 is duplicated for each graphics processing engine and / or graphics acceleration module 3346 and can be initialized by a hypervisor or operating system. Each of these duplicated registers can be included in accelerator integration slice 3390. Exemplary registers that can be initialized by a hypervisor are shown in Table 1. Table 1 - Hypervisor-initialized registers 1 Slice-Steuerregister 2 Reale Adresse (RA) Geplante Prozesse Gebietszeiger 3 Autoritätsmaskeüberschreibungsregister 4 Unterbrechungsvektortabelleneintragsoffset 5 Unterbrechungsvektortabelleneintragslimit 6 Zustandsregister 7 Logische Partitionierungs-ID 8 Real Address (RA) Hypervisor Accelerator Usage Entry Pointer 9 Memory description register

[0306] Examples of registers that can be initialized by an operating system are shown in Table 2. Table 2 - Operating system initialized registers 1 Process and thread identification 2 Effective Address (EA) Context Memory / Recovery Pointer 3 Virtual Address (VA) Accelerator Usage Entry Pointer 4 Virtual Address (VA) Memory Segment Table Pointer 5 Authority mask 6 Work descriptor

[0307] In one embodiment, each WD 3384 is specific to a single graphics acceleration module 3346 and / or a single graphics processing engine. It contains all the information a graphics processing engine needs to perform work, or it can be a pointer to a memory location where an application has set up a command queue for work to be completed.

[0308] In at least one embodiment, integration slice 3390 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0309] Fig.Figures 34A-29B illustrate exemplary graphics processing units (GPUs) according to at least one embodiment. In at least one embodiment, each of the exemplary GPUs can be manufactured using one or more IP cores. In addition to what is illustrated, other logic and circuitry can be included in at least one embodiment, including additional GPUs / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the exemplary GPUs are intended for use within a system-on-a-chip (SoC).

[0310] Fig. Figure 34A illustrates an exemplary graphics processor 3410 of an integrated SoC circuit which can be manufactured using one or more IP cores, according to at least one embodiment. Fig.Figure 34B illustrates an additional exemplary graphics processor 3440 of an integrated SoC circuit that can be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, graphics processor 3410 is of Fig. 34A is a low-performance graphics processor core. In at least one embodiment, graphics processor 3440 is of Fig. 34B is a graphics processor core with higher performance.

[0311] In at least one embodiment, the graphics processor 3410 includes a vertex processor 3405 and one or more fragment processors 3415A-2915N (e.g., 3415A, 3415B, 3415C, 3415D to 3415N-1 and 3415N). In at least one embodiment, the graphics processor 3410 can execute different shader programs via separate logic, such that the vertex processor 3405 is optimized to perform operations for vertex shader programs, while one or more fragment processors 3415A-2915N perform fragment shading operations (e.g., pixel shading operations) for fragment or pixel shader programs. In at least one embodiment, the vertex processor 3405 performs a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data.In at least one embodiment, fragment processor(s) 3415A-2915N use primitive and vertex data generated by vertex processor 3405 to create a frame buffer that is displayed on a display device. In at least one embodiment, fragment processor(s) 3415A-2915N are optimized to execute fragment shader programs, as provided in an OpenGL API, which can be used to perform operations similar to a pixel shader program, as provided in a Direct3D API.

[0312] In at least one embodiment, the graphics processor 3410 additionally includes one or more MMU(s) 3420A-2920B, cache(s) 3425A-2925B, and circuit assembly(s) 3430A-2930B. In at least one embodiment, one or more MMU(s) 3420A-2920B provide a mapping from virtual to physical addresses for the graphics processor 3410, including for the vertex processor 3405 and / or fragment processor(s) 3415A-2915N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more cache(s) 3425A-2925B.In at least one embodiment, one or more MMU(s) 3420A-2920B can be synchronized with other MMUs within a system, including one or more MMUs associated with one or more application processor(s), image processors, and / or video processors, so that each processor can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit clusters 3430A-2930B enable the 3410 graphics processor to connect to other IP cores within a SoC, either via an internal bus of the SoC or via a direct connection.

[0313] In at least one embodiment, graphics processor 3440 includes one or more MMU(s) 3420A-2920B, caches 3425A-2925B and circuit assemblies 3430A-2930B of graphics processor 3410. Fig.34A. In at least one embodiment, the graphics processor 3440 includes one or more shader core(s) 3455A-2955N (e.g., 3455A, 3455B, 3455C, 3455D, 3455E, 3455F to 3455N-1 and 3455N) that provide a unified shader core architecture in which a single core or type can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary.In at least one embodiment, the graphics processor 3440 includes an inter-core task manager 3445, which acts as a thread dispatcher to send execution threads to one or more shader cores 3455A-2955N, and a tiling unit 3458 to accelerate tiling operations for tiling-based rendering, in which rendering operations for a scene are divided into image space, for example to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0314] In at least one embodiment, graphics processors 3410 or 3440 can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0315] Fig.Figure 35A illustrates a graphics core 3500 according to at least one embodiment. In at least one embodiment, the graphics core 3500 can be integrated within a graphics processor 2410. Fig. 24. In at least one embodiment, graphics core 3500 can be a unified shader core 2955A-2955N as shown in Fig.29B. In at least one embodiment, graphics core 3500 includes a shared instruction cache 3502, a texture unit 3518, and a cache / shared memory 3520, which are common execution resources within graphics core 3500. In at least one embodiment, graphics core 3500 can include multiple slices 3501A-3001N or partitioning for each core, and a graphics processor can include multiple instances of graphics core 3500. Slices 3501A-3001N can include support logic, including a local instruction cache 3504A-3004N, a thread scheduler 3506A-3006N, a thread dispatcher 3508A-3008N, and a set of registers 3510A-3010N.In at least one embodiment, slices 3501A-3001N can include a set of additional functional units (“AFUs”) 3512A-3012N, floating-point units (“FPUs”) 3514A-3014N, arithmetic logic units (“ALUs”) 3516-3016N, address calculation units (“ACUs”) 3513A-3013N, double-precision floating-point units (“DPFPUs”) 3515A-3015N and matrix processing units (“MPUs”) 3517A-3017N.

[0316] In at least one embodiment, FPUs 3514A–3014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while DPFPUs 3515A–3015N perform double-precision (64-bit) floating-point operations. In at least one embodiment, ALUs 3516A–3016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, MPUs 3517A–3017N can also be configured for mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, MPUs 3517-3017N can perform a variety of matrix operations to accelerate CUDA programs, including enabling support for accelerated general matrix-matrix multiplication (“GEMM”).In at least one embodiment, AFUs 3512A -3012N can perform additional logic operations not supported by floating-point or integer units, including trigonometric operations (e.g. sine, cosine, etc.).

[0317] Fig.Figure 35B illustrates a general-purpose graphics processing unit (“GPGPU”) 3530 according to at least one embodiment. In at least one embodiment, GPGPU 3530 is highly parallel and suitable for use on a multi-chip module. In at least one embodiment, GPGPU 3530 can be configured to allow highly parallel computing operations to be performed by an array of GPUs. In at least one embodiment, GPGPU 3530 can be directly linked to other instances of GPGPU 3530 to create a multi-GPU cluster to improve execution time for CUDA programs. In at least one embodiment, GPGPU 3530 includes a host interface 3532 to enable communication with a host processor. In at least one embodiment, host interface 3532 is a PCIe interface.In at least one embodiment, host interface 3532 can be a vendor-specific communications interface or communications fabric. In at least one embodiment, GPGPU 3530 receives instructions from a host processor and uses a global scheduler 3534 to distribute execution threads associated with these instructions to a set of compute clusters 3536A-3036H. In at least one embodiment, compute clusters 3536A-3036H share a cache memory 3538. In at least one embodiment, cache memory 3538 can serve as a higher-level cache for cache memories within compute clusters 3536A-3036H.

[0318] In at least one embodiment, GPGPU 3530 includes memory 3544A-3044B coupled to compute clusters 3536A-3036H via a set of memory controllers 3542A-3042B. In at least one embodiment, memory 3544A-3044B can include various types of memory devices, including DRAM or graphics random-access memory, such as synchronous graphics random-access memory (“SGRAM”), including graphics double-rate data rate (“GDDR”) memory.

[0319] In at least one embodiment, computing clusters 3536A-3036H each include a set of graphics cores, such as graphics core 3500 from Fig.35A, which can include several types of integer and floating-point logic units capable of performing arithmetic operations with a range of precisions, including those suitable for calculations associated with CUDA programs. For example, in at least one embodiment, at least one subset of floating-point units in each of computation clusters 3536A–3036H can be configured to perform 16-bit or 32-bit floating-point operations, while another subset of floating-point units can be configured to perform 64-bit floating-point operations.

[0320] In at least one embodiment, multiple instances of GPGPU 3530 can be configured to operate as a computing cluster. In at least one embodiment, computing clusters 3536A-3036H can implement any technically feasible communication techniques for synchronization and data exchange. In at least one embodiment, multiple instances of GPGPU 3530 communicate via host interface 3532. In at least one embodiment, GPGPU 3530 includes an I / O hub 3539 that couples GPGPU 3530 to a GPU link 3540, enabling a direct connection to other instances of GPGPU 3530. In at least one embodiment, GPU link 3540 is coupled to a dedicated GPU-to-GPU bridge, enabling communication and synchronization between multiple instances of GPGPU 3530.In at least one embodiment, GPU Link 3540 connects to a high-speed network to transmit and receive data to and from other GPGPUs 3530 or parallel processors. In at least one embodiment, multiple instances of GPGPU 3530 reside in separate data processing systems and communicate via a network device accessible through Host Interface 3532. In at least one embodiment, GPU Link 3540 can be configured to provide a connection to a host processor in addition to, or as an alternative to, Host Interface 3532. In at least one embodiment, GPGPU 3530 can be configured to execute a CUDA program.

[0321] In at least one embodiment, the 3500 graphics core or 3530 GPGPU can be cooled by a coolant distribution system that implements a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described.

[0322] Fig. Figure 36A illustrates a parallel processor 3600 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 3600 can be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (“ASICs”), or FPGAs.

[0323] In at least one embodiment, the parallel processor 3600 includes a parallel processing unit 3602. In at least one embodiment, the parallel processing unit 3602 includes an I / O unit 3604, which enables communication with other devices, including other instances of the parallel processing unit 3602. In at least one embodiment, the I / O unit 3604 can be directly connected to other devices. In at least one embodiment, the I / O unit 3604 connects to other devices using a hub or switch interface, such as a memory hub 605. In at least one embodiment, connections between the memory hub 605 and the I / O unit 3604 form a communication link.In at least one embodiment, I / O unit 3604 connects to a host interface 3606 and a memory crossbar 3616, wherein host interface 3606 receives commands aimed at performing processing operations, and memory crossbar 3616 receives commands aimed at performing memory operations.

[0324] In at least one embodiment, when host interface 3606 receives a command buffer via I / O unit 3604, host interface 3606 can direct work operations to execute these commands to a front end 3608. In at least one embodiment, front end 3608 couples to a scheduler 3610, which is configured to distribute commands or other work items to a processing array 3612. In at least one embodiment, scheduler 3610 ensures that processing array 3612 is properly configured and in a valid state before distributing tasks to processing array 3612. In at least one embodiment, scheduler 3610 is implemented via firmware logic running on a microcontroller.In at least one embodiment, the microcontroller-implemented scheduler 3610 is configurable to perform complex scheduling and workload distribution operations with coarse and fine granularity, enabling fast preemption and context switching of threads running on the processing array 3612. In at least one embodiment, host software can schedule workloads to the processing array 3612 via one of several graphics processing doorbells. In at least one embodiment, workloads can then be automatically distributed across the processing array 3612 by scheduler 3610 logic within a microcontroller that incorporates the scheduler 3610.

[0325] In at least one embodiment, processing array 3612 can contain up to "N" clusters (e.g., cluster 3614A, cluster 3614B, up to cluster 3614N). In at least one embodiment, each cluster 3614A-3614N of processing array 3612 can execute a large number of concurrent threads. In at least one embodiment, scheduler 3610 can allocate work to clusters 3614A-3614N of processing array 3612 using various scheduling and / or workload allocation algorithms that can vary depending on the workload generated for each type of program or computation. In at least one embodiment, scheduling can be handled dynamically by scheduler 3610 or can be partially assisted by compiler logic during the compilation of program logic configured for execution by processing array 3612.In at least one embodiment, different clusters 3614A-3114N of processing array 3612 can be assigned to process different types of programs or to perform different types of calculations.

[0326] In at least one embodiment, processing array 3612 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing array 3612 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing array 3612 can include logic to perform processing tasks, including filtering video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.

[0327] In at least one embodiment, processing array 3612 is configured to perform parallel graphics processing operations. In at least one embodiment, processing array 3612 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing array 3612 may be configured to execute graphics processing-related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 3602 may transfer data from system memory via I / O unit 3604 for processing. In at least one embodiment, data transferred during processing may be written to on-chip memory (e.g., the memory chip) during processing.B. a parallel processor memory 3622) stored and then written back to system memory.

[0328] In at least one embodiment, when parallel processing unit 3602 is used to perform graphics processing, scheduler 3610 can be configured to divide a processing workload into approximately equal-sized tasks to better facilitate the distribution of graphics processing operations across multiple clusters 3614A-3114N of processing array 3612. In at least one embodiment, portions of processing array 3612 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to produce a rendered image for display.In at least one embodiment, intermediate data generated by one or more of the clusters 3614A-3114N can be stored in buffers to allow intermediate data to be transferred between the clusters 3614A-3114N for further processing.

[0329] In at least one embodiment, processing array 3612 can receive processing tasks to be executed via scheduler 3610, which receives commands defining processing tasks from frontend 3608. In at least one embodiment, processing tasks can include indices of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how data is to be processed (e.g., which program is to be executed). In at least one embodiment, scheduler 3610 can be configured to retrieve indices corresponding to tasks or can receive indices from frontend 3608. In at least one embodiment, frontend 3608 can be configured to ensure that processing array 3612 is configured in a valid state before a workload originating from incoming command buffers (e.g., batch buffers, push buffers, etc.) is executed.) is specified, is initiated.

[0330] In at least one embodiment, each of one or more instances of a parallel processing unit 3602 can be coupled to parallel processor memory 3622. In at least one embodiment, parallel processor memory 3622 can be accessed via a memory crossbar 3616, which can receive memory requests from a processing array 3612 and an I / O unit 3604. In at least one embodiment, the memory crossbar 3616 can access parallel processor memory 3622 via a memory interface 3618. In at least one embodiment, the memory interface 3618 can include several partitioning units (e.g., a partitioning unit 3620A, partitioning unit 3620B, up to partitioning unit 3620N), each of which can be coupled to a portion (e.g., a memory unit) of parallel processor memory 3622.In at least one embodiment, a number of partitioning units 3620A–3120N is configured to be equal to a number of storage units, such that a first partitioning unit 3620A has a corresponding first storage unit 3624A, a second partitioning unit 3620B has a corresponding storage unit 3624B, and an Nth partitioning unit 3620N has a corresponding Nth storage unit 3624N. In at least one embodiment, a number of partitioning units 3620A–3120N can be different from the number of storage devices.

[0331] In at least one embodiment, memory units 3624A-3124N can include various types of memory devices, including DRAM or graphics random-accessible memory, such as SGRAM, including GDDR memory. In at least one embodiment, memory units 3624A-3124N can also include 3D stacked memory, including, but not limited to, high-bandwidth memory (“HBM”). In at least one embodiment, render targets, such as frame buffers or texture maps, can be stored across memory units 3624A-3124N, allowing partitioning units 3620A-3120N to write portions of each render target in parallel to efficiently utilize available bandwidth of parallel processor memory 3622.In at least one embodiment, a local instance of parallel processor memory 3622 can be excluded in favor of a unified memory design that uses system memory in conjunction with local cache memory.

[0332] In at least one embodiment, each of the clusters 3614A-3114N of processing array 3612 can process data written to each of the storage units 3624A-3124N within parallel processor memory 3622. In at least one embodiment, memory crossbar 3616 can be configured to transfer an output from each cluster 3614A-3114N to any partitioning unit 3620A-3120N or to another cluster 3614A-3114N, which can perform additional processing operations on an output. In at least one embodiment, each cluster 3614A-3114N with memory interface 3618 can communicate through memory crossbar 3616 to read from or write to various external storage devices.In at least one embodiment, memory crossbar 3616 has a connection to memory interface 3618 for communication with I / O unit 3604, as well as a connection to a local instance of parallel processor memory 3622, which allows processing units within different clusters 3614A-3114N to communicate with system memory or other memory that is not local to parallel processing unit 3602. In at least one embodiment, memory crossbar 3616 can use virtual channels to separate traffic flows between clusters 3614A-3114N and partitioning units 3620A-3120N.

[0333] In at least one embodiment, multiple instances of Parallel Processing Unit 3602 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of Parallel Processing Unit 3602 can be configured to work together in a compatible manner, even if different instances include different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of Parallel Processing Unit 3602 can include floating-point units with higher precision relative to other instances.In at least one embodiment, systems incorporating one or more instances of Parallel Processing Unit 3602 or Parallel Processor 3600 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld personal computers, servers, workstations, game consoles and / or embedded systems.

[0334] Fig. Figure 36B illustrates a processing cluster 3694 according to at least one embodiment. In at least one embodiment, processing cluster 3694 is contained within a parallel processing unit. In at least one embodiment, processing cluster 3694 is one of processing clusters 3614A - 3114N of Fig.36. In at least one embodiment, Processing Cluster 3694 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program running on a particular set of input data. In at least one embodiment, single-instruction multiple data ("SIMD") instruction output techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single-instruction multiple thread ("SIMT") techniques are used to support the parallel execution of a large number of generally synchronized threads, using a common instruction unit that is configured to issue instructions to a set of processing engines within each Processing Cluster 3694.

[0335] In at least one embodiment, the operation of processing cluster 3694 can be controlled via a pipeline manager 3632, which distributes processing tasks to parallel SIMT processors. In at least one embodiment, pipeline manager 3632 receives instructions from scheduler 3610. Fig.36 and manages the execution of these instructions via a graphics multiprocessor 3634 and / or a texture unit 3636. In at least one embodiment, the graphics multiprocessor 3634 is an exemplary instance of a parallel SIMT processor. However, in at least one embodiment, different types of parallel SIMT processors of different architectures may be included within the processing cluster 3694. In at least one embodiment, one or more instances of the graphics multiprocessor 3634 may be included within the processing cluster 3694. In at least one embodiment, the graphics multiprocessor 3634 can process data, and a data crossbar 3640 can be used to distribute processed data to one of several possible destinations, including other shader units.In at least one embodiment, Pipelinemanager 3632 can facilitate the distribution of processed data by specifying targets for processed data to be distributed via data crossbar 3640.

[0336] In at least one embodiment, each graphics multiprocessor 3634 within processing cluster 3694 can include an identical set of functional execution logic (e.g., arithmetic logic units, load / store units (“LSUs”), etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner, allowing new instructions to be issued before previous instructions have fully executed. In at least one embodiment, functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same hardware functional unit can be used to perform different operations, and any combination of functional units can be present.

[0337] In at least one embodiment, instructions transmitted to processing cluster 3694 constitute a thread. In at least one embodiment, a set of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 3634. In at least one embodiment, a thread group can contain fewer threads than the number of processing engines within graphics multiprocessor 3634.In at least one embodiment, if a thread group contains fewer threads than the number of processing engines, one or more of the processing engines can be idle during cycles in which that thread group is being processed. In at least one embodiment, a thread group can also contain more threads than the number of processing engines within the 3634 graphics multiprocessor. In at least one embodiment, if a thread group contains more threads than the number of processing engines within the 3634 graphics multiprocessor, processing can be performed over successive clock cycles. In at least one embodiment, multiple thread groups can be executed concurrently on the 3634 graphics multiprocessor.

[0338] In at least one embodiment, the graphics multiprocessor 3634 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 3634 can dispense with an internal cache and use a cache memory (e.g., L1 cache 3648) within processing clusters 3694. In at least one embodiment, each graphics multiprocessor 3634 also has access to Level 2 ("L2") caches within partitioning units (e.g., partitioning units 3620A-3120N of Fig.36A), which are shared by all processing clusters 3694 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 3634 can also access off-chip global memory, which may include one or more local parallel processor memories and / or system memories. In at least one embodiment, any memory outside of the parallel processing unit 3602 can be used as global memory. In at least one embodiment, the processing cluster 3694 includes multiple instances of the graphics multiprocessor 3634 that can share common instructions and data, which can be stored in the L1 cache 3648.

[0339] In at least one embodiment, each processing cluster 3694 can include an MMU 3645 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 3645 can be located within memory interface 3618 of Fig.36 reside. In at least one embodiment, MMU 3645 includes a set of page table entries (“PTEs”) used to map a virtual address to a physical address of a tile and optionally a cache row index. In at least one embodiment, MMU 3645 may include address translation lookaside buffers (“TLBs”) or caches that may reside within graphics multiprocessor 3634, L1 cache 3648, or processing cluster 3694. In at least one embodiment, a physical address is processed to distribute surface data access locality to allow efficient request nesting between partitioning units. In at least one embodiment, a cache row index may be used to determine whether a request for a cache row is a hit or a miss.

[0340] In at least one embodiment, processing cluster 3694 can be configured such that each graphics multiprocessor 3634 is coupled to a texture unit 3636 for performing texture mapping operations, such as determining texture sampling positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 3634 and retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 3634 outputs a processed task to data crossbar 3640 to provide a processed task to another processing cluster 3694 for further processing or to store a processed task in an L2 cache, local parallel processor memory, or system memory via memory crossbar 3616.In at least one embodiment, a pre-raster operation unit (“preROP”) 3642 is configured to receive data from graphics multiprocessor 3634 and to forward data to ROP units which may be located with partitioning units as described herein (e.g. partitioning units 3620A-3120N of . Fig. 36). In at least one embodiment, PreROP 3642 can perform color blending optimizations, organize pixel color data, and perform address translations.

[0341] Fig. Figure 36C illustrates a graphics multiprocessor 3696 according to at least one embodiment. In at least one embodiment, graphics multiprocessor 3696 is graphics multiprocessor 3634. Fig.36B. In at least one embodiment, the graphics multiprocessor 3696 is coupled to the pipeline manager 3632 of the processing cluster 3694. In at least one embodiment, the graphics multiprocessor 3696 comprises an execution pipeline, including, but not limited to, an instruction cache 3652, an instruction unit 3654, an address mapping unit 3656, a register file 3658, one or more GPGPU cores 3662, and one or more LSUs 3666. The GPGPU cores 3662 and LSUs 3666 are coupled to the cache memory 3672 and shared memory 3670 via a memory and cache verb 3668.

[0342] In at least one embodiment, instruction cache 3652 receives a stream of instructions for executing pipeline manager 3632. In at least one embodiment, instructions are cached in instruction cache 3652 and sent for execution by instruction unit 3654. In at least one embodiment, instruction unit 3654 can send instructions as thread groups (e.g., warps), with each thread of a thread group being assigned to a different execution unit within GPGPU core 3662. In at least one embodiment, an instruction can access any address space from a local, shared, or global address space by specifying an address within a unified address space.In at least one embodiment, address mapping unit 3656 can be used to translate addresses in a unified address space into a different memory address that can be accessed by LSUs 3666.

[0343] In at least one embodiment, register file 3658 provides a set of registers for functional units of the graphics multiprocessor 3696. In at least one embodiment, register file 3658 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 3662, LSUs 3666) of the graphics multiprocessor 3696. In at least one embodiment, register file 3658 is subdivided between each of the functional units, such that each functional unit is allocated a dedicated portion of register file 3658. In at least one embodiment, register file 3658 is subdivided between different thread groups executed by the graphics multiprocessor 3696.

[0344] In at least one embodiment, GPGPU cores 3662 can each include FPUs and / or integer ALUs used to execute instructions from graphics multiprocessors 3696. GPGPU cores 3662 can be similar or different in architecture. In at least one embodiment, a first set of GPGPU cores 3662 includes a single-precision FPU and an integer ALU, while a second set of GPGPU cores 3662 includes a double-precision FPU. In at least one embodiment, FPUs can implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 3696 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copy rectangle or pixel blending operations.In at least one embodiment, one or more of the GPGPU cores 3662 can also include fixed or special function logic.

[0345] In at least one embodiment, GPGPU cores 3662 include SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 3662 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores 3662 can be generated at compile time by a shader compiler or automatically when programs written and compiled for single-program multiple data (SPMD) or SIMT architectures are executed. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction.For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel over a single SIMD8 logic unit.

[0346] In at least one embodiment, the memory and cache assembly 3668 is a network that connects each functional unit of the graphics multiprocessor 3696 to the register file 3658 and to shared memory 3670. In at least one embodiment, the memory and cache assembly 3668 is a crossbar assembly that allows the LSU 3666 to implement load and store operations between shared memory 3670 and the register file 3658. In at least one embodiment, the register file 3658 can operate at the same frequency as GPGPU cores 3662, thus enabling very low-latency data transfer between GPGPU cores 3662 and the register file 3658. In at least one embodiment, shared memory 3670 can be used to enable communication between threads running on functional units within the graphics multiprocessor 3696.In at least one embodiment, cache memory 3672 can, for example, be used as a data cache to temporarily store texture data communicated between functional units and texture unit 3636. In at least one embodiment, shared memory 3670 can also be used as a program-managed cache. In at least one embodiment, threads running on GPGPU cores 3662 can programmatically store data within shared memory in addition to automatically cached data stored within cache memory 3672.

[0347] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, a GPU can be communicatively coupled to host processor / cores via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, a GPU can be integrated on the same package or chip as cores and communicatively coupled to cores via a processor bus / interconnect that is internal to a package or chip. In at least one embodiment, regardless of how a GPU is connected, processor cores can assign work to a GPU in the form of sequences of instructions / directives contained in a workload diagram (WD).In at least one embodiment, a GPU then uses dedicated switching technology / logic to efficiently process these commands / instructions.

[0348] In at least one embodiment, different processors and processor cores can be used in Fig. 36A-C are cooled by a coolant distribution system implementing a physics-informed machine learning system, as in at least one embodiment described in Fig. 1-5 is described. General arithmetic

[0349] The following figures represent, without limitation, exemplary software constructs within general computing that can be used to implement at least one embodiment.

[0350] Fig.Figure 37 illustrates a software stack of a programming platform according to at least one embodiment. In at least one embodiment, a programming platform is a platform for utilizing hardware on a computing system to accelerate computational tasks. A programming platform may be accessible to software developers through libraries, compiler directives, and / or extensions of programming languages ​​in at least one embodiment. In at least one embodiment, a programming platform may be, but is not limited to, CUDA, Radeon Open Compute Platform (“ROCm”), OpenCL (OpenCL™ is developed by the Khronos group), SYCL, or Intel One API.

[0351] In at least one embodiment, a software stack 3700 of a programming platform provides an execution environment for an application 3701. In at least one embodiment, the application 3701 can include any computer software capable of being launched on the software stack 3700. In at least one embodiment, the application 3701 can include, but is not limited to, an artificial intelligence (“AI”) / machine learning (“ML”) application, a high-performance computing (“HPC”) application, a virtual desktop infrastructure (“VDI”), or a data center workload.

[0352] In at least one embodiment, application 3701 and software stack 3700 run on hardware 3707. Hardware 3707 can, in at least one embodiment, include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support a programming platform. In at least one embodiment, such as with CUDA, software stack 3700 can be vendor-specific and compatible only with devices from a single vendor or vendors. In at least one embodiment, such as with OpenCL, software stack 3700 can be used with devices from different vendors. In at least one embodiment, hardware 3707 includes a host that is connected to one or more devices that can be accessed to perform computing tasks via application programming interface (API) calls.A device within Hardware 3707 may, in at least one embodiment, include a GPU, FPGA, AI engine or other computing device (but may also include a CPU) and its memory, but is not limited to this, in contrast to a host within Hardware 3707, which may include a CPU (but may also include a computing device) and its memory, but is not limited to this.

[0353] In at least one embodiment, the software stack 3700 of a programming platform includes, without limitation, a number of libraries 3703, a runtime 3705, and a device kernel driver 3706. Each of the libraries 3703 may contain data and program code that can be used by computer programs and exploited during software development in at least one embodiment. In at least one embodiment, libraries 3703 may contain, but are not limited to, previously written code and subroutines, classes, values, type specifications, configuration data, documentation, auxiliary data, and / or message templates. In at least one embodiment, libraries 3703 include functions optimized for execution on one or more device types.In at least one embodiment, libraries 3703 may include, but are not limited to, functions for performing mathematical, deep learning, and / or other types of operations on devices. In at least one embodiment, libraries 3303 are associated with corresponding APIs 3302, which may include one or more APIs that expose functions implemented in libraries 3303.

[0354] In at least one embodiment, Application 3701 is written as source code which is compiled into executable code, as shown below in conjunction with Fig.Section 37 is discussed in more detail. Executable code of application 3701 can, in at least one embodiment, run at least partially on an execution environment provided by software stack 3700. In at least one embodiment, during the execution of application 3701, code may be required to run on a device, as opposed to a host. In such a case, runtime 3705 can be called to load and start the necessary code on a device, in at least one embodiment. In at least one embodiment, runtime 3705 can include any technically feasible runtime system capable of supporting the execution of application S01.

[0355] In at least one embodiment, Runtime 3705 is implemented as one or more runtime libraries associated with corresponding APIs shown as API 3704 or APIs 3704. One or more such runtime libraries may, without limitation, include, in at least one embodiment, functions for memory management, execution control, device management, error handling, and / or synchronization. In at least one embodiment, memory management functions may include, but are not limited to, functions for allocating, unallocating, and copying device memory, as well as for transferring data between host memory and device memory.In at least one embodiment, execution control functions may include, but are not limited to, functions for starting a function (sometimes called a "kernel" when a function is a global function that can be called by a host) on a device and setting attribute values ​​in a buffer maintained by a runtime library for a given function to be executed on a device.

[0356] Runtime libraries and corresponding APIs (3704) can be implemented in at least one embodiment in any technically feasible way. In at least one embodiment, one (or any number of) APIs can expose a low-level set of functions for fine-grained control of a device, while another (or any number of) APIs can expose a higher-level set of such functions. In at least one embodiment, a high-level runtime API can be built upon a low-level API. In at least one embodiment, one or more runtime APIs can be language-specific APIs layered on top of a language-independent runtime API.

[0357] In at least one embodiment, Device Kernel Driver 3706 is configured to facilitate communication with an underlying device. In at least one embodiment, Device Kernel Driver 3706 can provide low-level functionality upon which APIs, such as API 3704 or APIs 3704, and / or other software depend. In at least one embodiment, Device Kernel Driver 3706 can be configured to compile intermediate representation ("IR") code into binary code at runtime. For CUDA, Device Kernel Driver 3706 can compile parallel-threaded execution ("PTX") IR code, which is not hardware-specific, into binary code for a specific target device at runtime (with caching of compiled binary code), sometimes referred to as "finalizing" code, in at least one embodiment.This can allow finalized code to run on a target device that may not have existed when source code was originally compiled into PTX code, in at least one embodiment. Alternatively, in at least one embodiment, device source code can be compiled offline into binary code without requiring the device kernel driver to compile 3706 IR code at runtime.

[0358] Fig. Figure 38 illustrates a CUDA implementation of Software Stack 3200 from Fig.32 according to at least one embodiment. In at least one embodiment, a CUDA software stack 3800, on which an application 3801 can be started, includes CUDA libraries 3803, a CUDA runtime 3805, a CUDA driver 3807, and a device kernel driver 3808. In at least one embodiment, the CUDA software stack 3800 is executed on hardware 3809, which may include a GPU that supports CUDA and is developed by NVIDIA Corporation of Santa Clara, CA.

[0359] In at least one embodiment, application 3801, CUDA runtime 3805, and device kernel driver 3808 can perform similar functionalities to application 3201, runtime 3205, and device kernel driver 3206, respectively, as described above in conjunction with Fig.32. In at least one embodiment, CUDA Driver 3807 includes a library (libcuda.so) that implements a CUDA Driver API 3806. Similar to a CUDA Runtime API 3804 implemented by a CUDA Runtime Library (cudart), CUDA Driver API 3806, in at least one embodiment, can expose, among other things, functions for memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability. In at least one embodiment, CUDA Driver API 3806 differs from CUDA Runtime API 3804 in that CUDA Runtime API 3804 simplifies device code management by providing implicit initialization, context (analogous to a process) management, and module management (analogous to dynamically loaded libraries).In contrast to the high-level CUDA runtime API 3804, the CUDA driver API 3806, in at least one embodiment, is a low-level API that provides more granular control of a device, particularly with respect to contexts and module loading. In at least one embodiment, the CUDA driver API 3806 can expose context management functions not exposed by the CUDA runtime API 3804. In at least one embodiment, the CUDA driver API 3806 is also language-independent and supports, for example, OpenCL in addition to the support provided by the CUDA runtime API 3804. Furthermore, in at least one embodiment, development libraries, including the CUDA runtime 3805, can be considered separate from driver components, including the user-mode CUDA driver 3807 and the kernel-mode device driver 3808 (sometimes referred to as the "display" driver).

[0360] In at least one embodiment, CUDA libraries 3803 may include, but are not limited to, mathematical libraries, deep learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries that can utilize parallel computing applications, such as Application 3801. In at least one embodiment, CUDA libraries 3803 may include mathematical libraries such as, among others, a cuBLAS library, which is an implementation of Basic Linear Algebra Subprograms (“BLAS”) for performing linear algebra operations, a cuFFT library for computing Fast Fourier Transforms (“FFTs”), and a cuRAND library for generating random numbers.In at least one embodiment, CUDA libraries can include 3803 deep learning libraries, such as a cuDNN library of primitives for deep neural networks and a TensorRT platform for high-performance deep learning inference.

[0361] Fig. Figure 39 illustrates a ROCm implementation of software stack 3200 from Fig. 32 according to at least one embodiment. In at least one embodiment, an ROCm software stack 3900, on which an application 3901 can be started, includes a language runtime 3903, a system runtime 3905, a thunk 3907, an ROCm kernel driver 3908, and a device kernel driver 3909. In at least one embodiment, the ROCm software stack 3900 runs on hardware 3910, which may include a GPU that supports ROCm and is developed by AMD Corporation of Santa Clara, CA.

[0362] In at least one embodiment, application 3901 can perform similar functionalities to application 3201, which are described above in conjunction with Fig. 32 is discussed. Additionally, language runtime 3903 and system runtime 3905 can perform similar functionalities to runtime 3205, which are discussed above in conjunction with Fig.32 is discussed in at least one embodiment. In at least one embodiment, language runtime 3903 and system runtime 3905 differ in that system runtime 3905 is a language-independent runtime that implements an ROCr system runtime API 3904 and makes use of a Heterogeneous System Architecture (“HAS”) runtime API. The HAS runtime API is a thin user-mode API that, in at least one embodiment, exposes interfaces for accessing and interacting with an AMD GPU, including functions for memory management, execution control via architected kernel dispatch, error handling, system and agent information, and runtime initialization and shutdown. In contrast to system runtime 3905, language runtime 3903, in at least one embodiment, is an implementation of a language-specific runtime API 3902 layered on top of an ROCr system runtime API 3904.In at least one embodiment, the language runtime API may include, but is not limited to, a Heterogeneous Compute Interface for Portability (“HIP”) language runtime API, a Heterogeneous Compute Compiler (“HCC”) language runtime API, or an OpenCL API. In particular, HIP language is an extension of the C++ programming language with functionally similar versions of CUDA mechanisms, and in at least one embodiment, an HIP language runtime API includes functions similar to those of CUDA runtime API 3804, described above in conjunction with... Fig. 38 is discussed, including functions for memory management, execution control, device management, error handling and synchronization.

[0363] In at least one embodiment, Thunk (ROCt) 3907 is an interface that can be used to interact with the underlying ROCm driver 3908. In at least one embodiment, the ROCm driver 3908 is a ROCk driver that is a combination of an AMDGPU driver and a HAS kernel driver (amdkfd). In at least one embodiment, the AMDGPU driver is a device kernel driver for GPUs, developed by AMD, that performs similar functionalities to the device kernel driver 3206 described above in conjunction with Fig. 32 is discussed. In at least one embodiment, the HAS kernel driver is a driver that allows different types of processors to share system resources more effectively via hardware features.

[0364] In at least one embodiment, various libraries (not shown) can be included in ROCm software stack 3900 via language runtime 3903 and provide functionality similar to CUDA libraries 3803, which are described above in conjunction with Fig. 38 are discussed. In at least one embodiment, various libraries may include, but are not limited to, mathematical, deep learning and / or other libraries, such as, among others, a hipBLAS library that implements functions similar to those of CUDA cuBLAS, and a rocFFT library for calculating FFTs, similar to CUDA cuFFT.

[0365] Fig. Figure 40 illustrates an OpenCL implementation of Software Stack 3200 from Fig.32 according to at least one embodiment. In at least one embodiment, an OpenCL software stack 4000, on which an application 4001 can be started, includes an OpenCL frame 4005, an OpenCL runtime 4006, and a device kernel driver 4008. In at least one embodiment, the OpenCL software stack 4000 is executed on hardware 3309 that is not vendor-specific. Because OpenCL is supported by devices developed by different vendors, specific OpenCL drivers may be required to operate compatibly with hardware from such vendors in at least one embodiment.

[0366] In at least one embodiment, application 4001, OpenCL runtime 4006, device kernel driver 4008, and hardware 4009 can perform similar functionalities to application 3201, runtime 3205, device kernel driver 3206, and hardware 3207, respectively, as described above in conjunction with Fig.32 are discussed. In at least one embodiment, application 4001 further includes an OpenCL kernel 4002 with code to be executed on a device.

[0367] In at least one embodiment, OpenCL defines a "platform" that allows a host to control devices connected to it. In at least one embodiment, an OpenCL framework provides a platform layer API and a runtime API, shown as Platform API 4003 and Runtime API 4007. In at least one embodiment, Runtime API 4005 uses contexts to manage kernel execution on devices. In at least one embodiment, each identified device can be associated with a specific context that Runtime API 4007 can use to manage instruction queues, program objects, and kernel objects, and to share memory objects, among other things, for that device.In at least one embodiment, Platform API 4003 exposes functions that, among other things, allow the use of device contexts to select and initialize devices, submit work to devices via command queues, and enable data transfer to and from devices. Additionally, the OpenCL framework provides various built-in functions (not shown), including mathematical functions, relational functions, and image processing functions, among others, in at least one embodiment.

[0368] In at least one embodiment, a 4004 compiler is also included in the 4007 OpenCL frame. Source code can be compiled offline before an application is executed, or online during application execution, in at least one embodiment. Unlike CUDA and ROCm, OpenCL applications in at least one embodiment can be compiled online by the 4004 compiler, which is included to be representative of any number of compilers that can be used to compile source code and / or IR code, such as Standard Portable Intermediate Representation (SPIR-V) code, into binary code. Alternatively, in at least one embodiment, OpenCL applications can be compiled offline before such applications are executed.

[0369] Fig.Figure 41 illustrates software supported by a programming platform according to at least one embodiment. In at least one embodiment, a programming platform 4104 is configured to support various programming models 4103, libraries and / or middleware 4102, and frameworks 4101 upon which an application 4100 can rely. In at least one embodiment, the application 4100 can be an AI / ML application implemented, for example, using a deep learning framework such as MXNet, PyTorch, or TensorFlow, which may rely on libraries such as cuDNN, NVIDIA Collective Communications Library (“NCCL”), and / or NVIDIA Developer Data Loading Library (“DALI”) CUDA libraries to provide accelerated computing on underlying hardware.

[0370] In at least one embodiment, programming platform 4104 can be a CUDA, ROCm or OpenCL platform, each of which is described above in conjunction with Fig. 33, Fig. 34 and Fig. 35. In at least one embodiment, programming platform 4104 supports several programming models 4103, which are abstractions of an underlying computer system that allow expressions of algorithms and data structures. Programming models 4103 can expose features of underlying hardware to improve performance, in at least one embodiment. In at least one embodiment, programming models 4103 can include, but are not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism (“C++AMP”), Open Multi-Processing (“OpenMP”), Open Accelerators (“OpenACC”), and / or Vulcan Compute.

[0371] In at least one embodiment, libraries and / or middleware 4102 provide implementations of abstractions of programming models 4104. In at least one embodiment, such libraries include data and programming code that can be used by computer programs and during software development. In at least one embodiment, such middleware includes software that provides services for applications beyond those available from programming platform 4104. In at least one embodiment, libraries and / or middleware 4102 may include, but are not limited to, cuBLAS, cuFFT, cuRAND, and other CUDA libraries or rocBLAS, rocFFT, rocRAND, and other ROCm libraries.Additionally, libraries and / or middlewares 4102 may, in at least one embodiment, include NCCL and ROCm Communication Collectives Library (“RCCL”) libraries that provide communication routines for GPUs, an MIOpen library for deep learning acceleration, and / or a proprietary library for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.

[0372] In at least one embodiment, application frameworks 4101 depend on libraries and / or middlewares 4102. In at least one embodiment, each application framework 4101 is a software framework used to implement a standard structure of application software. An AI / ML application can be implemented using a framework such as Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks in at least one embodiment.

[0373] Fig.Figure 42 illustrates compiling code for execution on a programming platform of Fig. 37...

Claims

A processor comprising: one or more circuits to cause one or more neural networks trained to simulate one or more types of data center hardware, each to jointly simulate a data center cooling system with the one or more types of data center hardware. The processor according to claim 1, wherein at least one of the one or more neural networks is trained using a dataset comprising one or more data center conditions, input and output characteristics of a corresponding type of data center hardware, and performance data, wherein the data center conditions comprise one or more of: a number of server racks, a heat load for each server rack, estimated capital and / or operating costs, and wherein the performance data comprise cooling performance and configuration data specific to the corresponding type of data input hardware. The processor according to claim 1, wherein the one or more neural networks comprise a first neural network trained to simulate a coolant distribution unit (CDU), wherein the first neural network further comprises one or more sub-neural networks, and wherein the one or more sub-neural network comprises a first sub-neural network to simulate a CDU pump, a second sub-neural network to simulate a heat exchanger, a third sub-neural network to simulate one or more valves and filters, and a fourth sub-neural network to simulate CDU controls. The processor according to claim 1, wherein the one or more neural networks comprise a second neural network trained to simulate the operation of an air cooling unit that cools a hot airflow from the data center cooling system. The processor according to claim 1, wherein the one or more neural networks corresponding to the one or more types of data center hardware are integrated based at least partially on a data input cooling system architecture. The processor according to claim 1, wherein the one or more circuits are configured to dynamically update an integration of the one or more neural networks to jointly simulate an updated data center cooling system in response to a change in the data center cooling system. The processor according to claim 6, wherein the modification of the data center cooling system comprises one or more of: the addition of new data center hardware; the removal of one or more types of data center hardware; the replacement of one or more types of data center hardware; a modification of an arrangement of one or more types of data center hardware; and updated control parameters associated with one or more types of data center hardware. The processor according to claim 7, wherein at least one neural network of the one or more neural networks is caused to automatically generate one or more control parameters for at least one type of data center hardware based at least partially on one or more feedback from one or more simulation results of the data center cooling system. A system comprising: one or more processors to cause one or more neural networks trained to simulate one or more types of data center hardware, each to jointly simulate a data center cooling system with the one or more types of data center hardware. The system according to claim 9, wherein at least one of the one or more neural networks is trained using a dataset comprising one or more data center conditions, input and output characteristics of a corresponding type of data center hardware, and performance data, and wherein the performance data includes cooling performance and configuration data specific to the corresponding type of data input hardware, and wherein the data center conditions include one or more of: a number of server racks, a heat load for each server rack, estimated capital and / or operating costs for each server rack, and wherein the performance data includes cooling performance and configuration data specific to the corresponding type of data input hardware. The system according to claim 9, wherein the one or more neural networks comprise a first neural network trained to simulate the CDU, and wherein the first neural network further comprises one or more sub-neural networks, and wherein the one or more sub-neural networks comprise a first sub-neural network to simulate a CDU pump, a second sub-neural network to simulate a heat exchanger, a third sub-neural network to simulate one or more valves and filters, and a fourth sub-neural network to simulate CDU controls. The system according to claim 9, wherein one or more neural networks corresponding to one or more types of data center hardware are integrated based at least partially on a data input cooling system architecture. The system according to claim 9, wherein the one or more circuits are configured to dynamically update an integration of one or more neural networks to jointly simulate an updated data center cooling system in response to a change in the data center cooling system. The system according to claim 13, wherein the modification of the data center cooling system comprises one or more of: the addition of new data center hardware; the removal of one or more types of data center hardware; the replacement of one or more types of data center hardware; a modification of an arrangement of one or more types of data center hardware; and updated control parameters associated with one or more types of data center hardware. The system according to claim 14, wherein at least one neural network of the one or more neural networks is caused to automatically generate one or more control parameters for at least one type of data center hardware based at least partially on one or more feedback from one or more simulation results of the data center cooling system. A method comprising: inducing one or more neural networks trained to simulate one or more types of data center hardware to jointly simulate a data center cooling system with the one or more types of data center hardware. The method according to claim 16, wherein at least one of the one or more neural networks is trained using a dataset comprising one or more data center conditions, input and output characteristics of a corresponding type of data center hardware and performance data, and wherein the performance data comprises cooling performance and configuration data for the corresponding type of data input hardware, and wherein the data center conditions comprise one or more of: a number of server racks, a heat load for each server rack, and wherein the performance data comprises cooling performance and configuration data specific to the corresponding type of data input hardware. The method according to claim 16, wherein the one or more neural networks comprise a first neural network trained to simulate the CDU, and wherein the first neural network further comprises one or more sub-neural networks, and wherein the one or more sub-neural networks comprise a first sub-neural network to simulate a CDU pump, a second sub-neural network to simulate a heat exchanger, a third sub-neural network to simulate one or more valves and filters, and a fourth sub-neural network to simulate CDU controls. The method according to claim 16, wherein one or more neural networks corresponding to one or more types of data center hardware are integrated based at least partially on a data input cooling system architecture. The method according to claim 16, further comprising: dynamically updating an integration of one or more neural networks to jointly simulate an updated data center cooling system in response to a change in the data center cooling system, and wherein the change in the data center cooling system comprises one or more of: the addition of new data center hardware; the removal of one or more types of data center hardware; the replacement of one or more types of data center hardware; a modification of an arrangement of one or more types of data center hardware; and updated control parameters associated with one or more types of data center hardware.