Liquid flow distribution using one or more neural networks
By adopting a neural network-based cooling system in the data center, using the NNCFD system to predict future states and dynamically adjust liquid flow and temperature, the temperature regulation challenges caused by changes in coolant flow in existing cooling systems are solved, achieving finer temperature control and higher cooling efficiency.
Patent Information
- Application Number
- CN202111628582.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2021-12-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-28
AI Technical Summary
Existing cooling systems have imbalanced designs in data centers, resulting in changes in coolant flow, increasing the challenges of temperature regulation, especially in systems lacking fine-grained flow regulation.
Using a neural network-based cooling system, the neural network computing fluid dynamics (NNCFD) system is used to predict the future state of the computing environment, dynamically adjust the liquid flow and temperature to achieve finer temperature control.
By predicting future states, neural network cooling systems can actively adjust the cooling system, reduce temperature fluctuations, improve cooling efficiency, and avoid hardware damage and data loss.
Smart Images

Figure CN114698335B_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment is directed to processing resources for performing and facilitating artificial intelligence. For example, at least one embodiment is directed to a processor or computing system for training a neural network according to the various new techniques described herein. Background Art
[0002] In computing environments such as data centers, control infrastructure is used to manage aspects such as coolant distribution and temperature control. Various sensors or monitoring components may be used to determine the current state of the computing environment, and this state information may be used to determine adjustments to be made to the control infrastructure. For cooling systems, this may include adjusting the flow of air or liquid through specific cooling components. The purely reactive nature of these approaches may result in undesirable events such as excessive thermal events, which may result in hardware damage or data loss. Additionally, the unbalanced design of these cooling systems can result in variations in coolant flow throughout the system, making temperature regulation more challenging, especially for systems that lack fine-grained flow regulation at various points throughout the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1A and 1B illustrates a computing environment according to at least one embodiment;
[0004] Figure 2A and 2B illustrates temperature distribution in a data center according to at least one embodiment;
[0005] Figure 3 illustrates a server rack with a cooling system supporting a neural network according to at least one embodiment;
[0006] Figure 4A , 4B 4C and 4D illustrate a data center and server racks with a cooling system supporting a neural network according to at least one embodiment;
[0007] Figure 5A and 5B A process for regulating liquid flow in a cooling system according to at least one embodiment is shown;
[0008] Figure 6 A distributed system according to at least one embodiment is shown;
[0009] Figure 7 illustrates an exemplary data center in accordance with at least one embodiment;
[0010] Figure 8 illustrates a client-server network according to at least one embodiment;
[0011] Fig. 9 illustrates a computer network according to at least one embodiment;
[0012] Fig. 10A illustrates a networked computer system according to at least one embodiment;
[0013] Fig. 10B illustrates a networked computer system according to at least one embodiment;
[0014] Fig. 10C illustrates a networked computer system according to at least one embodiment;
[0015] Fig.11 One or more components of a system environment according to at least one embodiment are shown, wherein a service may be provided as a third party web service;
[0016] Fig.12 illustrates a cloud computing environment according to at least one embodiment;
[0017] Fig.13 A set of functional abstraction layers provided by a cloud computing environment according to at least one embodiment is shown;
[0018] Fig.14 A supercomputer at a chip level is shown according to at least one embodiment;
[0019] Fig.15 illustrates a supercomputer at a rack module level according to at least one embodiment;
[0020] Fig.16 illustrates a rack-scale supercomputer according to at least one embodiment;
[0021] Fig.17 illustrates a supercomputer at an overall system level according to at least one embodiment;
[0022] Fig.18A Inference and / or training logic according to at least one embodiment is shown;
[0023] Fig.18B Inference and / or training logic according to at least one embodiment is shown;
[0024] Fig.19 illustrates the training and deployment of a neural network according to at least one embodiment;
[0025] Fig. 20 illustrates the architecture of a network system according to at least one embodiment;
[0026] Fig.21 illustrates the architecture of a network system according to at least one embodiment;
[0027] Fig. 22 A control plane protocol stack is shown according to at least one embodiment;
[0028] Fig.23 illustrates a user plane protocol stack according to at least one embodiment;
[0029] Fig.24 Components of a core network according to at least one embodiment are shown;
[0030] Fig.25 Components of a system supporting network function virtualization (NFV) according to at least one embodiment are shown;
[0031] Fig.26 A processing system according to at least one embodiment is shown;
[0032] Fig. 27 A computer system according to at least one embodiment is shown;
[0033] Fig.28 A system according to at least one embodiment is shown;
[0034] Fig.29 An exemplary integrated circuit according to at least one embodiment is shown;
[0035] Fig.30 A computing system according to at least one embodiment is shown;
[0036] Fig.31 An APU is shown according to at least one embodiment;
[0037] Fig.32 A CPU according to at least one embodiment is shown;
[0038] Fig.33 An exemplary accelerator integrated slice is shown in accordance with at least one embodiment;
[0039] Figures 34A-34B An exemplary graphics processor is shown in accordance with at least one embodiment;
[0040] Fig.35A illustrates a graphics core according to at least one embodiment;
[0041] Fig.35B A GPGPU is shown according to at least one embodiment;
[0042] Fig.36A A parallel processor according to at least one embodiment is shown;
[0043] Fig.36Billustrates a processing cluster according to at least one embodiment;
[0044] Fig.36C A graphics multiprocessor is shown in accordance with at least one embodiment;
[0045] Fig.37 A software stack for a programming platform according to at least one embodiment is shown;
[0046] Fig.38 According to at least one embodiment, Fig.37 CUDA implementation of the software stack;
[0047] Fig.39 According to at least one embodiment, Fig.37 ROCm implementation of the software stack;
[0048] Fig.40 According to at least one embodiment, Fig.37 OpenCL implementation of the software stack;
[0049] Fig.41 illustrates software supported by a programming platform according to at least one embodiment; and
[0050] Fig.42 A method for performing Figure 37-40 Compiled code that is executed on the programming platform of DETAILED DESCRIPTION
[0051] In at least one embodiment, the computing environment may include various computing devices and control systems, such as Figure 1A100. In at least one embodiment, the data center 100 may include one or more rooms 102 having racks 110 and auxiliary equipment for accommodating one or more servers on one or more server trays. In at least one embodiment, the data center 100 is supported by a cooling tower 104 located outside the data center 100. In at least one embodiment, the cooling tower 104 dissipates heat from within the data center 100 by acting on the main cooling loop 106. In at least one embodiment, a cooling distribution unit (CDU) 112 is used between the main cooling loop 106 and the second or auxiliary cooling loop 108 to enable heat to be extracted from the second or auxiliary cooling loop 108 to the main cooling loop 106. In at least one embodiment, the auxiliary cooling loop 108 can access different pipes entering the server tray as needed. In at least one embodiment, the loops 106, 108 are shown as line diagrams, but ordinary technicians will recognize that one or more pipe features can be used. In at least one embodiment, flexible polyvinyl chloride (PVC) pipes can be used with associated piping systems to move fluids along each of the loops 106, 108. In at least one embodiment, one or more coolant pumps may be used to maintain a pressure differential within the loops 106, 108 to enable coolant to move based on temperature sensors in various locations, including in the room, in one or more racks 110, and / or in server boxes or server trays within the racks 110.
[0052] In at least one embodiment, the coolant in the primary cooling loop 106 and the auxiliary cooling loop 108 can be at least water (deionized water or other) and an additive, such as ethylene glycol or propylene glycol. In operation, in at least one embodiment, each of the primary cooling loop and the auxiliary cooling loop has its own coolant. In at least one embodiment, the coolant in the auxiliary cooling loop can be dedicated to the requirements of the components in the server tray or rack 110. In at least one embodiment, the CDU 112 is capable of complex control of the coolant in the loops 106, 108 independently or simultaneously. In at least one embodiment, the CDU can be adapted to control the flow rate so that the coolant is appropriately distributed to extract the heat generated within the rack 110. In at least one embodiment, more flexible ducting 114 is provided from the auxiliary cooling loop 108 to enter each server tray and provide coolant to the electrical and / or computing components.
[0053] In at least one embodiment, electrical and / or computing components are used interchangeably to refer to heat generating components that benefit from a data center cooling system.In at least one embodiment, the tubing 118 forming part of the auxiliary cooling loop 108 may be referred to as one or more room manifolds.
[0054] In at least one embodiment, the line 116 extending from the line 118 may also be part of the auxiliary cooling loop 108, but may be referred to as a row manifold. In at least one embodiment, the line 114 enters the rack as part of the auxiliary cooling loop 108, but may be referred to as a rack cooling manifold. In at least one embodiment, the row manifold 116 extends along the rows in the data center 100 to all racks. In at least one embodiment, the control of the auxiliary cooling loop 108 including the manifolds 118, 116, and 114 may be improved. In at least one embodiment, a chiller 120 may be provided in the main cooling loop within the data center 102 to support cooling before the cooling tower. In at least one embodiment, to the extent that additional loops are present in the main control loop, these additional loops may provide cooling external to the rack and external to the auxiliary cooling loop and may be combined with the main cooling loop.
[0055] In at least one embodiment, in operation, heat generated within the server trays of the rack 110 can be transferred to the coolant exiting the rack 110 via the flexible tubing of the row manifold 114 of the auxiliary cooling loop 108. In at least one embodiment, the secondary coolant (in the auxiliary cooling loop 108) from the CDU 112 for cooling the rack 110 moves toward the rack 110. In at least one embodiment, the secondary coolant from the CDU 112 is transferred from one side of the room manifold with tubing 118 to one side of the rack 110 via the row manifold 116 and passes through one side of the server tray via tubing 114. In at least one embodiment, the spent or returned secondary coolant (or the exiting secondary coolant that carries heat away from the computing components) exits from the other side of the server tray (such as entering the left side of the rack and exiting the right side of the rack for the server tray after circulating through the server tray or through the components on the server tray). In at least one embodiment, the spent second coolant exiting the server tray or rack 110 comes out of a different side (such as the exit side) of the tubes 114 and moves to a parallel, also exit side, of the row manifold 116. In at least one embodiment, the spent second coolant moves from the row manifold 116 in a parallel portion of the row manifold 118 that travels in an opposite direction to the incoming second coolant (which may also be the refreshed second coolant), and toward the CDU 112.
[0056] In at least one embodiment, the spent secondary coolant exchanges its heat with the primary coolant in the primary cooling loop 106 via the CDU 112. In at least one embodiment, the spent secondary coolant is refreshed (such as relatively cool when compared to the temperature of the spent second coolant stage) and is ready to circulate back through the auxiliary cooling loop 108 to the computing components. In at least one embodiment, various flow and temperature control features in the CDU 112 enable control of the heat exchanged from the spent secondary coolant or the flow of secondary coolant into and out of the CDU 112. The CDU 112 is also capable of controlling the flow of the primary coolant in the primary cooling loop 106.
[0057] In at least one embodiment, the operation of components within a data center can produce patterns of temperature changes and heat flow. In at least one embodiment, a cooling system can utilize one or more fluids (e.g., liquid or air) to attempt to manage temperature changes and levels across such a computing environment. In at least one embodiment, fluid-based cooling of high-density servers can be used to manage sudden high heat changes caused by varying computing loads or other such factors. In at least one embodiment, since the demand may change or tend to range from minimum to maximum different cooling requirements, a suitable cooling system must be used to meet these requirements in an economical manner. In at least one embodiment, a liquid cooling system can be used for medium to high cooling requirements. In at least one embodiment, different cooling requirements also reflect different thermal characteristics of the data center. In at least one embodiment, the heat generated from these components, servers, and racks is cumulatively referred to as a thermal characteristic or cooling requirement because the cooling requirement must fully address the thermal characteristic.
[0058] In at least one embodiment, temperature and airflow may also be varied within a computing device, such as for Figure 1B 150. In at least one embodiment, a pattern of airflow 152 can be determined within a computing device, which is shown here without at least a top cover to show the flow between various components within the device. Various modeling methods can be used to determine airflow, power distribution, temperature changes, and other such aspects of such a device during operation.
[0059] In at least one embodiment, sensors and other such devices may be used to measure aspects such as fluid flow, pressure, airflow, and temperature within a server, rack, data center, or other such component or environment. In at least one embodiment, data from these aspects may be used to model the current state of the computing environment, such as for Figure 2A200. In at least one embodiment, there are a plurality of racks 202 or other groupings of computer servers, each of which includes one or more fans or other such elements that can force air out of those servers. In at least one embodiment, this can be modeled as a pattern 204 of airflow and associated temperatures. In at least one embodiment, there can be various events that can lead to unacceptable conditions, such as periods of abnormally high computing load or external thermal events, such as Figure 2B In at least one embodiment, darker shaded areas may correspond to areas of higher temperature, and it can be seen that in Figure 2B In at least one embodiment, this may result in temperatures exceeding a maximum operating temperature threshold, and therefore may result in damage to at least some of the computing devices or related infrastructure. In at least one embodiment, the use of a neural network to dynamically predict such occurrences or events in near real time, and to be able to make adjustments proactively rather than reactively, can help reduce the occurrence or extent of such events, or at least reduce the impact of such events on the servers or related infrastructure. In at least one embodiment, this may include adjustments to be made to one or more flow control valves of a cooling system to regulate the flow of liquid to maintain a desired operating temperature in the data center or environment.
[0060] In at least one embodiment, a cooling system is disclosed that addresses thermal features in an associated computing or data center device, such as in a graphics processing unit (GPU), in a switch, in a dual in-line memory module (DIMM), or a central processing unit (CPU). In addition, in at least one embodiment, the associated computing or data center device may include a processing card having one or more GPUs, switches, or CPUs thereon. In at least one embodiment, each of these GPUs, switches, and CPUs may be a thermal feature of the computing device. In at least one embodiment, the GPU, CPU, or switch may have one or more cores, and each core may be a thermal feature.
[0061] In at least one embodiment, such a cooling system can utilize one or more neural networks with various systems, devices, components, boards or cards within a computing environment to provide predictions of one or more future states of the environment. In at least one embodiment, this may involve the use of a neural network computational fluid dynamics (NNCFD) system that provides the ability to predict future states within a computing environment, such as temperature, pressure or flow rate at one or more locations in one or more future time periods. In at least one embodiment, these predictions can be inferences made using data about the performance of data center equipment and information such as temperature, airflow and pressure at various locations throughout the environment. In at least one embodiment, the control device may include or execute an NNCFD application that is capable of collecting various heat, fluid flow, power and environmental data from various locations throughout the computing environment, such as a data center, a group of data centers, a server rack or a single server, and other devices and systems discussed and suggested herein. In at least one embodiment, NNCFD can be trained to continue learning over time, thereby providing accurate predictive modeling using various agents running in devices across the environment. In at least one embodiment, this can include one or more types of individual agents, such as a board or ASIC with built-in machine learning capabilities. In at least one embodiment, these individual agents can perform tasks such as regulating cooling (e.g., fan speed control or liquid flow control), power distribution (e.g., power distribution units (PDUs) or uninterruptible power supplies (UPS)), and facility infrastructure (e.g., chillers or cooling towers) control based at least in part on inferences from this predictive NNCFD to best anticipate any impending significant power or thermal events in order to respond to them smoothly and, in many cases, prevent them from occurring. In at least one embodiment, the NNCFD application can provide an intelligent, predictive data center building management system (BMS) or data center infrastructure management (DCIM). In at least one embodiment, such a system can be designed using specific hardware that enables machine learning capabilities to be directly incorporated into various computing devices or systems, such as NVIDIA's Jetson board, which can utilize one or more graphics processing units (GPUs) on the board to efficiently provide neural network-based reasoning and predictions.
[0062] In at least one embodiment, a neural network can be trained to solve fluid mechanics and dynamics problems that might otherwise be solved using a set of complex equations. In at least one embodiment, such equations are given by:
[0063]
[0064] Wherein the sum of the unstable component and the convective component is equal to the sum of the diffusion component and the generative component. In at least one embodiment, such an equation is a function of, for example, continuity, momentum, and energy. In at least one embodiment, the goal of solving such an equation is to obtain the fluid velocity value (air or liquid), temperature value, and pressure value at any point in the computing environment at any current or future time point. In at least one embodiment, a neural network-based architecture can be used to solve such computational mechanics or physics problems. In at least one embodiment, this can be achieved using a point cloud representation of three-dimensional (3D) geometry and a multi-physics network that maintains respect for the control partial differential equation (PDE). In at least one embodiment, the performance of such neural network-based modeling can be optimized for various configurations, such as multi-node configurations, multi-GPU configurations, automatic mixed precision (AMP) configurations, or accelerated linear algebra (XLA) configurations, such as those provided by NVIDIA. In at least one embodiment, such a neural network can be trained using a loss function that considers aspects such as partial differential equations (PDEs), boundary conditions (BCs), initial conditions (ICs), and environmental data. In at least one embodiment, such a network can be trained using an unsupervised, physically driven approach. In at least one embodiment, a model can be trained for each type of device to be used, and updated network parameters can be used, such as those obtained over time from further training on additional data. In at least one embodiment, a neural network of an appropriate type suitable for equation solving and physics-based analysis can be used, such as SIMNet from NVIDIA. In at least one embodiment, SIMNet provides a simulation toolkit based on physics and machine learning. In at least one embodiment, SIMNet provides a framework for modeling partial differential equations (PDEs) together with boundary conditions and initial conditions. In at least one embodiment, flow problems can be partially solved by modeling mass balance conditions as hard constraints and global constraints, which can improve accuracy and convergence characteristics. In at least one embodiment, for multi-physics problems across multiple domains, separate networks for different physics coupled at domain interfaces also provide acceptable performance.
[0065] In at least one embodiment, neural networks can be embedded in various devices to enable those devices to predict future states and make proactive adjustments based at least in part on those predictions. In at least one embodiment, this can include installing directly in a device such as Figure 3A neural network-enabled device in a liquid-cooled server rack 300 is shown. In at least one embodiment, the rack 300 may include a plurality of liquid-cooled servers 302 or other such devices. In at least one embodiment, a neural network-enabled rack manifold 304 may be included in the rack 300 to provide a flow of liquid into each liquid-cooled server 302 through an inlet valve 308 and return the liquid with heat removed from the server 302 through an outlet valve 310. In at least one embodiment, the rack manifold 304 may include a plurality of neural network-enabled auxiliary liquid distribution control boards 306. In at least one embodiment, there may be one such control board 306 for each inlet valve 308 and outlet valve 310 for each server. In at least one embodiment, the rack 300 may also include a neural network-enabled rack power distribution unit (PDU) 312, which may include a plurality of outlets designed to distribute power to computers or network devices within the rack 300. In at least one embodiment, there may be a neural network-enabled power distribution control board 314 associated with each outlet of the PDU 312. In at least one embodiment, sensors may capture information about temperature, airflow, or other such aspects of the computing environment inside and / or outside the rack 300, including inside and / or outside any individual server 302 located therein. In at least one embodiment, fluid cooling may remove a certain amount of heat from the servers 302, but due to factors such as varying loads and external temperature fluctuations, temperatures at various locations may vary and may reach or exceed temperature limits at which such devices can continue to operate properly. In at least one embodiment, attempts may be made to ensure that temperatures at specific locations remain below acceptable limits, where such locations may relate to junction or core temperatures of a processor (e.g., a CPU or GPU), a memory module, or a power supply.
[0066] In at least one embodiment, the neural networks in these control boards 306, 314 can be responsible for predicting values such as coolant flow or airflow and temperature at these different locations at one or more future time points. In at least one embodiment, the predictions from these networks can be used by the corresponding boards, manifolds, PDUs or other such devices to make adjustments to ensure appropriate temperatures and flows and keep the equipment running normally. In at least one embodiment, these adjustments can include adjusting the amount of fluid or air flowing into or out of the server or equipment, and the temperature of the fluid or air. In at least one embodiment, this can also include adjusting the amount of electricity provided to one or more of these devices. In at least one embodiment, these can include similar adjustments that will be made in other data centers, but can be based on future predictions rather than observed states, so that such systems can be active rather than passive. In at least one embodiment, the monitoring can occur continuously or at least at regular intervals to ensure appropriate continuous operation.
[0067] In at least one embodiment, each neural network in each type of component can be trained specifically for that type of component and can receive updated network parameters that can be generated as a result of further training or continued learning. In at least one embodiment, inferences generated by a single device can also be shared with other devices in the data center, which can help make more accurate predictions. In at least one embodiment, each control board 306, 314 can be a board that supports neural networks, an ASIC, or other such component. In at least one embodiment, the control board can have components of a general-purpose computer, including at least one processor 1014 and a memory 1016 with associated circuitry, such as for Fig. 10A , as described in computing device 1002 in . In at least one embodiment, a Jetson board from NVIDIA can be used, which is a complete system on module (SPM), including CPU, GPU, PMIC, DRAM and flash memory. In at least one embodiment, such a module is also expandable to provide additional functions or capabilities. In at least one embodiment, each such board can be used as a small artificial intelligence (AI) or a computer supporting machine learning. In at least one embodiment, multiple such boards can be used in parallel, such as on the rack manifold 304, so as to process data from multiple high-resolution sensors at the same time. In at least one embodiment, these different neural networks can make predictions that can be shared across servers, racks or data centers, for example, in order to provide accurate predictions of different locations at one or more future time points. In at least one embodiment, appropriate adjustments can then be made appropriately to maintain temperature and other parameters at appropriate levels or values. In at least one embodiment, the liquid distribution control board 306 can make predictions, which can be used only to adjust the control for the relevant flow value, or these predictions can be shared so that other adjustments can also be made at least in part based on these predictions. In at least one embodiment, a control system for a data center can collect predictions from these different boards or networks to make adjustments that are more suitable for the entire data center. In at least one embodiment, adjustments can then be made at the server, rack, pod, data center, or other such level.
[0068] In at least one embodiment, a neural network-enabled method for liquid cooling racks at a data center level may utilize Figure 4AComponents of the system 400 shown in . In at least one embodiment, an external cooling unit 402 can provide liquid at a determined temperature to a data center cooling distribution unit 404. In at least one embodiment, the unit 404 can include a set of neural network-enabled distribution unit control boards 406 that can perform inferences and make adjustments based on future values of the inferences. In at least one embodiment, this can include adjusting the temperature of liquid to be received from the external cooling unit 402, or adjusting the flow of liquid into and / or out of the external unit 402. In at least one embodiment, this can also include adjusting the flow into a manifold 408 for a row of liquid-cooled racks 412, as well as the flow into and out of these racks, which can be determined at least in part based on inferences from a neural network-enabled control board 410 built into the row manifold 410. In at least one embodiment, there can be different levels of flow into and out of different liquid-cooled racks. In at least one embodiment, there can also be different flows into individual servers in a rack, such as with respect to servers in the rack. Figure 3 As described. In at least one embodiment, a similar approach can be used for air-cooled data centers. In at least one embodiment, temperature and airflow can be measured and predicted for future values at, for example, 30 seconds, 1 minute, 5 minutes, or 1 hour in the future, which can then adjust air temperature, airflow, power distribution, and other such aspects. In at least one embodiment, various types of input data can be used to make these predictions, such as server load, current temperature, current pressure, flow rate, power consumption, and other relevant data for various locations in the data center.
[0069] In at least one embodiment, these neural network-enabled boards can control the amount, supply, and return of liquid flows at different locations, such as racks, servers, or GPU levels. In at least one embodiment, these boards can be used to predict, control, and create optimal solutions for such data centers to maintain desired operating conditions without causing equipment loss or downtime due to temperature issues or thermal events. In at least one embodiment, this can be based on current load and expected load, as well as measured or detected network conditions. In at least one embodiment, this approach can also be applied to other types of equipment, such as UPS power supplies. In at least one embodiment, any device with a logic board or component can be injected with this small neural network code for these and other such purposes. In at least one embodiment, such a network can access all devices available in a unit or device, so that in a device such as a PDU, the board can regulate or completely cut off power to a server node, such as where a leak is detected or predicted and the relevant equipment may be damaged if it continues to operate. In at least one embodiment, the operation can be terminated normally or the power can be shut down to avoid damage. In at least one embodiment, the execution of the program can be terminated alternatively to at least avoid data loss.
[0070] In at least one embodiment, the cooling system can utilize a direct return design, such as Figure 4A As shown. In at least one embodiment, the coolant from the CDU 404 is directed to the individual server racks along the supply manifold 408. In at least one embodiment, the first server rack (A) along the flow will receive the liquid first, and the last server rack (E) along the flow will receive the liquid last. In at least one embodiment, the return manifold 410 will be used to receive the heated liquid passing through these racks and return the heated liquid to the CDU 404. In at least one embodiment, the heated liquid returned from the first server rack (A) will have the shortest return path, while the last server rack (E) has the longest return path. In at least one embodiment, this results in the total liquid path 414 of the first server rack (A) being much shorter than the total liquid path 416 of the last server rack (E) 416. In at least one embodiment, there may be significant changes in aspects such as pressure and flow throughout the circuit of the cooling system. In at least one embodiment, this can include lower pressures near the end of the cooling circuit and higher pressures in the middle of the circuit. In at least one embodiment, such an unbalanced design can result in unacceptable variations in flow rates throughout a data center or other environment because the different path lengths will result in different pressures, temperatures, and other conditions that affect the flows. In at least one embodiment, the different amounts of time required for the liquid to travel along these different paths can also result in flow variations.
[0071] In at least one embodiment, a reverse return based design may be utilized, such as Figure 4B In at least one embodiment, the system includes a reverse return manifold 422 that receives heated liquid first from the first server rack (A) and last from the last server rack (E), but rather than directing the flow directly back to the CDU 406, the liquid flows from the first server to the last server, or in a reverse return manifold 422. Figure 4B 404. As shown, the liquid for each server rack will follow a path 424 of similar length, at least through the corresponding manifold. In at least one embodiment, these path lengths can be effectively equal, allowing for slight variations based on factors such as placement and component type. In at least one embodiment, the liquid for each server rack will flow through substantially the entire length of the supply manifold and the return manifold, rather than as Figure 4AIn at least one embodiment, this causes the liquid for all server racks to experience these same changes in temperature, pressure, and other environmental parameters encountered throughout the operation of the manifold. In at least one embodiment, this balanced design can help reduce the flow rate variation of liquid delivered to different locations throughout the cooling system. In at least one embodiment, this can help ensure that the flow rate variation remains within a desired range, or less than a maximum variation threshold. In at least one embodiment, such a balanced design can help ensure that equal and sufficient liquid is delivered to all servers and racks in a data center through the various cooling components of the liquid cooling system.
[0072] In at least one embodiment, in various data center collectors, row manifolds, and rack manifolds, all supply fluid lines and return fluid lines between one or more CDUs and one or more direct-to-chip (D2C) cold plate cooling loops or heat exchangers in immersed liquid cooling blades can utilize reverse returns rather than direct returns. In at least one embodiment, the electronic valve control device at each outlet of these collectors, row manifolds, and rack manifolds can have the ability to both monitor and control the flow of primary and / or auxiliary fluids in and out of these CDUs, which flow to cold plates attached to components such as CPUs, GPUs, and switches, as well as other D2C liquid cooling components and blade immersed servers. In at least one embodiment, the liquid cooling system can utilize a combination of self-balancing flow design in all primary and auxiliary loops with electronic monitoring and control utilizing AI or machine learning, based at least in part on data received from various flow rate, pressure, temperature, power, and other environmental or operational sensors. In at least one embodiment, this approach can provide improved or optimized liquid distribution of liquid cooling components, such as servers and processors, in an environment such as a data center.
[0073] In at least one embodiment, neural network-based flow control can be used with this design to provide further variation minimization or flow consistency. In at least one embodiment, the combination of self-balancing with electronic monitoring and control can provide a very balanced and optimized coolant flow for each of the servers, racks, processors, or other such components without these components conflicting with each other and without requiring extensive control schemes. In at least one embodiment, flow controllers can be located at different locations throughout the cooling system, such as can be used to direct coolant into and out of components such as racks, servers, and electronic components, where the liquid can be provided using paths such as row manifolds and rack manifolds. In at least one embodiment, these flow controllers can allow incremental control (e.g., partially open and closed) to provide an accurate, self-balancing method. In at least one embodiment, a self-balancing design with finely adjustable control can provide accurate balancing of liquid flow at various outlets (if not all) of the liquid cooling system. In at least one embodiment, this approach can be used to provide a flow rate of, for example, 50 liters per minute (lpm) with a tolerance, threshold, or variation range of + / - 5 lpm, whereas a design without this approach might have a flow rate of approximately 60 lpm at one location and approximately 40 lpm at another location due to differences such as pressure gradients and gravity effects, which could result in an unacceptable flow imbalance across the system.
[0074] In at least one embodiment, the control system 440 can be as follows Figure 4CImplementation shown. In at least one embodiment, a reverse return row manifold design can be used to provide a balance of pressure and flow in the system. In at least one embodiment, a controllable flow valve 446 can be used for each inlet and / or outlet location, such as the inlet and outlet from the row manifold to each rack. In at least one embodiment, the flow manager 442 can receive data from various sensors, devices, or data sources in the relevant environment. In at least one embodiment, the flow manager 442 can provide at least some of the data as input to one or more neural networks (e.g., deep neural networks (DNNs)) of the reasoning module 444. In at least one embodiment, sensors and valves can be calibrated for a type of fluid used. In at least one embodiment, there can be one flow manager 442 for the entire data center or environment, or there can be flow managers for individual rows, racks, or servers. In at least one embodiment, a neural network can also be used at each individual flow control valve because it can communicate with one or more individual flow controllers. In at least one embodiment, the flow manager 442 can receive one or more inferences from the reasoning module 444, wherein these inferences can include adjustments or settings for individual flow control valves. In at least one embodiment, the flow manager 442 may then send this or related data or instructions to the individual flow control valves 446 or other such components to make small adjustments to help balance the flow in the system.
[0075] In at least one embodiment, this approach can also be used at the rack level or server level, such as Figure 4D In at least one embodiment, the liquid-cooled rack may have a rack manifold with a reverse return design so that the liquid path lengths of each server in the rack are also equal. In at least one embodiment, the size of the manifold may be much smaller than the size of the corresponding rack, such as Figure 4D The exploded view is shown for clarity of explanation. In at least one embodiment, the flow manager 462 can utilize the reasoning module 464 to determine the adjustments to be made to the individual flow control valves 466 within the rack to balance the flow between the individual servers in the rack. In at least one embodiment, the flow manager can also provide such instructions to the flow control valves or components within the server, such as can be used to direct coolant to a specific cold plate, processor, or other internal server component. In at least one embodiment, the use of these intelligent control valves can provide a well-balanced pipeline distribution and controllability. In at least one embodiment, each of these intelligent control valves 466 can fine-tune itself appropriately, which can be done under the guidance of the flow manager. In at least one embodiment, these flow valves can utilize PID or PIV servo control, where PIV allows control based on position and velocity errors.
[0076] In at least one embodiment, one or more neural networks may be trained to reason about the amount of liquid cooling required for various components in various situations. In at least one embodiment, this may include determining how much additional cooling may be required for a GPU that has gone from a more normal 300 watts of operating power to 500 watts, which may require 1.7 lpm instead of 1.0 lpm. In at least one embodiment, the neural network may reason about the adjustments to be made to the liquid flow to the component, and may be able to proactively make accurate adjustments to avoid the significant temperature fluctuations that result. In at least one embodiment, there may be one or more operating tables that may relate factors such as proportional flow rate to power consumption, which may be used to train these models or to determine adjustments based on inferred state data. In at least one embodiment, these tables may also be used for different types of components, or different models of the same type of components, to account for differences inherent in the components themselves. In at least one embodiment, reasoning may be performed at different levels of granularity, from the overall data center level to the internal server component level such as a GPU. In at least one embodiment, data may be collected by a component such as a board management controller (BMC) or server and provided to a traffic manager via a network port, which may send instructions to one or more traffic controllers to make one or more appropriate adjustments.
[0077] In at least one embodiment, a neural network can predict the state of fluid at different locations at one or more future times based at least in part on available information. In at least one embodiment, this can include a snapshot of current information, or can include at least some data from the recent past. In at least one embodiment, these predictions can then be used to adjust aspects such as flow rate. In at least one embodiment, a cooling distribution unit can adjust liquid flow or coolant temperature. In at least one embodiment, such a device can employ these predictions and be proactive in order to prevent any undesirable events before they occur. In at least one embodiment, a variety of sensors can provide information about the current state of the computing environment. In at least one embodiment, this can include the use of sensors such as temperature sensors, load sensors, flow sensors, or pressure sensors, which can be collected instantaneously or historically. In at least one embodiment, predictions made based on data from these sensors can be compared to one or more thresholds, ranges, or other operating criteria to determine whether any changes should be made. In at least one embodiment, this can include making adjustments to prevent unacceptable temperature increases at specific locations within a data center or other such environment. In at least one embodiment, this can include closing valves or increasing coolant flow, adjusting coolant temperature, or sounding alarms, and other such remedial measures.
[0078] In at least one embodiment, the Figure 5A 500 for adjusting one or more environmental control components. In at least one embodiment, sensor data is captured 502 for one or more locations in a computing environment, which may include pressure, temperature, and flow rate data, for example. In at least one embodiment, this may include data captured by sensors in racks, servers, cooling units, or other such devices. In at least one embodiment, at least some of the sensor data may be provided 504 as input to one or more neural networks, such as may be utilized by one or more flow manager modules within or associated with one or more devices or systems in the environment. In at least one embodiment, valve settings may be received 506 from one or more neural networks to be applied to one or more flow control valves to keep the liquid flow in the cooling system within a maximum variation in the environment. In at least one embodiment, flow data may be received from one or more neural networks instead, and adjustments may be determined based at least in part on the inferred flow data. In at least one embodiment, one or more adjustments may be made to one or more flow control valves based on these inferred or determined valve settings (508). In at least one embodiment, the process may continue to maintain a consistent and balanced flow throughout the liquid cooling system.
[0079] In at least one embodiment, the Figure 5B 5. A process 550 for regulating liquid flow in a cooling system is shown. In at least one embodiment, inputs related to environmental parameters of one or more components (e.g., processors, servers, or racks) within a data center are received 552. In at least one embodiment, one or more neural networks can be used to infer one or more adjustments to be made to maintain consistency of liquid flow in the data center. In at least one embodiment, these one or more adjustments can then be performed 556 for one or more control valves or other such components of the data center.
[0080] Servers and Data Centers
[0081] The following figures illustrate, but are not limited to, exemplary network server and data center based systems that may be used to implement at least one embodiment.
[0082] Figure 6A distributed system 600 is shown in accordance with at least one embodiment. In at least one embodiment, the distributed system 600 includes one or more client computing devices 602, 604, 606, and 608 configured to execute and operate client applications, such as network (web) browsers, proprietary clients, and / or variations thereof, over one or more networks 610. In at least one embodiment, a server 612 may be communicatively coupled to remote client computing devices 602, 604, 606, and 608 via the network 610.
[0083] In at least one embodiment, the server 612 may be adapted to run one or more services or software applications, such as services and applications that can manage session activities for single sign-on (SSO) access across multiple data centers. In at least one embodiment, the server 612 may also provide other services, or software applications, which may include non-virtualized and virtualized environments. In at least one embodiment, these services may be provided to users of client computing devices 602, 604, 606, and / or 608 as web-based services or cloud services or under a software as a service (SaaS) model. In at least one embodiment, users operating client computing devices 602, 604, 606, and / or 608 may in turn utilize one or more client applications to interact with the server 612 to utilize the services provided by these components.
[0084] In at least one embodiment, software components 618, 620, and 622 of system 600 are implemented on server 612. In at least one embodiment, one or more components of system 600 and / or the services provided by these components may also be implemented by one or more of client computing devices 602, 604, 606, and / or 608. In at least one embodiment, a user operating a client computing device may then utilize one or more client applications to use the services provided by these components. In at least one embodiment, these components may be implemented in hardware, firmware, software, or a combination thereof. It should be understood that a variety of different system configurations are possible, which may differ from distributed system 600. Therefore, Figure 6 The illustrated embodiment is at least one embodiment of a distributed system for implementing an embodiment system and is not intended to be limiting.
[0085] In at least one embodiment, client computing devices 602, 604, 606, and / or 608 may include different types of computing systems. In at least one embodiment, client computing devices may include portable handheld devices (e.g., Cellular phone, computing tablets, personal digital assistants (PDAs), or wearable devices (e.g., Google head mounted display), running software (such as Microsoft Windows ) and / or various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS and / or their variants). In at least one embodiment, the device can support different applications, such as different Internet-related applications, email, short message service (SMS) applications, and can use various other communication protocols. In at least one embodiment, the client computing device can also include a general-purpose personal computer, which in at least one embodiment includes running various versions of Microsoft Apple and / or a personal computer and / or laptop computer running Linux operating system.
[0086] In at least one embodiment, the client computing device may be a computer running various commercially available or any workstation computer running any of the UNIX-like operating systems, including but not limited to various GNU / Linux operating systems, such as Google Chrome OS. In at least one embodiment, client computing devices may also include electronic devices capable of communicating over one or more networks 610, such as thin client computers, Internet-enabled gaming systems (e.g., with or without gesture input devices for Microsoft's Xbox gaming console), and / or personal messaging devices. Figure 6 The distributed system 600 in FIG. 6 is shown as having four client computing devices, but any number of client computing devices may be supported. Other devices (such as devices with sensors, etc.) may interact with the server 612 .
[0087] In at least one embodiment, the network 610 in the distributed system 600 can be any type of network capable of supporting data communications using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internetwork Packet Exchange), AppleTalk, and / or variants thereof. In at least one embodiment, the network 610 can be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, and / or any other wireless protocols), and / or any combination of these and / or other networks.
[0088] In at least one embodiment, server 612 may be comprised of one or more general purpose computers, dedicated server computers (including, in at least one embodiment, PC (personal computer) servers, In at least one embodiment, the server 612 may include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization. In at least one embodiment, one or more flexible logical storage device pools may be virtualized to maintain virtual storage devices for the server. In at least one embodiment, the virtual network may be controlled by the server 612 using software defined networking. In at least one embodiment, the server 612 may be suitable for running one or more services or software applications.
[0089] In at least one embodiment, the server 612 can run any operating system, and any commercially available server operating system. In at least one embodiment, the server 612 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Server, database server and / or variants thereof. In at least one embodiment, exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variants thereof.
[0090] In at least one embodiment, server 612 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client computing devices 602, 604, 606, and 608. In at least one embodiment, data feeds and / or event updates may include, but are not limited to, data received from one or more third-party information sources and continuous data streams. feed, Updates or real-time updates, which may include real-time events related to sensor data applications, financial quoters, network performance measurement tools (e.g., network monitoring and business management applications), clickstream analysis tools, automobile traffic monitoring, and / or changes thereof. In at least one embodiment, the server 612 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of the client computing devices 602, 604, 606, and 608.
[0091] In at least one embodiment, the distributed system 600 may also include one or more databases 614 and 616. In at least one embodiment, the database may provide a mechanism for storing information such as user interaction information, usage pattern information, adaptation rule information, and other information. In at least one embodiment, the databases 614 and 616 may reside in various locations. In at least one embodiment, one or more of the databases 614 and 616 may reside on a non-transient storage medium local to the server 612 (and / or residing in the server 612). In at least one embodiment, the databases 614 and 616 may be remote from the server 612 and communicate with the server 612 via a network-based connection or a dedicated connection. In at least one embodiment, the databases 614 and 616 may reside in a storage area network (SAN). In at least one embodiment, any necessary files for performing functions attributed to the server 612 may be appropriately stored locally on the server 612 and / or remotely stored. In at least one embodiment, the databases 614 and 616 may include a relational database, such as a database suitable for storing, updating, and retrieving data in response to SQL formatted commands.
[0092] Figure 7 An exemplary data center 700 is shown in accordance with at least one embodiment. In at least one embodiment, data center 700 includes, but is not limited to, a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0093] In at least one embodiment, Figure 7 As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources ("node CRs") 716(1)-716(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 716(1)-716(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays ("FPGAs"), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 716(1)-716(N) may be a server having one or more of the above-mentioned computing resources.
[0094] In at least one embodiment, the grouped computing resources 714 may include a separate grouping (not shown) of node CRs housed in one or more racks, or many racks (also not shown) housed in data centers at various geographic locations. The separate grouping of node CRs within the grouped computing resources 714 may include computing, networks, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0095] In at least one embodiment, resource coordinator 712 may configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may include a software design infrastructure ("SDI") management entity for data center 700. In at least one embodiment, resource coordinator 712 may include hardware, software, or some combination thereof.
[0096] In at least one embodiment, Figure 7As shown, the framework layer 720 includes, but is not limited to, a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, the framework layer 720 may include a framework that supports software 752 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 752 or the application 742 may include a web-based service software or application, such as a service or application provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be, but is not limited to, a free and open source software network application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and a distributed file system 738 for supporting large-scale data processing. In at least one embodiment, the resource manager 736 can manage clustered or grouped computing resources mapped to or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 736 can coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0097] In at least one embodiment, the software 752 included in the software layer 730 may include software used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0098] In at least one embodiment, the one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, CUDA applications, 5G network applications, artificial intelligence applications, data center applications, and / or variations thereof.
[0099] In at least one embodiment, any of the configuration manager 734, resource manager 736, and resource coordinator 712 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of the data center 700 from making potentially bad configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0100] Figure 8 A client-server network 804 formed by a plurality of interconnected network server computers 802 is shown according to at least one embodiment. In at least one embodiment, each network server computer 802 stores data accessible to other network server computers 802 and client computers 806 and networks 808 linked to the wide area network 804. In at least one embodiment, the configuration of the client-server network 804 can change over time as client computers 806 and one or more networks 808 are connected and disconnected from the network 804, and as one or more trunk server computers 802 are added to or removed from the network 804. In at least one embodiment, the client-server network includes client computers 806 and networks 808 when such client computers 806 and networks 808 are connected to the network server computers 802. In at least one embodiment, the term computer includes any device or machine capable of accepting data, applying a prescribed process to the data, and providing a result of the process.
[0101] In at least one embodiment, the client-server network 804 stores information accessible to the network server computer 802, the remote network 808, and the client computer 806. In at least one embodiment, the network server computer 802 is formed by a mainframe computer, a minicomputer, and / or a microcomputer each having one or more processors. In at least one embodiment, the server computer 802 is linked together by wired and / or wireless transmission media (such as wires, optical fiber cables) and / or microwave transmission media, satellite transmission media, or other conductive, optical or electromagnetic wave transmission media. In at least one embodiment, the client computer 806 accesses the network server computer 802 through similar wired or wireless transmission media. In at least one embodiment, the client computer 806 can be linked to the client-server network 804 using a modem and a standard telephone communication network. In at least one embodiment, alternative operator systems (such as cable and satellite communication systems) can also be used to link to the client-server network 804. In at least one embodiment, other private or time-shared operator systems can be used. In at least one embodiment, the network 804 is a global information network, such as the Internet. In at least one embodiment, the network is a private intranet that uses similar protocols to the Internet but with added security measures and restricted access controls. In at least one embodiment, the network 804 is a private or semi-private network that uses a proprietary communication protocol.
[0102] In at least one embodiment, the client computer 806 is any end-user computer, and may also be a mainframe computer, a minicomputer, or a microcomputer having one or more microprocessors. In at least one embodiment, the server computer 802 may sometimes be used as a client computer to access another server computer 802. In at least one embodiment, the remote network 808 may be a local area network, a network added to a wide area network through an independent service provider (ISP) for the Internet, or another group of computers interconnected by a wired or wireless transmission medium with a fixed or time-varying configuration. In at least one embodiment, the client computer 806 may be linked to the network 804 independently or through the remote network 808 and access the network 804.
[0103] Fig. 9A computer network 908 is shown that connects one or more computing machines according to at least one embodiment. In at least one embodiment, the network 908 can be any type of electrically connected computer group, including, for example, the following network: the Internet, an intranet, a local area network (LAN), a wide area network (WAN), or an interconnection combination of these network types. In at least one embodiment, the connection within the network 908 can be a remote modem, Ethernet (IEEE 802.3), a token ring (IEEE802.5), a fiber distributed data link interface (FDDI), an asynchronous transfer mode (ATM), or any other communication protocol. In at least one embodiment, the computing device linked to the network can be a desktop, a server, a portable, a handheld, a set-top box, a personal digital assistant (PDA), a terminal, or any other desired type or configuration. In at least one embodiment, depending on their functionality, the network-connected device can be widely varied in terms of processing power, internal memory, and other performances.
[0104] In at least one embodiment, the communication within the network and the communication to or from the computing device connected to the network can be wired or wireless. In at least one embodiment, the network 908 can include at least in part the world-wide public Internet, which is usually connected to multiple users according to the transmission control protocol / Internet protocol (TCP / IP) specification according to the client-server model. In at least one embodiment, the client-server network is the dominant model for communication between two computers. In at least one embodiment, the client computer ("client") issues one or more commands to the server computer ("server"). In at least one embodiment, the server fulfills the client command by accessing available network resources and returning information to the client according to the client command. In at least one embodiment, the client computer system and the network resources residing on the network server are assigned network addresses for identification during communication between the elements of the network. In at least one embodiment, the communication from other network-connected systems to the server will include the network address of the relevant server / network resource as part of the communication, so that the appropriate destination of the data / request is identified as the recipient. In at least one embodiment, when the network 908 includes the global Internet, the network address is an IP address in TCP / IP format, which can at least partially route data to an email account, website, or other Internet tool residing on the server. In at least one embodiment, information and services residing on the network server can be available to a web browser of a client computer through a domain name (e.g., www.site.com) that maps to the IP address of the network server.
[0105] In at least one embodiment, a plurality of clients 902, 904, and 906 are connected to a network 908 via respective communication links. In at least one embodiment, each of these clients may access the network 908 via any desired form of communication, such as via a dial-up modem connection, a cable link, a digital subscriber line (DSL), a wireless or satellite link, or any other form of communication. In at least one embodiment, each client may communicate using any machine (e.g., a personal computer (PC), a workstation, a dedicated terminal, a personal data assistant (PDA), or other similar device) that is compatible with the network 908. In at least one embodiment, the clients 902, 904, and 906 may or may not be located in the same geographic area.
[0106] In at least one embodiment, multiple servers 910, 912, and 914 are connected to a network 918 to serve clients communicating with the network 918. In at least one embodiment, each server is typically a powerful computer or device that manages network resources and responds to client commands. In at least one embodiment, the server includes a computer-readable data storage medium that stores program instructions and data, such as a hard drive and a RAM memory. In at least one embodiment, servers 910, 912, and 914 run applications that respond to client commands. In at least one embodiment, server 910 can run a web server application that responds to client requests for HTML pages, and can also run a mail server application that receives and routes emails. In at least one embodiment, other applications can also be run on server 910, such as an FTP server or a media server for streaming audio / video data to a client. In at least one embodiment, different servers can be dedicated to performing different tasks. In at least one embodiment, server 910 can be a dedicated web server that manages resources related to a website for different users, and server 912 can be dedicated to providing email (email) management. In at least one embodiment, other servers may be dedicated to media (audio, video, etc.), file transfer protocol (FTP), or a combination of any two or more services typically available or provided over a network. In at least one embodiment, each server may be in a location that is the same or different from that of the other servers. In at least one embodiment, there may be multiple servers that perform mirroring tasks for users, thereby alleviating congestion or minimizing traffic directed to and from a single server. In at least one embodiment, servers 910, 912, 914 are under the control of a web hosting provider in the business of maintaining and delivering third-party content over network 918.
[0107] In at least one embodiment, a web hosting provider delivers services to two different types of clients. In at least one embodiment, one type, which may be referred to as a browser, requests content from a server 910, 912, 914, such as a web page, an email message, a video clip, etc. In at least one embodiment, a second type, which may be referred to as a user, hires a web hosting provider to maintain network resources, such as a website, and make them available to the browser. In at least one embodiment, the user contracts with the web hosting provider to make available memory space, processor capacity, and communication bandwidth for the network resources they desire, depending on the amount of server resources the user desires to utilize.
[0108] In at least one embodiment, in order for the web hosting provider to serve both clients, an application that manages the network resources hosted by the server must be appropriately configured. In at least one embodiment, the program configuration process involves defining a set of parameters that at least partially controls the application's response to browser requests and also at least partially defines the server resources available to a particular user.
[0109] In one embodiment, the intranet server 916 communicates with the network 908 via a communication link. In at least one embodiment, the intranet server 916 communicates with a server manager 918. In at least one embodiment, the server manager 918 includes a database of application configuration parameters used in the servers 910, 912, 914. In at least one embodiment, a user modifies the database 920 via the intranet 916, and the server manager 918 interacts with the servers 910, 912, 914 to modify the application parameters so that they match the contents of the database. In at least one embodiment, a user logs into the intranet 916 by connecting to the intranet 916 via the computer 902 and entering authentication information such as a username and password.
[0110] In at least one embodiment, when a user wishes to log in to a new service or modify an existing service, the intranet server 916 authenticates the user and provides the user with an interactive screen display / control panel that allows the user to access the configuration parameters of a specific application. In at least one embodiment, a plurality of modifiable text boxes describing aspects of the configuration of the user's website or other network resources are presented to the user. In at least one embodiment, if the user desires to increase the memory space reserved for his website on the server, a field in which the user specifies the desired memory space is provided to the user. In at least one embodiment, in response to receiving the information, the intranet server 916 updates the database 920. In at least one embodiment, the server manager 918 forwards the information to the appropriate server and uses the new parameters during the operation of the application. In at least one embodiment, the intranet server 916 is configured to provide the user with access to the configuration parameters of the hosted network resources (e.g., web pages, emails, FTP sites, media sites, etc.) that the user has signed with a web hosting service provider.
[0111] Fig. 10A A networked computer system 1000A according to at least one embodiment is shown. In at least one embodiment, the networked computer system 1000A includes a plurality of nodes or personal computers ("PCs") 1002, 1018, 1020. In at least one embodiment, the personal computer or node 1002 includes a processor 1014, a memory 1016, a camera 1004, a microphone 1006, a mouse 1008, a speaker 1010, and a monitor 1012. In at least one embodiment, the PCs 1002, 1018, 1020 can each run one or more desktop servers of an internal network within a given company, for example, or can be servers of a general network that is not limited to a specific environment. In at least one embodiment, each PC node of the network has a server, so that each PC node of the network represents a specific network server with a specific network URL address. In at least one embodiment, each server defaults to a default web page for a user of the server, which itself can contain embedded URLs pointing to further subpages of the user on the server, or to other servers on the network or pages on other servers.
[0112] In at least one embodiment, nodes 1002, 1018, 1020 and other nodes of the network are interconnected via medium 1022. In at least one embodiment, medium 1022 can be a communication channel such as an integrated services digital network ("ISDN"). In at least one embodiment, the various nodes of the networked computer system can be connected through various communication media, including a local area network ("LAN"), a plain old telephone line ("POTS") (sometimes referred to as a public switched telephone network ("PSTN")), and / or variants thereof. In at least one embodiment, the various nodes of the network can also constitute computer system users interconnected via a network such as the Internet. In at least one embodiment, each server on the network (running from a specific node of the network at a given instance) has a unique address or identification within the network, which can be specified according to a URL.
[0113] In at least one embodiment, multiple multipoint conferencing units ("MCUs") may thus be used to transmit data to and from various nodes or "endpoints" of a conferencing system. In at least one embodiment, the nodes and / or MCUs may be interconnected via ISDN links or through a local area network ("LAN"), in addition to various other communication media (such as, nodes connected via the Internet). In at least one embodiment, the nodes of a conferencing system may generally be connected directly to a communication medium (such as a LAN) or through an MCU, and the conferencing system may include other nodes or elements, such as routers, servers, and / or variations thereof.
[0114] In at least one embodiment, processor 1014 is a general purpose programmable processor. In at least one embodiment, the processor of a node of networked computer system 1000A may also be a dedicated video processor. In at least one embodiment, different peripherals and components of a node (such as those of node 1002) may be different from those of other nodes. In at least one embodiment, node 1018 and node 1020 may be configured to be the same or different from node 1002. In at least one embodiment, the node may be implemented on any suitable computer system other than a PC system.
[0115] Fig. 10BA networked computer system 1000B according to at least one embodiment is shown. In at least one embodiment, system 1000B shows a network (such as LAN 1024), which can be used to interconnect various nodes that can communicate with each other. In at least one embodiment, attached to LAN 1024 are multiple nodes, such as PC nodes 1026, 1028, 1030. In at least one embodiment, the node can also be connected to the LAN via a network server or other device. In at least one embodiment, system 1000B includes other types of nodes or elements, including routers, servers, and nodes for at least one embodiment.
[0116] Fig. 10C A networked computer system 1000C is shown according to at least one embodiment. In at least one embodiment, system 1000C shows a WWW system with communications across a backbone communication network (such as the Internet 1032) that can be used to interconnect various nodes of the network. In at least one embodiment, the WWW is a set of protocols that operate on top of the Internet and allow a graphical interface system to operate on it to access information through the Internet. In at least one embodiment, attached to the Internet 1032 in the WWW are multiple nodes, such as PCs 1040, 1042, 1044. In at least one embodiment, the nodes interface with other nodes of the WWW through WWW HTTP servers (such as servers 1034, 1036). In at least one embodiment, PC 1044 can be a PC that forms a node of the network 1032, and PC 1044 itself runs its server 1036, although for illustrative purposes, it is not shown in FIG. Fig. 10C PC 1044 and server 1036 are shown separately in FIG.
[0117] In at least one embodiment, the WWW is a distributed type of application characterized by WWW HTTP, the protocol of the WWW, which runs on top of the Internet's Transmission Control Protocol / Internet Protocol ("TCP / IP"). In at least one embodiment, the WWW can therefore be characterized by a set of protocols (i.e., HTTP) running on the Internet as its "backbone."
[0118] In at least one embodiment, a web browser is an application running on a node of the network in a WWW-type compatible network system that allows users of a particular server or node to view such information and, therefore, allows users to search for graphics and text-based files linked together using hypertext links embedded in documents or files available from servers on a network that understands HTTP. In at least one embodiment, when a user retrieves a given web page of a first server associated with a first node using another server on a network such as the Internet, the retrieved document may have different hypertext links embedded therein, and a local copy of the page is created locally on the retrieving user. In at least one embodiment, when a user clicks on a hypertext link, the locally stored information associated with the selected hypertext link is typically sufficient to allow the user's machine to open a connection through the Internet to the server indicated by the hypertext link.
[0119] In at least one embodiment, more than one user can be coupled to each HTTP server via a LAN (such as LAN 1038, such as shown with respect to WWW HTTP server 1034). In at least one embodiment, system 1000C may also include other types of nodes or elements. In at least one embodiment, the WWW HTTP server is an application running on a machine such as a PC. In at least one embodiment, each user may be considered to have a unique "server", as shown with respect to PC 1044. In at least one embodiment, a server may be considered to be a server such as WWW HTTP server 1034, which provides access to a network for a LAN or multiple nodes or multiple LANs. In at least one embodiment, there are multiple users, each user having a desktop PC or a node of a network, each desktop PC potentially establishing a server for its user. In at least one embodiment, each server is associated with a specific network address or URL, which, when accessed, provides a default web page for the user. In at least one embodiment, a web page may contain further links (embedded URLs) pointing to further subpages of the user on the server, or to other servers on the network or to pages on other servers on the network.
[0120] Cloud Computing and Services
[0121] The following figures illustrate, but are not limited to, exemplary cloud-based systems that may be used to implement at least one embodiment.
[0122] In at least one embodiment, cloud computing is a style of computing in which dynamically scalable and usually virtualized resources are provided as services over the Internet. In at least one embodiment, users do not need to have knowledge of the technical infrastructure that supports them, expertise in the technical infrastructure, or control over the technical infrastructure, which can be referred to as "in the cloud". In at least one embodiment, cloud computing merges infrastructure into services, platform as a service, software as a service, and other variations with common themes that rely on the Internet to meet the computing needs of users. In at least one embodiment, a typical cloud deployment (such as in a private cloud (e.g., an enterprise network)) or a data center (DC) in a public cloud (e.g., the Internet) may consist of thousands of servers (or alternatively, VMs), hundreds of Ethernet, Fiber Channel or Fiber Channel over Ethernet (FCoE) ports, switching and storage infrastructure, etc. In at least one embodiment, the cloud may also consist of a network service infrastructure, such as an IPsec VPN hub, a firewall, a load balancer, a wide area network (WAN) optimizer, etc. In at least one embodiment, remote subscribers can securely access cloud applications and services by connecting via a VPN tunnel (e.g., an IPsec VPN tunnel).
[0123] In at least one embodiment, cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be quickly provisioned and released with minimal management effort or service provider interaction.
[0124] In at least one embodiment, cloud computing is characterized by on-demand self-service, where consumers can automatically and unilaterally provision computing capabilities, such as server time and network storage, as needed, without human interaction with each service provider. In at least one embodiment, cloud computing is characterized by broad network access, where capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). In at least one embodiment, cloud computing is characterized by resource pooling, where a provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically signed and reallocated based on consumer demand. In at least one embodiment, there is a sense of location independence, as consumers typically have no control or knowledge of the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0125] In at least one embodiment, resources include storage, processing, memory, network bandwidth, and virtual machines. In at least one embodiment, cloud computing is characterized by rapid elasticity, where capacity can be quickly and elastically provisioned (automatically in some cases) to quickly scale down and quickly release to quickly scale up. In at least one embodiment, the capacity available for provisioning generally appears to the consumer to be unlimited and can be purchased in any quantity at any time. In at least one embodiment, cloud computing is characterized by measured services, where the cloud system automatically controls and optimizes resource usage by utilizing metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). In at least one embodiment, resource usage can be monitored, controlled, and reported, thereby providing transparency to both the provider and consumer of the utilized services.
[0126] In at least one embodiment, cloud computing may be associated with a variety of services. In at least one embodiment, cloud software as a service (SaaS) may refer to a service that provides consumers with the ability to use a provider's applications running on a cloud infrastructure. In at least one embodiment, the applications may be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure including networks, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0127] In at least one embodiment, cloud platform as a service (PaaS) may refer to a service where the capability provided to the consumer is to deploy consumer-created or acquired applications onto a cloud infrastructure, where the applications are created using programming languages and tools supported by the provider. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure including networks, servers, operating systems, or storage, but has control over the deployed applications and possibly the configuration of the application hosting environment.
[0128] In at least one embodiment, cloud infrastructure as a service (IaaS) may refer to a service in which the capabilities provided to the consumer are to provide processing, storage, networking, and other basic computing resources on which the consumer can deploy and run arbitrary software, which may include operating systems and applications. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).
[0129] In at least one embodiment, cloud computing can be deployed in different ways. In at least one embodiment, a private cloud may refer to a cloud infrastructure that is operated only for an organization. In at least one embodiment, a private cloud may be managed by an organization or a third party and may exist on or off-premises. In at least one embodiment, a community cloud may refer to a cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., missions, security requirements, policies, and compliance considerations). In at least one embodiment, a community cloud may be managed by an organization or a third party and may exist on or off-premises. In at least one embodiment, a public cloud may refer to a cloud infrastructure that is available to the general public or a large industry group and owned by an organization that provides cloud services. In at least one embodiment, a hybrid cloud may refer to a cloud infrastructure that is a component of two or more clouds (private, community, or public), which remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds). In at least one embodiment, a cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability.
[0130] Fig.11 One or more components of a system environment 1100 are shown, according to at least one embodiment, where services may be provided as third-party network services. In at least one embodiment, the third-party network may be referred to as a cloud, a cloud network, a cloud computing network, and / or variations thereof. In at least one embodiment, the system environment 1100 includes one or more client computing devices 1104, 1106, and 1108, which may be used by users to interact with a third-party network infrastructure system 1102 that provides third-party network services (which may be referred to as cloud computing services). In at least one embodiment, the third-party network infrastructure system 1102 may include one or more computers and / or servers.
[0131] It should be understood that Fig.11 The third-party network infrastructure system 1102 depicted in FIG. 1 may have other components in addition to those depicted. Further, Fig.11 An embodiment of a third-party network infrastructure system is depicted. In at least one embodiment, the third-party network infrastructure system 1102 may have Fig.11 More or fewer components may be depicted, two or more components may be combined, or there may be a different configuration or arrangement of components.
[0132] In at least one embodiment, the client computing devices 1104, 1106, and 1108 can be configured to operate a client application, such as a web browser, a proprietary client application, or some other application that can be used by a user of the client computing device to interact with the third-party network infrastructure system 1102 to use services provided by the third-party network infrastructure system 1102. Although the exemplary system environment 1100 is shown with three client computing devices, any number of client computing devices can be supported. In at least one embodiment, other devices, such as devices with sensors, etc., can interact with the third-party network infrastructure system 1102. In at least one embodiment, one or more networks 1110 can facilitate communication and data exchange between the client computing devices 1104, 1106, and 1108 and the third-party network infrastructure system 1102.
[0133] In at least one embodiment, the services provided by the third-party network infrastructure system 1102 may include a host of services that are available on-demand to users of the third-party network infrastructure system. In at least one embodiment, a variety of services may also be provided, including but not limited to online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database management and processing, managed technical support services, and / or variations thereof. In at least one embodiment, the services provided by the third-party network infrastructure system may be dynamically scalable to meet the needs of its users.
[0134] In at least one embodiment, a specific instantiation of a service provided by the third-party network infrastructure system 1102 may be referred to as a "service instance". In at least one embodiment, generally, any service available to a user from a third-party network service provider system via a communication network (such as the Internet) is referred to as a "third-party network service". In at least one embodiment, in a public third-party network environment, the servers and systems that make up the third-party network service provider system are different from the customer's own on-premises servers and systems. In at least one embodiment, the third-party network service provider system can host applications, and users can subscribe to and use the applications on demand via a communication network (such as the Internet).
[0135] In at least one embodiment, services in a computer network third-party network infrastructure may include protected computer network access to storage, hosted databases, hosted network servers, software applications, or other services provided to users by a third-party network provider. In at least one embodiment, services may include password-protected access to remote storage devices on a third-party network via the Internet. In at least one embodiment, services may include hosted relational databases and scripting language middleware engines based on web services for private use by networked developers. In at least one embodiment, services may include access to an email software application hosted on a third-party network provider's website.
[0136] In at least one embodiment, the third-party network infrastructure system 1102 may include a set of applications, middleware and database service providers delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available and secure manner. In at least one embodiment, the third-party network infrastructure system 1102 may also provide computing and analysis services related to "big data". In at least one embodiment, the term "big data" is generally used to refer to extremely large data sets that can be stored and manipulated by analysts and researchers in order to visualize a large amount of data, detect trends, and / or otherwise interact with data. In at least one embodiment, big data and related applications can be hosted and / or manipulated by the infrastructure system at many levels and at different scales. In at least one embodiment, tens, hundreds or thousands of processors linked in parallel can act on such data in order to present the data or simulate external forces on the data or its represented content. In at least one embodiment, these data sets may involve structured data (such as structured data organized in a database or otherwise according to a structured model) and / or unstructured data (e.g., emails, images, data blobs (binary large objects), web pages, complex event processing). In at least one embodiment, by leveraging the ability of an embodiment to focus more (or fewer) computing resources toward a target relatively quickly, third-party network infrastructure systems may be better available to perform tasks on large data sets based on demand from businesses, government agencies, research organizations, private individuals, groups of like-minded individuals or organizations, or other entities.
[0137] In at least one embodiment, the third-party network infrastructure system 1102 can be adapted to automatically provide, manage, and track customer subscriptions to services provided by the third-party network infrastructure system 1102. In at least one embodiment, the third-party network infrastructure system 1102 can provide third-party network services via different deployment models. In at least one embodiment, services can be provided under a public third-party network model, where the third-party network infrastructure system 1102 is owned by an organization that sells third-party network services and makes the services available to the general public or different industry enterprises. In at least one embodiment, services can be provided under a private third-party network model, in which the third-party network infrastructure system 1102 operates only for a single organization and can provide services to one or more entities within the organization. In at least one embodiment, third-party network services can also be provided under a community third-party network model, where the third-party network infrastructure system 1102 and the services provided by the third-party network infrastructure system 1102 are shared by several organizations in a related community. In at least one embodiment, third-party network services can also be provided under a hybrid third-party network model, which is a combination of two or more different models.
[0138] In at least one embodiment, the services provided by the third-party network infrastructure system 1102 may include one or more services provided under a software as a service (SaaS) category, a platform as a service (PaaS) category, an infrastructure as a service (IaaS) category, or other service categories including hybrid services. In at least one embodiment, a customer via a subscription order may subscribe to one or more services provided by the third-party network infrastructure system 1102. In at least one embodiment, the third-party network infrastructure system 1102 then performs processing to provide the services in the customer's subscription order.
[0139] In at least one embodiment, the services provided by the third-party network infrastructure system 1102 may include, but are not limited to, application services, platform services, and infrastructure services. In at least one embodiment, application services may be provided by the third-party network infrastructure system via a SaaS platform. In at least one embodiment, the SaaS platform may be configured to provide third-party network services belonging to the SaaS category. In at least one embodiment, the SaaS platform may provide the ability to build and deliver a set of on-demand applications on an integrated development and deployment platform. In at least one embodiment, the SaaS platform may manage and control the underlying software and infrastructure used to provide SaaS services. In at least one embodiment, by utilizing the services provided by the SaaS platform, customers may utilize applications executed on a third-party network infrastructure system. In at least one embodiment, customers may obtain application services without the need for customers to purchase separate licenses and support. In at least one embodiment, a variety of different SaaS services may be provided. In at least one embodiment, this may include, but is not limited to, services that provide solutions for sales performance management, enterprise integration, and business flexibility for large organizations.
[0140] In at least one embodiment, platform services may be provided by the third-party network infrastructure system 1102 via a PaaS platform. In at least one embodiment, the PaaS platform may be configured to provide third-party network services that fall into the PaaS category. In at least one embodiment, platform services may include, but are not limited to, services that enable organizations to merge existing applications on a shared common architecture, and the ability to build new applications that utilize shared services provided by the platform. In at least one embodiment, the PaaS platform may manage and control the underlying software and infrastructure used to provide PaaS services. In at least one embodiment, customers may obtain PaaS services provided by the third-party network infrastructure system 1102 without requiring the customer to purchase separate licenses and support.
[0141] In at least one embodiment, by utilizing the services provided by the PaaS platform, customers can adopt programming languages and tools supported by the third-party network infrastructure system and also control the deployed services. In at least one embodiment, the platform services provided by the third-party network infrastructure system may include database third-party network services, middleware third-party network services, and third-party network services. In at least one embodiment, the database third-party network service may support a shared service deployment model that enables organizations to aggregate database resources and provide database as a service to customers in the form of a database third-party network. In at least one embodiment, in a third-party network infrastructure system, the middleware third-party network service can provide customers with a platform to develop and deploy different business applications, and the third-party network service can provide customers with a platform to deploy applications.
[0142] In at least one embodiment, a variety of different infrastructure services can be provided by an IaaS platform in a third-party network infrastructure system. In at least one embodiment, the infrastructure services facilitate the management and control of underlying computing resources (such as storage, network, and other basic computing resources) by customers utilizing services provided by SaaS platforms and PaaS platforms.
[0143] In at least one embodiment, the third-party network infrastructure system 1102 may also include infrastructure resources 1130 for providing resources for providing various services to customers of the third-party network infrastructure system. In at least one embodiment, the infrastructure resources 1130 may include a pre-integrated and optimized combination of hardware (such as servers, storage, and networking resources) for performing services and other resources provided by the PaaS platform and the SaaS platform.
[0144] In at least one embodiment, resources in the third party network infrastructure system 1102 can be shared by multiple users and dynamically reallocated as needed. In at least one embodiment, resources can be allocated to users in different time zones. In at least one embodiment, the third party network infrastructure system 1102 can enable a first group of users in a first time zone to utilize the resources of the third party network infrastructure system for a specified number of hours, and then enable the same resources to be reallocated to another group of users in a different time zone, thereby maximizing resource utilization.
[0145] In at least one embodiment, a plurality of internal shared services 1132 shared by different components or modules of the third-party network infrastructure system 1102 may be provided to enable services provided by the third-party network infrastructure system 1102. In at least one embodiment, these internal shared services may include, but are not limited to, security and identity services, integration services, enterprise library services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services for enabling third-party network support, email services, notification services, file transfer services, and / or variations thereof.
[0146] In at least one embodiment, the third-party network infrastructure system 1102 can provide comprehensive management of third-party network services (e.g., SaaS, PaaS, and IaaS services) in the third-party network infrastructure system. In at least one embodiment, the third-party network management functionality can include capabilities for provisioning, managing, and tracking customer subscriptions received by the third-party network infrastructure system 1102 and / or variations thereof.
[0147] In at least one embodiment, Fig.11As shown, third-party network management functionality may be provided by one or more modules, such as an order management module 1120, an order coordination module 1122, an order provisioning module 1124, an order management and monitoring module 1126, and an identity management module 1128. In at least one embodiment, these modules may include or be provided using one or more computers and / or servers, which may be general purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0148] In at least one embodiment, at step 1134, a customer using a client device (such as client computing device 1104, 1106, or 1108) may interact with third-party network infrastructure system 1102 by requesting one or more services provided by third-party network infrastructure system 1102 and placing an order for a subscription to one or more services provided by third-party network infrastructure system 1102. In at least one embodiment, the customer may access a third-party network user interface (UI), such as third-party network UI 1112, third-party network UI 1114, and / or third-party network UI 1116, and place an order via these UIs. In at least one embodiment, the order information received by third-party network infrastructure system 1102 in response to the customer placing an order may include information identifying the customer and one or more services provided by third-party network infrastructure system 1102 that the customer wants to subscribe to.
[0149] In at least one embodiment, at step 1136, the order information received from the customer can be stored in order database 1118. In at least one embodiment, if this is a new order, a new record can be created for the order. In at least one embodiment, order database 1118 can be one of several databases operated by third party network infrastructure system 1118 and in conjunction with other system elements.
[0150] In at least one embodiment, at step 1138, the order information may be forwarded to order management module 1120, which may be configured to perform billing and accounting functions associated with the order, such as validating the order and, upon validation, booking an order.
[0151] In at least one embodiment, at step 1140, information about the order may be transmitted to an order coordination module 1122, which is configured to coordinate the provisioning of services and resources for the order placed by the customer. In at least one embodiment, the order coordination module 1122 may use the services of the order provisioning module 1124 for provisioning. In at least one embodiment, the order coordination module 1122 enables management of the business processes associated with each order and applies business logic to determine whether the order should continue to be fulfilled.
[0152] In at least one embodiment, at step 1142, when an order for a new subscription is received, the order coordination module 1122 sends a request to the order provisioning module 1124 to allocate resources and configure the resources required to satisfy the subscription order. In at least one embodiment, the order provisioning module 1124 implements resource allocation for the services ordered by the customer. In at least one embodiment, the order provisioning module 1124 provides a level of abstraction between the third-party network services provided by the third-party network infrastructure system 1100 and the physical implementation layer for provisioning resources for providing the requested services. In at least one embodiment, this enables the order coordination module 1122 to be isolated from implementation details, such as whether services and resources are actually provisioned in real time, or are pre-provisioned and allocated / assigned only upon request.
[0153] In at least one embodiment, once the services and resources are provisioned, a notification may be sent to the subscribing client indicating that the requested service is now ready for use, step 1144. In at least one embodiment, information (e.g., a link) may be sent to the client enabling the client to begin using the requested service.
[0154] In at least one embodiment, at step 1146, the orders subscribed by the customer may be managed and tracked by the order management and monitoring module 1126. In at least one embodiment, the order management and monitoring module 1126 may be configured to collect usage statistics regarding the customer's use of the subscription service. In at least one embodiment, statistics may be collected for the amount of storage used, the amount of data transferred, the number of users, and the amount and / or changes in system power-up time and system power-down time.
[0155] In at least one embodiment, the third-party network infrastructure system 1100 may include an identity management module 1128 configured to provide identity services, such as access management and authorization services in the third-party network infrastructure system 1100. In at least one embodiment, the identity management module 1128 may control information about customers who wish to utilize services provided by the third-party network infrastructure system 1102. In at least one embodiment, such information may include information authenticating the identities of such customers and information describing which actions those customers are authorized to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). In at least one embodiment, the identity management module 1128 may also include managing descriptive information about each customer and information about how and by whom the descriptive information may be accessed and modified.
[0156] Fig.12 A cloud computing environment 1202 is shown in accordance with at least one embodiment. In at least one embodiment, the cloud computing environment 1202 includes one or more computer systems / servers 1204 with which computing devices such as personal digital assistants (PDAs) or cell phones 1206A, desktop computers 1206B, laptop computers 1206C, and / or automobile computer systems 1206N communicate. In at least one embodiment, this allows infrastructure, platforms, and / or software to be provided as a service from the cloud computing environment 1202 so that each client does not need to maintain such resources individually. It should be understood that Fig.12 The types of computing devices 1206A-N shown in are intended to be illustrative only, and the cloud computing environment 1202 may communicate with any type of computerized device over any type of network and / or network / addressable connection (eg, using a web browser).
[0157] In at least one embodiment, computer system / server 1204, which may be represented as a cloud computing node, may operate with numerous other general purpose or special purpose computing system environments or configurations. In at least one embodiment, computing systems, environments, and / or configurations that may be suitable for use with computer system / server 1204 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices, and / or variations thereof.
[0158] In at least one embodiment, computer system / server 1204 can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. In at least one embodiment, program modules include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. In at least one embodiment, exemplary computer system / server 1204 can be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In at least one embodiment, in a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.
[0159] Fig.13 The cloud computing environment 1202 ( Fig.12 ) provides a set of functional abstraction layers. It should be understood in advance that Fig.13 The components, layers, and functions shown in are intended to be illustrative only, and the components, layers, and functions may vary.
[0160] In at least one embodiment, the hardware and software layer 1302 includes hardware and software components. In at least one embodiment, the hardware components include mainframes, servers based on various RISC (Reduced Instruction Set Computer) architectures, various computing systems, supercomputing systems, storage devices, networks, networking components, and / or variants thereof. In at least one embodiment, the software components include network application server software, various application server software, various database software, and / or variants thereof.
[0161] In at least one embodiment, virtualization layer 1304 provides an abstraction layer from which the following exemplary virtual entities can be provided: virtual servers, virtual storage, virtual networks (including virtual private networks), virtual applications, virtual clients, and / or variations thereof.
[0162] In at least one embodiment, the management layer 1306 provides various functions. In at least one embodiment, resource provisioning provides dynamic acquisition of computing resources and other resources for performing tasks within a cloud computing environment. In at least one embodiment, metering provides usage tracking when resources are utilized within a cloud computing environment, as well as billing or invoicing for the consumption of these resources. In at least one embodiment, resources may include application software licenses. In at least one embodiment, security provides authentication for users and tasks, as well as protection of data and other resources. In at least one embodiment, a user interface provides access to the cloud computing environment for both users and system administrators. In at least one embodiment, service level management provides cloud computing resource allocation and management so that the required service level is met. In at least one embodiment, service level agreement (SLA) management provides pre-arrangement and acquisition of cloud computing resources, anticipating future demand for the cloud computing resources according to the SLA.
[0163] In at least one embodiment, workload layer 1308 provides functionality that utilizes a cloud computing environment. In at least one embodiment, workloads and functionality that can be provided from this layer include: mapping and navigation, software development and management, educational services, data analysis and processing, transaction processing, and service delivery.
[0164] Supercomputing
[0165] The following figures illustrate, but are not limited to, exemplary supercomputer-based systems that may be used to implement at least one embodiment.
[0166] In at least one embodiment, a supercomputer may refer to a hardware system that exhibits significant parallelism and includes at least one chip, wherein the chips in the system are interconnected by a network and are placed in a hierarchically organized housing. In at least one embodiment, a large hardware system that fills a computer room with several racks is at least one embodiment of a supercomputer, each rack containing several boards / rack modules, each board / rack module containing several chips all interconnected by an extensible network. In at least one embodiment, a single rack of such a large hardware system is at least one other embodiment of a supercomputer. In at least one embodiment, a single chip that exhibits significant parallelism and includes several hardware components can also be considered a supercomputer because as feature size may decrease, the amount of hardware that can be combined in a single chip may also increase.
[0167] Fig.14A chip-level supercomputer according to at least one embodiment is shown. In at least one embodiment, the main calculation is performed in a finite state machine (1404) called a thread unit inside an FPGA or ASIC chip. In at least one embodiment, a task and synchronization network (1402) connects the finite state machine and is used to dispatch threads and perform operations in the correct order. In at least one embodiment, a memory network (1406, 1410) is used to access multi-level partitioned on-chip cache levels (1408, 1412). In at least one embodiment, a memory controller (1416) and an off-chip memory network (1414) are used to access off-chip memory. In at least one embodiment, an I / O controller (1418) is used for cross-chip communication when the design is not suitable for a single logic chip.
[0168] Fig.15 A supercomputer at the rack module level is shown according to at least one embodiment. In at least one embodiment, within the rack module, there are multiple FPGA or ASIC chips (1502) connected to one or more DRAM cells (1504) that make up the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its neighboring FPGA / ASIC chips using a wide bus on the board with differential high-speed signaling (1506). In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.
[0169] Fig.16 A rack-scale supercomputer is shown in accordance with at least one embodiment. Fig.17 An overall system-level supercomputer according to at least one embodiment is shown. In at least one embodiment, see Fig.16 and Fig.17, between rack modules in a rack and across racks of the entire system, high-speed serial optical or copper cables (1602, 1702) are used to implement a scalable, possibly incomplete hypercube network. In at least one embodiment, one of the FPGA / ASIC chips of the accelerator is connected to a host system (1704) via a PCI-Express connection. In at least one embodiment, the host system includes a host microprocessor (1708) on which the software portion of the application runs and a memory consisting of one or more host memory DRAM units (1706) that are consistent with the memory on the accelerator. In at least one embodiment, the host system can be a separate module on one of the racks, or can be integrated with one of the modules of the supercomputer. In at least one embodiment, a circular topology of cube connections provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, a small group of FPGA / ASIC chips on a rack module can act as a single hypercube node, so that the total number of external links per group is increased compared to a single chip. In at least one embodiment, the group includes chips A, B, C, and D on a rack module, and the rack module has an internal wide differential bus connecting A, B, C, and D in a ring organization. In at least one embodiment, there are 12 serial communication cables connecting the rack module to the outside world. In at least one embodiment, chip A on the rack module is connected to serial communication cables 0, 1, 2. In at least one embodiment, chip B is connected to cables 3, 4, 5. In at least one embodiment, chip C is connected to 6, 7, 8. In at least one embodiment, chip D is connected to 9, 10, 11. In at least one embodiment, the entire group {A, B, C, D} that makes up the rack module can form a hypercube node within a supercomputer system, with up to 212=4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, in order for chip A to send a message out on link 4 of the group {A, B, C, D}, the message must first be routed to chip B using the on-board differential wide bus connection. In at least one embodiment, messages arriving on link 4 to the group {A, B, C, D} destined for chip A (i.e., arriving at B) must also first be routed to the correct destination chip (A) inside the group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes may also be implemented.
[0170] AI
[0171] The following figures illustrate, but are not limited to, exemplary artificial intelligence-based systems that may be used to implement at least one embodiment.
[0172] Fig.18AInference and / or training logic 1815 is shown for performing reasoning and / or training operations associated with one or more embodiments. Fig.18A and / or Fig.18B Provide details about the inference and / or training logic 1815.
[0173] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, code and / or data storage 1801 for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 1815 may include or be coupled to code and / or data storage 1801 for storing graph code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 1801 stores weight parameters and / or input / output data for each layer of a neural network that is trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 1801 may be included with other on-chip or off-chip data storage devices, including the processor's L1, L2, or L3 cache memory or system memory.
[0174] In at least one embodiment, any portion of code and / or data storage 1801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 1801 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or code and / or data storage 1801 is internal or external to a processor, in at least one embodiment, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of a neural network, or some combination of these factors.
[0175] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, code and / or data storage 1805 for storing reverse and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 1805 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during reverse back propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 1815 may include or be coupled to the code and / or data storage 1805 to store graph code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).
[0176] In at least one embodiment, the code (such as graph code) causes the weights or other parameter information to be loaded into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, any portion of the code and / or data storage 1805 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 1805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1805 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage device. In at least one embodiment, the choice of whether the code and / or data storage 1805 is internal or external to the processor, in at least one embodiment, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0177] In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be separate storage structures. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be a combined storage structure. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 1801 and code and / or data storage 1805 may be included with other on-chip or off-chip data storage (including the processor's L1, L2, or L3 cache or system memory).
[0178] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1810, including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or as directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 1820, which is a function of input / output and / or weight parameter data stored in code and / or data storage 1801 and / or code and / or data storage 1805. In at least one embodiment, activations stored in activation storage 1820 are generated based on linear algebra and / or matrix-based math performed by ALU 1810 in response to executing instructions or other code, where weight values stored in code and / or data store 1805 and / or data store 1801 are used as operands along with other values (such as bias values, gradient information, momentum values, or other parameters or hyperparameters), any or all of which may be stored in code and / or data store 1805 or code and / or data store 1801 or in another storage on or off-chip.
[0179] In at least one embodiment, one or more ALUs 1810 are included within one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 1810 may be external to a processor or other hardware logic devices or circuits (e.g., a coprocessor) that use them. In at least one embodiment, ALUs 1810 may be included within an execution unit of a processor or otherwise within an ALU bank accessible by an execution unit of a processor, the execution units of which may be within the same processor or distributed between different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data storage 1801, code and / or data storage 1805, and activation storage 1820 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 1820 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In addition, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry and retrieved and / or processed using the processor's retrieval, decoding, scheduling, execution, retirement, and / or other logic circuitry.
[0180] In at least one embodiment, activation storage 1820 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage device. In at least one embodiment, activation storage 1820 may be completely or partially within or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 1820 is internal or external to the processor, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training of the neural network, or some combination of these factors.
[0181] In at least one embodiment, Fig.18A The inference and / or training logic 1815 shown in FIG. 1 may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Fig.18A The inference and / or training logic 1815 shown in FIG. 1 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).
[0182] Fig.18B Inference and / or training logic 1815 is shown in accordance with at least one embodiment. In at least one embodiment, the reasoning and / or training logic 1815 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise used exclusively in conjunction with weight values or other information corresponding to one or more neuron layers within a neural network. In at least one embodiment, Fig.18B The inference and / or training logic 1815 shown in FIG. 1 may be combined with an application specific integrated circuit (ASIC) (e.g., from Google Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Fig.18B The inference and / or training logic 1815 shown in FIG. 1 may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 1815 includes, but is not limited to, code and / or data storage 1801 and code and / or data storage 1805, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Fig.18B In at least one embodiment described in , each of code and / or data storage 1801 and code and / or data storage 1805 is associated with a dedicated computing resource (e.g., computing hardware 1802 and computing hardware 1806), respectively. In at least one embodiment, each of computing hardware 1802 and computing hardware 1806 includes one or more ALUs that perform mathematical functions (such as linear algebraic functions) solely on the information stored in code and / or data storage 1801 and code and / or data storage 1805, respectively, with the results being stored in activation storage 1820.
[0183] In at least one embodiment, each code and / or data storage 1801 and 1805 and corresponding computing hardware 1802 and 1806, respectively, corresponds to a different layer of the neural network, such that the resulting activation from one storage / computation pair 1801 / 1802 in the code and / or data storage 1801 and computing hardware 1802 is provided as input to the next storage / computation pair 1805 / 1806 in the code and / or data storage 1805 and computing hardware 1806, so as to mirror the conceptual organization of the neural network. In at least one embodiment, each of the storage / computation pairs 1801 / 1802 and 1805 / 1806 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) after or in parallel with the storage / computation pairs 1801 / 1802 and 1805 / 1806 may be included in the inference and / or training logic 1815.
[0184] Fig.19 The training and deployment of a deep neural network according to at least one embodiment is shown. In at least one embodiment, an untrained neural network 1906 is trained using a training data set 1902. In at least one embodiment, the training framework 1904 is a PyTorch framework, while in other embodiments, the training framework 1904 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 and enables it to be trained using the processing resources described herein to generate a trained neural network 1908. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training may be performed in a supervised, partially supervised, or unsupervised manner.
[0185] In at least one embodiment, untrained neural network 1906 is trained using supervised learning, where training data set 1902 includes inputs paired with expected outputs for the inputs, or where training data set 1902 includes inputs with known outputs, and the outputs of neural network 1906 are manually graded. In at least one embodiment, untrained neural network 1906 is trained in a supervised manner, and inputs from training data set 1902 are processed and the resulting outputs are compared to a set of expected or desired outputs. In at least one embodiment, errors are then back-propagated through untrained neural network 1906. In at least one embodiment, training framework 1904 adjusts the weights that control untrained neural network 1906. In at least one embodiment, training framework 1904 includes tools for monitoring how well untrained neural network 1906 converges toward a model (such as trained neural network 1908) that is suitable for generating correct answers (such as results 1914) based on input data (such as new data set 1912). In at least one embodiment, the training framework 1904 repeatedly trains the untrained neural network 1906 while adjusting the weights using a loss function and an adjustment algorithm (such as stochastic gradient descent) to refine the output of the untrained neural network 1906. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 until the untrained neural network 1906 achieves a desired accuracy. In at least one embodiment, the trained neural network 1908 can then be deployed to implement any number of machine learning operations.
[0186] In at least one embodiment, untrained neural network 1906 is trained using unsupervised learning, where untrained neural network 1906 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training data set 1902 will include input data without any associated output data or "ground truth" data. In at least one embodiment, untrained neural network 1906 can learn groupings within training data set 1902 and can determine how individual inputs are related to untrained data set 1902. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in trained neural network 1908 that is capable of performing operations useful in reducing the dimensionality of new data set 1912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new data set 1912 that deviate from the normal pattern of new data set 1912.
[0187] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mixture of labeled and unlabeled data is included in the training data set 1902. In at least one embodiment, the training framework 1904 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1908 to adapt to new data sets 1912 without forgetting the knowledge infused into the trained neural network 1408 during initial training.
[0188] 5G Network
[0189] The following figures illustrate, but are not limited to, exemplary 5G network-based systems that may be used to implement at least one embodiment.
[0190] Fig. 20 The architecture of a system 2000 of a network according to at least one embodiment is shown. In at least one embodiment, the system 2000 is shown to include user equipment (UE) 2002 and UE 2004. In at least one embodiment, UE 2002 and 2004 are shown as smart phones (e.g., handheld touch screen mobile computing devices that can connect to one or more cellular networks), but may also include any mobile or non-mobile computing device, such as a personal digital assistant (PDA), a pager, a laptop computer, a desktop computer, a wireless handheld device, or any computing device that includes a wireless communication interface.
[0191] In at least one embodiment, any one of UE 2002 and UE 2004 may include an Internet of Things (IoT) UE, which may include a network access layer designed for low-power IoT applications that utilize short-lived UE connections. In at least one embodiment, the IoT UE may utilize technologies such as for exchanging data with an MTC server or device via a public land mobile network (PLMN), proximity-based services (ProSe) or device-to-device (D2D) communication, a sensor network, or an IoT network, such as machine-to-machine (M2M) or machine-type communication (MTC). In at least one embodiment, the M2M or MTC data exchange may be a machine-initiated data exchange. In at least one embodiment, the IoT network describes interconnected IoT UEs, which may include uniquely identifiable embedded computing devices (within the Internet infrastructure) with short-lived connections. In at least one embodiment, the IoT UE may execute background applications (e.g., keep-alive messages, status updates, etc.) to facilitate connectivity to the IoT network.
[0192] In at least one embodiment, UE 2002 and UE 2004 may be configured to connect (e.g., communicatively couple) to a radio access network (RAN) 2016. In at least one embodiment, RAN 2016 may be an evolved universal mobile telecommunications system (UMTS) terrestrial radio access network (E-UTRAN), a NextGen RAN (NG RAN), or some other type of RAN in at least one embodiment. In at least one embodiment, UE 2002 and UE 2004 utilize connection 2012 and connection 2014, respectively, each connection including a physical communication interface or layer. In at least one embodiment, connections 2012 and 2014 are shown as air interfaces for achieving communicatively coupling, and may be consistent with a cellular communication protocol, such as a global system for mobile communications (GSM) protocol, a code division multiple access (CDMA) network protocol, a push-to-talk (PTT) protocol, a cellular PTT (POC) protocol, a universal mobile telecommunications system (UMTS) protocol, a 3GPP long term evolution (LTE) protocol, a fifth generation (5G) protocol, a new radio (NR) protocol, and variations thereof.
[0193] In at least one embodiment, the UEs 2002 and 2004 may also directly exchange communication data via the ProSe interface 2006. In at least one embodiment, the ProSe interface 2006 may alternatively be referred to as a side link interface, which includes one or more logical channels, including but not limited to a physical side link control channel (PSCCH), a physical side link shared channel (PSSCH), a physical side link discovery channel (PSDCH), and a physical side link broadcast channel (PSBCH).
[0194] In at least one embodiment, UE 2004 is shown as being configured to access access point (AP) 2010 via connection 2008. In at least one embodiment, connection 2008 may include a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein AP 2010 would include wireless fidelity. In at least one embodiment, AP 2010 is shown connected to the Internet and not to the core network of the wireless system.
[0195] In at least one embodiment, the RAN 2016 may include one or more access nodes that enable connections 2012 and 2014. In at least one embodiment, these access nodes (ANs) may be referred to as base stations (BS), NodeBs, evolved NodeBs (eNBs), next generation NodeBs (gNBs), RAN nodes, etc., and may include ground stations (e.g., ground access points) or satellite stations that provide coverage within a geographic area (e.g., a cell). In at least one embodiment, the RAN 2016 may include one or more RAN nodes (e.g., macro RAN nodes 2018) for providing macro cells and one or more RAN nodes (e.g., low power (LP) RAN nodes 2020) for providing femto cells or pico cells (e.g., cells with smaller coverage areas, smaller user capacity, or higher bandwidth than macro cells).
[0196] In at least one embodiment, any of the RAN nodes 2018 and 2020 may terminate the air interface protocol and may be the first point of contact for the UEs 2002 and 2004. In at least one embodiment, any of the RAN nodes 2018 and 2020 may implement various logical functions of the RAN 2016, including but not limited to radio network controller (RNC) functions such as radio bearer management, uplink and downlink dynamic radio resource management, and data packet scheduling and mobility management.
[0197] In at least one embodiment, UE 2002 and UE 2004 may be configured to communicate with each other or with any of RAN node 2018 and RAN node 2020 through a multi-carrier communication channel using orthogonal frequency division multiplexing (OFDM) communication signals according to various communication techniques, such as but not limited to orthogonal frequency division multiple access (OFDMA) communication techniques (e.g., for downlink communication) or single carrier frequency division multiple access (SC-FDMA) communication techniques (e.g., for uplink and ProSe or sidelink communication), and / or variants thereof. In at least one embodiment, the OFDM signal may include multiple orthogonal subcarriers.
[0198] In at least one embodiment, the downlink resource grid can be used for downlink transmission from any one of the RAN nodes 2018 and 2020 to the UEs 2002 and 2004, while the uplink transmission can utilize similar techniques. In at least one embodiment, the grid can be a time-frequency grid called a resource grid or a time-frequency resource grid, which is a physical resource in the downlink in each time slot. In at least one embodiment, this time-frequency plane representation is a common practice in OFDM systems, which makes it intuitive for radio resource allocation. In at least one embodiment, each column and each row of the resource grid corresponds to an OFDM symbol and an OFDM subcarrier, respectively. In at least one embodiment, the duration of the resource grid in the time domain corresponds to a time slot in a radio frame. In at least one embodiment, the minimum time-frequency unit in the resource grid is represented as a resource element. In at least one embodiment, each resource grid includes a plurality of resource blocks, which describe the mapping of certain physical channels to resource elements. In at least one embodiment, each resource block includes a set of resource elements. In at least one embodiment, in the frequency domain, this can represent the minimum number of resources that can currently be allocated. In at least one embodiment, there are several different physical downlink channels that are transmitted using such resource blocks.
[0199] In at least one embodiment, the physical downlink shared channel (PDSCH) may carry user data and higher layer signaling to UEs 2002 and 2004. In at least one embodiment, the physical downlink control channel (PDCCH) may carry information about the transport format and resource allocation associated with the PDSCH channel, etc. In at least one embodiment, it may also inform UEs 2002 and 2004 of the transport format, resource allocation, and HARQ (Hybrid Automatic Repeat Request) information related to the uplink shared channel. In at least one embodiment, typically, downlink scheduling (allocation of control and shared channel resource blocks to UEs 2002 within a cell) may be performed at any one of the RAN nodes 2018 and 2020 based on channel quality information fed back from any one of the UEs 2002 and 2004. In at least one embodiment, downlink resource allocation information may be sent on a PDCCH for (e.g., allocated to) each of the UEs 2002 and 2004.
[0200] In at least one embodiment, the PDCCH may use control channel elements (CCEs) to transmit control information. In at least one embodiment, before being mapped to resource elements, the PDCCH complex symbols may first be organized into quaternions, which may then be permuted using a sub-block interleaver for rate matching. In at least one embodiment, each PDCCH may be transmitted using one or more of these CCEs, where each CCE may correspond to nine sets of four physical resource elements referred to as resource element groups (REGs). In at least one embodiment, four quadrature phase shift keying (QPSK) symbols may be mapped to each REG. In at least one embodiment, one or more CCEs may be used to transmit the PDCCH, depending on the size of the downlink control information (DCI) and the channel conditions. In at least one embodiment, there may be four or more different PDCCH formats (e.g., aggregation levels, L=1, 2, 4, or 8) defined in LTE with different numbers of CCEs.
[0201] In at least one embodiment, an enhanced physical downlink control channel (EPDCCH) using PDSCH resources may be used for control information transmission. In at least one embodiment, one or more enhanced control channel elements (ECCEs) may be used to transmit EPDCCH. In at least one embodiment, each ECCE may correspond to nine sets of four physical resource elements referred to as enhanced resource element groups (EREGs). In at least one embodiment, ECCEs may have other numbers of EREGs in some cases.
[0202] In at least one embodiment, the RAN 2016 is shown as being communicatively coupled to a core network (CN) 2038 via an S1 interface 2022. In at least one embodiment, the CN 2038 may be an evolved packet core (EPC) network, a NextGen packet core (NPC) network, or some other type of CN. In at least one embodiment, the S1 interface 2022 is divided into two parts: an S1-U interface 2026, which carries traffic data between the RAN nodes 2018 and 2020 and the serving gateway (S-GW) 2030; and an S1-mobility management entity (MME) interface 2024, which is a signaling interface between the RAN nodes 2018 and 2020 and the MME 2028.
[0203] In at least one embodiment, CN 2038 includes MME 2028, S-GW 2030, Packet Data Network (PDN) Gateway (P-GW) 2034 and Home Subscriber Server (HSS) 2032. In at least one embodiment, MME 2028 can be functionally similar to the control plane of a traditional Serving General Packet Radio Service (GPRS) Support Node (SGSN). In at least one embodiment, MME 2028 can manage mobility aspects in access, such as gateway selection and tracking area list management. In at least one embodiment, HSS 2032 can include a database for network users, which includes subscription-related information for supporting network entities to handle communication sessions. In at least one embodiment, CN 2038 can include one or more HSS 2032, depending on the number of mobile users, the capacity of the equipment, the organization of the network, etc. In at least one embodiment, HSS 2032 can provide support for routing / roaming, authentication, authorization, naming / addressing resolution, location dependency, etc.
[0204] In at least one embodiment, the S-GW 2030 may terminate the S1 interface 2022 towards the RAN 2016 and route data packets between the RAN 2016 and the CN 2038. In at least one embodiment, the S-GW 2030 may be a local mobility anchor for inter-RAN node handovers and may also provide an anchor for inter-3GPP mobility. In at least one embodiment, other responsibilities may include lawful interception, charging, and some policy enforcement.
[0205] In at least one embodiment, the P-GW 2034 may terminate the SGi interface towards the PDN.
[0206] In at least one embodiment, the P-GW 2034 can route data packets between the EPC network 2038 and an external network (such as a network including an application server 2040 (or referred to as an application function (AF))) via an Internet Protocol (IP) interface 2042. In at least one embodiment, the application server 2040 can be an element that uses a core network (e.g., a UMTS packet service (PS) domain, a LTE PS data service, etc.) to provide applications that use IP bearer resources. In at least one embodiment, the P-GW 2034 is shown as being communicatively coupled to the application server 2040 via an IP communication interface 2042. In at least one embodiment, the application server 2040 can also be configured to support one or more communication services (e.g., voice over Internet protocol (VoIP) sessions, PTT sessions, group communication sessions, social network services, etc.) for UEs 2002 and 2004 via CN 2038.
[0207] In at least one embodiment, the P-GW 2034 may also be a node for policy enforcement and charging data collection. In at least one embodiment, the policy and charging enforcement function (PCRF) 2036 is a policy and charging control element of the CN 2038. In at least one embodiment, in a non-roaming scenario, a single PCRF may exist in the home public land mobile network (HPLMN) associated with the UE's Internet Protocol Connectivity Access Network (IP-CAN) session. In at least one embodiment, in a roaming scenario with local traffic breakout, there may be two PCRFs associated with the UE's IP-CAN session: the home PCRF (H-PCRF) within the HPLMN and the visited PCRF (V-PCRF) within the visited public land mobile network (VPLMN). In at least one embodiment, the PCRF 2036 may be communicatively coupled to the application server 2040 via the P-GW 2034. In at least one embodiment, the application server 2040 may signal the PCRF 2036 to indicate a new service flow and select appropriate quality of service (QoS) and charging parameters. In at least one embodiment, the PCRF 2036 can supply this rule to a policy and charging enforcement function (PCEF) (not shown) with an appropriate traffic flow template (TFT) and QoS class (QCI) identifier, which initiates the QoS and charging specified by the application server 2040.
[0208] Fig.21 The architecture of a system 2100 of a network according to some embodiments is shown. In at least one embodiment, the system 2100 is shown to include a UE 2102, a 5G access node or RAN node (shown as (R)AN node 2108), a user plane function (shown as UPF 2104), a data network (DN 2106), which in at least one embodiment can be an operator service, Internet access or a third party service, and a 5G core network (5GC) (shown as CN 2110).
[0209] In at least one embodiment, CN 2110 includes an authentication server function (AUSF 2114); a core access and mobility management function (AMF 2112); a session management function (SMF 2118); a network exposure function (NEF 2116); a policy control function (PCF 2122); a network function (NF) repository function (NRF 2120); a unified data management (UDM 2124); and an application function (AF 2126). In at least one embodiment, CN 2110 may also include other elements not shown, such as a structured data storage network function (SDSF), an unstructured data storage network function (UDSF), and variations thereof.
[0210] In at least one embodiment, UPF 2104 may act as an anchor point for intra-RAT and inter-RAT mobility, an external PDU session point interconnected to DN 2106, and a branch point supporting multi-homing PDU sessions. In at least one embodiment, UPF 2104 may also perform packet routing and forwarding, packet inspection, user plane portion of policy rule implementation, legal interception of packets (UP collection); service usage reporting, QoS processing for the user plane (e.g., packet filtering, gating, UL / DL rate enforcement), uplink service verification (e.g., SDF to QoS flow mapping), transport level packet marking in uplink and downlink, and downlink packet buffering and downlink data notification triggering. In at least one embodiment, UPF 2104 may include an uplink classifier to support routing of service flows to data networks. In at least one embodiment, DN 2106 may represent various network operator services, Internet access, or third-party services.
[0211] In at least one embodiment, the AUSF 2114 may store data for authentication of the UE 2102 and handle authentication-related functions. In at least one embodiment, the AUSF 2114 may facilitate a common authentication framework for various access types.
[0212] In at least one embodiment, AMF 2112 may be responsible for registration management (e.g., for registering UE 2102, etc.), connection management, reachability management, mobility management, and lawful interception of AMF-related events, as well as access authentication and authorization. In at least one embodiment, AMF 2112 may provide transmission of SM messages for SMF 2118 and act as a transparent proxy for routing SM messages. In at least one embodiment, AMF 2112 may also provide UE 2102 with SMS function (SMSF) ( Fig.21 In at least one embodiment, the AMF 2112 may include a security context management (SCM) function that receives keys from the SEA that it uses to derive access network-specific keys. In addition, in at least one embodiment, the AMF 2112 may be a termination point for the RAN CP interface (N2 reference point), a termination point for NAS (NI) signaling, and perform NAS encryption and integrity protection.
[0213] In at least one embodiment, the AMF 2112 may also support NAS signaling with the UE 2102 over the N3 interworking function (IWF) interface. In at least one embodiment, the N3IWF may be used to provide access to untrusted entities. In at least one embodiment, the N3IWF may be the termination point of the N2 and N3 interfaces for the control plane and the user plane, respectively, and may therefore process N2 signaling from the SMF and AMF for PDU sessions and QoS, encapsulate / decapsulate packets for IPSec and N3 tunnels, mark N3 user plane packets in the uplink, and implement QoS corresponding to the N3 packet marking taking into account the QoS requirements associated with such marking received over N2. In at least one embodiment, the N3IWF may also relay uplink and downlink control plane NAS (NI) signaling between the UE 2102 and the AMF 2112, and relay uplink and downlink user plane packets between the UE 2102 and the UPF 2104. In at least one embodiment, the N3IWF also provides a mechanism for establishing an IPsec tunnel with UE 2102.
[0214] In at least one embodiment, SMF 2118 may be responsible for session management (e.g., session establishment, modification, and release, including tunnel maintenance between UPF and AN nodes); UE IP address allocation and management (including optional authorization); selection and control of UP functions; configuration of traffic steering at UPF to route traffic to the appropriate destination; interface termination towards policy control function; control part of policy enforcement and QoS; lawful interception (for SM events and interface to LI system); termination of the SM part of NAS messages; downlink data notification; initiator of AN-specific SM information, which is sent to AN via AMF on N2; determination of SSC mode for session. In at least one embodiment, SMF 2118 may include the following roaming functions: handling local implementation to apply QoS SLAB (VPLMN); charging data collection and charging interface (VPLMN); lawful interception (in VPLMN for SM events and interface to LI system); support interaction with external DN to transport signaling for PDU session authorization / authentication by external DN.
[0215] In at least one embodiment, the NEF 2116 may provide a means for securely exposing services and capabilities provided by 3GPP network functions to third parties, internal exposure / re-exposure, application functions (e.g., AF 2126), edge computing or fog computing systems, etc. In at least one embodiment, the NEF 2116 may authenticate, authorize and / or throttle the AF. In at least one embodiment, the NEF 2116 may also convert information exchanged with the AF 2126 and information exchanged with internal network functions. In at least one embodiment, the NEF 2116 may convert between AF service identifiers and internal 5GC information. In at least one embodiment, the NEF 2116 may also receive information from other network functions (NFs) based on the exposed capabilities of the other network functions. In at least one embodiment, the information may be stored at the NEF 2116 as structured data, or at a data storage NF using a standardized interface. In at least one embodiment, the stored information may then be re-exposed by the NEF 2116 to other NFs and AFs, and / or used for other purposes, such as analysis.
[0216] In at least one embodiment, NRF 2120 can support service discovery functionality, receive NF discovery requests from NF instances, and provide information of discovered NF instances to NF instances. In at least one embodiment, NRF 2120 also maintains information of available NF instances and the services they support.
[0217] In at least one embodiment, PCF 2122 can provide policy rules to the control plane functions to implement them, and can also support a unified policy framework to manage network behavior. In at least one embodiment, PCF 2122 can also implement a front end (FE) for accessing subscription information related to policy decisions in the UDR of UDM 2124.
[0218] In at least one embodiment, the UDM 2124 may process subscription-related information to support network entities handling communication sessions, and may store subscription data for the UE 2102. In at least one embodiment, the UDM 2124 may include two parts, an application FE and a user data repository (UDR). In at least one embodiment, the UDM may include a UDM FE that is responsible for handling credentials, location management, subscription management, etc. In at least one embodiment, several different front ends may serve the same user in different transactions. In at least one embodiment, the UDM-FE accesses the sub-subscription information stored in the UDR and performs authentication credential processing; user identity processing; access authorization; registration / mobility management; and subscription management. In at least one embodiment, the UDR may interact with the PCF 2122. In at least one embodiment, the UDM 2124 may also support SMS management, where the SMS-FE implements similar application logic as described above.
[0219] In at least one embodiment, AF 2126 can provide application impact on service routing, access to network capability exposure (NCE), and interaction with policy framework for policy control. In at least one embodiment, NCE can be a mechanism that allows 5GC and AF 2126 to provide information to each other via NEF 2116, which can be used for edge computing implementation. In at least one embodiment, network operators and third-party services can be hosted near the attachment access point of UE 2102 to achieve efficient service delivery through reduced end-to-end delay and load on the transmission network. In at least one embodiment, for edge computing implementation, 5GC can select UPF 2104 close to UE 2102 and perform service guidance from UPF 2104 to DN 2106 via N6 interface. In at least one embodiment, this can be based on UE subscription data, UE location and information provided by AF 2126. In at least one embodiment, AF 2126 can affect UPF (re) selection and service routing. In at least one embodiment, based on operator deployment, the network operator may allow the AF 2126 to interact directly with the relevant NFs when the AF 2126 is considered a trusted entity.
[0220] In at least one embodiment, CN 2110 may include SMSF, which may be responsible for SMS subscription checking and verification, and relaying SM messages to / from UE 2102 to / from other entities, such as SMS-GMSC / IWMSC / SMS routers. In at least one embodiment, SMS may also interact with AMF 2112 and UDM 2124 for notification procedures that UE 2102 is available for SMS delivery (e.g., setting a UE unreachable flag and notifying UDM 2124 when UE 2102 is available for SMS).
[0221] In at least one embodiment, system 2100 may include the following service-based interfaces: Namf: a service-based interface exposed by AMF; Nsmf: a service-based interface exposed by SMF; Nnef: a service-based interface exposed by NEF; Npcf: a service-based interface exposed by PCF; Nudm: a service-based interface exposed by UDM; Naf: a service-based interface exposed by AF; Nnrf: a service-based interface exposed by NRF; and Nausf: a service-based interface exposed by AUSF.
[0222] In at least one embodiment, the system 2100 may include the following reference points: N1: a reference point between the UE and the AMF; N2: a reference point between the (R)AN and the AMF; N3: a reference point between the (R)AN and the UPF; N4: a reference point between the SMF and the UPF; and N6: a reference point between the UPF and the data network. In at least one embodiment, there may be more reference points and / or service-based interfaces between NF services in the NF, however, for clarity, these interfaces and reference points have been omitted. In at least one embodiment, the NS reference point may be between the PCF and the AF; the N7 reference point may be between the PCF and the SMF; the N11 reference point is between the AMF and the SMF, and so on. In at least one embodiment, the CN 2110 may include an Nx interface, which is an inter-CN interface between the MME and the AMF 2112, so as to enable interworking between the CN 2110 and the CN 7221.
[0223] In at least one embodiment, the system 2100 may include multiple RAN nodes (such as (R)AN node 2108), wherein an Xn interface is defined between two or more (R)AN nodes 2108 (e.g., gNBs) connected to the 5GC 410, between the (R)AN node 2108 (e.g., gNBs) and an eNB (e.g., a macro RAN node) connected to the CN 2110, and / or between two eNBs connected to the CN 2110.
[0224] In at least one embodiment, the Xn interface may include an Xn user plane (Xn-U) interface and an Xn control plane (Xn-C) interface. In at least one embodiment, the Xn-U may provide non-guaranteed delivery of user plane PDUs and support / provide data forwarding and flow control functions. In at least one embodiment, the Xn-C may provide management and error handling functions, functions for managing the Xn-C interface; mobility support for UE 2102 in connected mode (e.g., CM-CONNECTED), including functions for managing UE mobility in connected mode between one or more (R)AN nodes 2108. In at least one embodiment, the mobility support may include context transfer from the old (source) serving (R)AN node 2108 to the new (target) serving (R)AN node 2108; and control of the user plane tunnel between the old (source) serving (R)AN node 2108 and the new (target) serving (R)AN node 2108.
[0225] In at least one embodiment, the protocol stack of Xn-U may include a transport network layer built on an Internet Protocol (IP) transport layer and a GTP-U layer for carrying user plane PDUs on top of UDP and / or one or more IP layers. In at least one embodiment, the Xn-C protocol stack may include an application layer signaling protocol (referred to as the Xn Application Protocol (Xn-AP)) and a transport network layer built on the SCTP layer. In at least one embodiment, the SCTP layer may be on top of the IP layer. In at least one embodiment, the SCTP layer provides guaranteed delivery of application layer messages. In at least one embodiment, in the transport IP layer, point-to-point transmission is used to deliver signaling PDUs. In at least one embodiment, the Xn-U protocol stack and / or the Xn-C protocol stack may be the same or similar to the user plane and / or control plane protocol stacks shown and described herein.
[0226] Fig. 22 2004 ), RAN 2016 , and MME 2028 .
[0227] In at least one embodiment, the PHY layer 2202 may send or receive information used by the MAC layer 2204 over one or more air interfaces. In at least one embodiment, the PHY layer 2202 may also perform link adaptation or adaptive modulation and coding (AMC), power control, cell search (e.g., for initial synchronization and handover purposes), and other measurements used by higher layers (e.g., RRC layer 2210). In at least one embodiment, the PHY layer 2202 may further perform error detection on transport channels, forward error correction (FEC) encoding / decoding of transport channels, modulation / demodulation of physical channels, interleaving, rate matching, mapping to physical channels, and multiple-input multiple-output (MIMO) antenna processing.
[0228] In at least one embodiment, the MAC layer 2204 may perform mapping between logical channels and transport channels, multiplexing MAC service data units (SDUs) from one or more logical channels onto transport blocks (TBs) to be delivered to the PHY via transport channels, demultiplexing MAC SDUs from transport blocks (TBs) delivered from the PHY via transport channels to one or more logical channels, multiplexing MAC SDUs onto TBs, scheduling information reporting, error correction through hybrid automatic repeat request (HARD), and logical channel prioritization.
[0229] In at least one embodiment, the RLC layer 2206 can operate in multiple operation modes, including: transparent mode (TM), unacknowledged mode (UM), and acknowledged mode (AM). In at least one embodiment, the RLC layer 2206 can perform transmission of upper layer protocol data units (PDUs), error correction through automatic repeat request (ARQ) for AM data transmission, and concatenation, segmentation, and reassembly of RLC SDUs for UM and AM data transmission. In at least one embodiment, the RLC layer 2206 can also perform re-segmentation of RLC data PDUs for AM data transmission, reordering of RLC data PDUs for UM and AM data transmission, detection of duplicate data for UM and AM data transmission, discarding RLC SDUs for UM and AM data transmission, detecting protocol errors for AM data transmission, and performing RLC reconstruction.
[0230] In at least one embodiment, the PDCP layer 2208 can perform header compression and decompression of IP data, maintain the PDCP sequence number (SN), perform in-sequence delivery of higher layer PDUs when reestablishing lower layers, eliminate duplication of lower layer SDUs when reestablishing lower layers for radio bearers mapped on RLC AM, encrypt and decrypt control plane data, perform integrity protection and integrity verification on control plane data, control timer-based data discard, and perform security operations (e.g., encryption, decryption, integrity protection, integrity verification, etc.).
[0231] In at least one embodiment, the main services and functions of the RRC layer 2210 may include broadcasting of system information (e.g., included in a master information block (MIB) or a system information block (SIB) related to a non-access stratum (NAS)), broadcasting of system information related to an access stratum (AS), paging, establishment, maintenance, and release of an RRC connection between a UE and an E-UTRAN (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), establishment, configuration, maintenance, and release of point-to-point radio bearers, security functions including key management, inter-radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting. In at least one embodiment, the MIB and SIB may include one or more information elements (IEs), each of which may include a separate data field or data structure.
[0232] In at least one embodiment, the UE 2002 and the RAN 2016 may utilize a Uu interface (e.g., an LTE-Uu interface) to exchange control plane data via a protocol stack including a PHY layer 2202, a MAC layer 2204, an RLC layer 2206, a PDCP layer 2208, and an RRC layer 2210.
[0233] In at least one embodiment, the non-access stratum (NAS) protocol (NAS protocol 2212) forms the highest layer of the control plane between the UE 2002 and the MME 2028. In at least one embodiment, the NAS protocol 2212 supports the mobility and session management procedures of the UE 2002 to establish and maintain an IP connection between the UE 2002 and the P-GW 2034.
[0234] In at least one embodiment, the Si application protocol (Si-AP) layer (Si-AP layer 2222) can support the functions of the Si interface and include basic procedures (EP). In at least one embodiment, the EP is an interaction unit between the RAN 2016 and the CN 2028. In at least one embodiment, the S1-AP layer services may include two groups: UE associated services and non-UE associated services. In at least one embodiment, these services perform functions including but not limited to: E-UTRAN Radio Access Bearer (E-RAB) management, UE capability indication, mobility, NAS signaling transmission, RAN Information Management (RIM) and configuration transfer.
[0235] In at least one embodiment, a stream control transmission protocol (SCTP) layer (alternatively referred to as a stream control transmission protocol / Internet protocol (SCTP / IP) layer) (SCTP layer 2220) can ensure reliable delivery of signaling messages between the RAN 2016 and the MME 2028 based in part on the IP protocol supported by the IP layer 2218. In at least one embodiment, the L2 layer 2216 and the L1 layer 2214 can refer to communication links (e.g., wired or wireless) used by the RAN node and the MME to exchange information.
[0236] In at least one embodiment, the RAN 2016 and one or more MMEs 2028 may utilize an S1-MME interface to exchange control plane data via a protocol stack including an L1 layer 2214 , an L2 layer 2216 , an IP layer 2218 , an SCTP layer 2220 , and a Si-AP layer 2222 .
[0237] Fig.23 2300 is a diagram of a user plane protocol stack according to at least one embodiment. In at least one embodiment, the user plane 2300 is shown as a communication protocol stack between the UE 2002, the RAN 2016, the S-GW 2030, and the P-GW 2034. In at least one embodiment, the user plane 2300 can utilize the same protocol layer as the control plane 2200. In at least one embodiment, the UE 2002 and the RAN 2016 can utilize a Uu interface (e.g., an LTE-Uu interface) to exchange user plane data via a protocol stack including a PHY layer 2202, a MAC layer 2204, an RLC layer 2206, and a PDCP layer 2208.
[0238] In at least one embodiment, the general packet radio service (GPRS) tunneling protocol (GTP-U) layer (GTP-U layer 2304) for the user plane can be used to carry user data within the GPRS core network and between the radio access network and the core network. In at least one embodiment, the transmitted user data can be a packet in any format of IPv4, IPv6 or PPP format. In at least one embodiment, the UDP and IP security (UDP / IP) layer (UDP / IP layer 2302) can provide a checksum of data integrity, a port number for addressing different functions at the source and destination, and encryption and authentication of the selected data stream. In at least one embodiment, the RAN 2016 and the S-GW 2030 can utilize the S1-U interface to exchange user plane data via a protocol stack including L1 layer 2214, L2 layer 2216, UDP / IP layer 2302 and GTP-U layer 2304. In at least one embodiment, the S-GW 2030 and the P-GW 2034 may utilize an S5 / S8a interface to exchange user plane data via a protocol stack including an L1 layer 2214, an L2 layer 2216, a UDP / IP layer 2302, and a GTP-U layer 2304. In at least one embodiment, as described above with respect to Fig. 22 As discussed, the NAS protocol supports the mobility of UE 2002 and session management procedures to establish and maintain an IP connection between UE 2002 and P-GW 2034.
[0239] Fig.24 Components 2400 of a core network according to at least one embodiment are shown. In at least one embodiment, the components of CN 2038 may be implemented in one physical node or in separate physical nodes, the separate physical nodes including components for reading and executing instructions from a machine-readable medium or a computer-readable medium (e.g., a non-transitory machine-readable storage medium). In at least one embodiment, network function virtualization (NFV) is used to virtualize any or all of the above network node functions via executable instructions stored in one or more computer-readable storage media (described in further detail below). In at least one embodiment, a logical instantiation of CN 2038 may be referred to as a network slice 2402 (e.g., network slice 2402 is shown to include HSS 2032, MME 2028, and S-GW 2030). In at least one embodiment, a logical instantiation of a portion of CN 2038 may be referred to as a network sub-slice 2404 (e.g., network sub-slice 2404 is shown to include P-GW 2034 and PCRF 2036).
[0240] In at least one embodiment, the NFV architecture and infrastructure can be used to virtualize one or more network functions onto physical resources including a combination of industry standard server hardware, storage hardware, or switches, which network functions can alternatively be performed by dedicated hardware. In at least one embodiment, the NFV system can be used to perform a virtual or reconfigurable implementation of one or more EPC components / functions.
[0241] Fig.25 2500 for supporting network function virtualization (NFV) according to at least one embodiment. In at least one embodiment, the system 2500 is shown to include a virtualization infrastructure manager (shown as VIM 2502), a network function virtualization infrastructure (shown as NFVI 2504), a VNF manager (shown as VNFM 2506), a virtualized network function (shown as VNF 2508), an element manager (shown as EM 2510), a NFV orchestrator (shown as NFVO 2512), and a network manager (shown as NM 2514).
[0242] In at least one embodiment, the VIM 2502 manages resources of the NFVI 2504. In at least one embodiment, the NFVI 2504 may include physical or virtual resources and applications (including hypervisors) used to execute the system 2500. In at least one embodiment, the VIM 2502 may utilize the NFVI 2504 to manage the lifecycle of virtual resources (e.g., creation, maintenance, and teardown of virtual machines (VMs) associated with one or more physical resources), track VM instances, track performance, failures, and security of VM instances and associated physical resources, and expose VM instances and associated physical resources to other management systems.
[0243] In at least one embodiment, the VNFM 2506 can manage the VNF 2508. In at least one embodiment, the VNF 2508 can be used to perform EPC components / functions. In at least one embodiment, the VNFM 2506 can manage the lifecycle of the VNF 2508 and track the performance, failures, and security of the virtual aspects of the VNF 2508. In at least one embodiment, the EM 2510 can track the performance, failures, and security of the functional aspects of the VNF 2508. In at least one embodiment, tracking data from the VNFM 2506 and the EM 2510 can include, in at least one embodiment, performance measurement (PM) data used by the VIM 2502 or the NFVI 2504. In at least one embodiment, both the VNFM 2506 and the EM 2510 can scale up / down the number of VNFs of the system 2500.
[0244] In at least one embodiment, the NFVO 2512 may coordinate, authorize, release, and occupy resources of the NFVI 2504 in order to provide the requested service (e.g., to execute an EPC function, component, or slice). In at least one embodiment, the NM 2514 may provide an end-user functional package responsible for managing a network, which may include network elements with VNFs, non-virtualized network functions, or both (management of the VNFs may occur via the EM 2510).
[0245] Computer-based systems
[0246] The following figures set forth, but are not limiting of, exemplary computer-based systems that may be used to implement at least one embodiment.
[0247] Fig.26 A processing system 2600 is shown in accordance with at least one embodiment. In at least one embodiment, the system 2600 includes one or more processors 2602 and one or more graphics processors 2608, and may be a single processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 2602 or processor cores 2607. In at least one embodiment, the processing system 2600 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.
[0248] In at least one embodiment, the processing system 2600 may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console for gaming and media consoles. In at least one embodiment, the processing system 2600 is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 2600 may also include a wearable device coupled to or integrated in a wearable device, such as a smart watch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 2600 is a television or set-top box device having one or more processors 2602 and a graphical interface generated by one or more graphics processors 2608.
[0249] In at least one embodiment, one or more processors 2602 each include one or more processor cores 2607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2607 is configured to process a specific instruction set 2609. In at least one embodiment, the instruction set 2609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing through very long instruction words (VLIW). In at least one embodiment, multiple processor cores 2607 can each process a different instruction set 2609, which can include instructions that help emulate other instruction sets. In at least one embodiment, the processor core 2607 can also include other processing devices, such as a digital signal processor (DSP).
[0250] In at least one embodiment, the processor 2602 includes a cache memory (cache) 2604. In at least one embodiment, the processor 2602 can have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared between the various components of the processor 2602. In at least one embodiment, the processor 2602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which can share the logic between the processor cores 2607 using known cache coherence techniques. In at least one embodiment, the processor 2602 additionally includes a register file 2606, and the processor 2602 may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 2606 may include general registers or other registers.
[0251] In at least one embodiment, one or more processors 2602 are coupled to one or more interface buses 2610 to transmit communication signals, such as address, data, or control signals, between the processor 2602 and other components in the system 2600. In at least one embodiment, the interface bus 2610 can be a processor bus in one embodiment, such as a version of a direct media interface (DMI) bus. In at least one embodiment, the interface bus 2610 is not limited to a DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 2602 includes an integrated memory controller 2616 and a platform controller hub 2630. In at least one embodiment, the memory controller 2616 facilitates communication between storage devices and other components of the processing system 2600, while the platform controller hub (PCH) 2630 provides connections to input / output (I / O) devices through a local I / O bus.
[0252] In at least one embodiment, the memory device 2620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have appropriate performance to be used as a processor memory. In at least one embodiment, the memory device 2620 may be used as a system memory of the processing system 2600 to store data 2622 and instructions 2621 for use when one or more processors 2602 execute an application or process. In at least one embodiment, the memory controller 2616 is also coupled to an optional external graphics processor 2612, which may communicate with one or more graphics processors 2608 in the processor 2602 to perform graphics and media operations. In at least one embodiment, the display device 2611 may be connected to the processor 2602. In at least one embodiment, the display device 2611 may include one or more of the internal display devices, such as in a mobile electronic device or portable computer device or an external display device connected via a display interface (e.g., a display port (DisplayPort) or the like). In at least one embodiment, the display device 2611 may include a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.
[0253] In at least one embodiment, the platform controller hub 2630 enables peripheral devices to be connected to the storage device 2620 and the processor 2602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 2646, a network controller 2634, a firmware interface 2628, a wireless transceiver 2626, a touch sensor 2625, a data storage device 2624 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 2624 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 2625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or long-term evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2628 enables communication with the system firmware, and in at least one embodiment, can be a unified extensible firmware interface (UEFI). In at least one embodiment, the network controller 2634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2610. In at least one embodiment, the audio controller 2646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 2600 includes an optional legacy I / O controller 2640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the processing system 2600. In at least one embodiment, the platform controller hub 2630 can also be connected to one or more universal serial bus (USB) controllers 2642, which connect input devices such as a keyboard and mouse 2643 combination, a camera 2644, or other USB input devices.
[0254] In at least one embodiment, instances of the memory controller 2616 and the platform controller hub 2630 may be integrated into a discrete external graphics processor, such as the external graphics processor 2612. In at least one embodiment, the platform controller hub 2630 and / or the memory controller 2616 may be external to one or more processors 2602. In at least one embodiment, the processing system 2600 may include the external memory controller 2616 and the platform controller hub 2630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 2602.
[0255] Fig. 27A computer system 2700 is shown according to at least one embodiment. In at least one embodiment, the computer system 2700 can be a system with interconnected devices and components, a SOC, or some combination. In at least one embodiment, the computer system 2700 is formed by a processor 2702, which can include an execution unit for executing instructions. In at least one embodiment, the computer system 2700 can include, but is not limited to, components, such as the processor 2702, which employs an execution unit including logic to execute algorithms for process data. In at least one embodiment, the computer system 2700 can include a processor, such as the Intel Corporation of Santa Clara, California, available from Processor family, XeonTM, XScaleTM and / or StrongARMTM, Core TM or Nervana TM microprocessor, although other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 2700 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux in at least one embodiment), embedded software, and / or graphical user interfaces may also be used.
[0256] In at least one embodiment, the computer system 2700 can be used in other devices, such as handheld devices and embedded applications. Some of at least one embodiment of handheld devices include cellular phones, Internet Protocol (InternetProtocol) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors ("DSPs"), SoCs, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system that can execute one or more instructions according to at least one embodiment.
[0257] In at least one embodiment, computer system 2700 may include, but is not limited to, processor 2702, which may include, but is not limited to, one or more execution units 2708, which may be configured to execute Compute Unified Device Architecture (“CUDA”) ( Developed by NVIDIA Corporation of Santa Clara, California) program. In at least one embodiment, a CUDA program is at least a portion of a software application written in the CUDA programming language. In at least one embodiment, the computer system 2700 is a single processor desktop or server system. In at least one embodiment, the computer system 2700 may be a multi-processor system. In at least one embodiment, the processor 2702 may include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor that implements an instruction set combination, or any other processor device, in at least one embodiment, such as a digital signal processor. In at least one embodiment, the processor 2702 may be coupled to a processor bus 2710, which may transmit data signals between the processor 2702 and other components in the computer system 2700.
[0258] In at least one embodiment, processor 2702 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 2704. In at least one embodiment, processor 2702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may reside external to processor 2702. In at least one embodiment, processor 2702 may include a combination of internal and external caches. In at least one embodiment, register file 2706 may store different types of data in various registers, including but not limited to integer registers, floating point registers, status registers, and instruction pointer registers.
[0259] In at least one embodiment, an execution unit 2708, including but not limited to logic to perform integer and floating point operations, is also located in the processor 2702. The processor 2702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 2708 may include logic for processing a packed instruction set 2709. In at least one embodiment, by including the packed instruction set 2709 in the instruction set of the general purpose processor 2702, and the associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general purpose processor 2702. In at least one embodiment, many multimedia applications may be executed faster and more efficiently by using the full width of the processor's data bus to perform operations on packed data, which may not require the transfer of smaller units of data on the processor's data bus to perform one or more operations on one data element at a time.
[0260] In at least one embodiment, execution unit 2708 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 2700 may include, but is not limited to, memory 2720. In at least one embodiment, memory 2720 may be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage device. Memory 2720 may store instructions 2719 and / or data 2721 represented by data signals that may be executed by processor 2702.
[0261] In at least one embodiment, the system logic chip can be coupled to the processor bus 2710 and the memory 2720. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 2716, and the processor 2702 can communicate with the MCH 2716 via the processor bus 2710. In at least one embodiment, the MCH 2716 can provide a high-bandwidth memory path 2718 to the memory 2720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 2716 can initiate data signals between the processor 2702, the memory 2720, and other components in the computer system 2700, and bridge data signals between the processor bus 2710, the memory 2720, and the system I / O 2722.
[0262] In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 2716 can be coupled to the memory 2720 via a high bandwidth memory path 2718, and the graphics / video card 2712 can be coupled to the MCH 2716 via an Accelerated Graphics Port ("AGP") interconnect 2714.
[0263] In at least one embodiment, the computer system 2700 can use the system I / O 2722 as a proprietary hub interface bus to couple the MCH 2716 to the I / O controller hub ("ICH") 2730. In at least one embodiment, the ICH 2730 can provide direct connection to certain I / O devices through a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 2720, the chipset, and the processor 2702. Examples can include, but are not limited to, an audio controller 2729, a firmware hub ("Flash BIOS") 2728, a wireless transceiver 2726, a data store 2724, a traditional I / O controller 2723 including a user input 2725 and a keyboard interface, a serial expansion port 2777 (e.g., USB), and a network controller 2734. The data store 2724 can include a hard drive, a floppy drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0264] In at least one embodiment, Fig. 27 A system comprising interconnected hardware devices or "chips" is shown. In at least one embodiment, Fig. 27 An exemplary SoC may be shown. In at least one embodiment, Fig. 27 The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (eg, PCIe), or some combination thereof. In at least one embodiment, one or more components of system 2700 are interconnected using a compute express link (CXL) interconnect.
[0265] Fig.28 A system 2800 is shown in accordance with at least one embodiment. In at least one embodiment, the system 2800 is an electronic device utilizing a processor 2810. In at least one embodiment, the system 2800 can be, in at least one embodiment but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0266] In at least one embodiment, system 2800 may include, but is not limited to, a processor 2810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 2810 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a USB (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Fig.28 A system is shown, which includes interconnected hardware devices or "chips". In at least one embodiment, Fig.28 An exemplary SoC may be shown. In at least one embodiment, Fig.28 The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof.
[0267] In at least one embodiment, Fig.28 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.
[0268] In at least one embodiment, Fig.28 The display 2824, touch screen 2825, touch pad 2830, near field communication unit ("NFC") 2845, sensor hub 2840, thermal sensor 2846, fast chipset ("EC") 2835, trusted platform module ("TPM") 2838, BIOS / firmware / flash ("BIOS, FW Flash") 2822, DSP 2860, solid state disk ("SSD") or hard disk drive ("HDD") 2820, wireless local area network unit ("WLAN") 2850, Bluetooth unit 2852, wireless wide area network unit ("WWAN") 2856, global positioning system (GPS) 2855, camera ("USB 3.0 camera") 2854 (e.g., USB 3.0 camera) or low power double data rate ("LPDDR") memory unit ("LPDDR3") 2815 implemented in at least one embodiment of the LPDDR3 standard may be included. Each of these components may be implemented in any suitable manner.
[0269] In at least one embodiment, other components may be communicatively coupled to the processor 2810 through the components discussed above. In at least one embodiment, an accelerometer 2841, an ambient light sensor (“ALS”) 2842, a compass 2843, and a gyroscope 2844 may be communicatively coupled to the sensor hub 2840. In at least one embodiment, a thermal sensor 2839, a fan 2837, a keyboard 2846, and a touchpad 2830 may be communicatively coupled to the EC 2835. In at least one embodiment, a speaker 2863, an earphone 2864, and a microphone (“mic”) 2865 may be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 2864, which in turn may be communicatively coupled to the DSP 2860. In at least one embodiment, the audio unit 2864 may include, but is not limited to, an audio codec / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 2857 may be communicatively coupled to the WWAN unit 2856. In at least one embodiment, components such as the WLAN unit 2850 and the Bluetooth unit 2852 and the WWAN unit 2856 may be implemented as a next generation form factor (NGFF).
[0270] Fig.29 An exemplary integrated circuit 2900 according to at least one embodiment is shown. In at least one embodiment, the exemplary integrated circuit 2900 is a SoC, which can be manufactured using one or more IP cores. In at least one embodiment, the integrated circuit 2900 includes one or more application processors 2905 (e.g., CPU), at least one graphics processor 2910, and may additionally include an image processor 2915 and / or a video processor 2920, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 2900 includes peripheral or bus logic, which includes a USB controller 2925, a UART controller 2930, a SPI / SDIO controller 2935, and an I2S / I2C controller 2940. In at least one embodiment, the integrated circuit 2900 may include a display device 2945 coupled to one or more of a high-definition multimedia interface (HDMI) controller 2950 and a mobile industry processor interface (MIPI) display interface 2955. In at least one embodiment, storage may be provided by a flash subsystem 2960, including flash memory and a flash controller. In at least one embodiment, a memory interface may be provided via a memory controller 2965 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 2970 .
[0271] Fig.30A computing system 3000 is shown in accordance with at least one embodiment. In at least one embodiment, the computing system 3000 includes a processing subsystem 3001 having one or more processors 3002 and a system memory 3004 communicating via an interconnect path that may include a memory hub 3005. In at least one embodiment, the memory hub 3005 may be a separate component within a chipset component or may be integrated within one or more processors 3002. In at least one embodiment, the memory hub 3005 is coupled to an I / O subsystem 3011 via a communication link 3006. In at least one embodiment, the I / O subsystem 3011 includes an I / O hub 3007 that may enable the computing system 3000 to receive input from one or more input devices 3008. In at least one embodiment, the I / O hub 3007 may enable a display controller, included in one or more processors 3002, to provide output to one or more display devices 3010A. In at least one embodiment, the one or more display devices 3010A coupled to the I / O hub 3007 may include local, internal, or embedded display devices.
[0272] In at least one embodiment, the processing subsystem 3001 includes one or more parallel processors 3012 coupled to a memory hub 3005 via a bus or other communication link 3013. In at least one embodiment, the communication link 3013 can be one of many standard-based communication link technologies or protocols, such as but not limited to PCIe, or can be a communication interface or communication structure for a vendor. In at least one embodiment, one or more parallel processors 3012 form a parallel or vector processing system in a computational concentration, which can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, one or more parallel processors 3012 form a graphics processing subsystem that can output pixels to one of one or more display devices 3010A coupled via an I / O hub 3007. In at least one embodiment, one or more parallel processors 3012 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 3010B.
[0273] In at least one embodiment, a system storage unit 3014 can be connected to the I / O hub 3007 to provide a storage mechanism for the computing system 3000. In at least one embodiment, an I / O switch 3016 can be used to provide an interface mechanism to enable connections between the I / O hub 3007 and other components, such as a network adapter 3018 and / or a wireless network adapter 3019 that can be integrated into the platform, as well as various other devices that can be added through one or more additional devices 3020. In at least one embodiment, the network adapter 3018 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 3019 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices including one or more radios.
[0274] In at least one embodiment, computing system 3000 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and / or variations thereof, which may also be connected to I / O hub 3007. Fig.30 The communication paths that interconnect the various components in the system can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols (e.g., NVLink high-speed interconnect or interconnect protocol).
[0275] In at least one embodiment, one or more parallel processors 3012 include circuits optimized for graphics and video processing (including video output circuits in at least one embodiment) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 3012 include circuits optimized for general processing. In at least one embodiment, the components of the computing system 3000 can be integrated with one or more other system elements on a single integrated circuit. In at least one embodiment, one or more parallel processors 3012, memory hub 3005, processor 3002, and I / O hub 3007 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of the computing system 3000 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 3000 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules into a modular computing system. In at least one embodiment, the I / O subsystem 3011 and the display device 3010B are omitted from the computing system 3000.
[0276] Processing system
[0277] The following figures illustrate, but are not limited to, exemplary processing systems that may be used to implement at least one embodiment.
[0278] Fig.31 An accelerated processing unit ("APU") 3100 is shown in accordance with at least one embodiment. In at least one embodiment, the APU 3100 was developed by AMD, Inc. of Santa Clara, California. In at least one embodiment, the APU 3100 may be configured to execute application programs, such as CUDA programs. In at least one embodiment, the APU 3100 includes, but is not limited to, a core complex 3110, a graphics complex 3140, a fabric 3160, an I / O interface 3170, a memory controller 3180, a display controller 3192, and a multimedia engine 3194. In at least one embodiment, the APU 3100 may include, but is not limited to, any combination of any number of core complexes 3110, any number of graphics complexes 3140, any number of display controllers 3192, and any number of multimedia engines 3194. For purposes of illustration, multiple instances of similar objects are represented herein by reference numerals, where the reference numeral identifies the object and the number in parentheses identifies the desired instance.
[0279] In at least one embodiment, the core complex 3110 is a CPU, the graphics complex 3140 is a GPU, and the APU 3100 is a processing unit that is not limited to integrating 3110 and 3140 onto a single chip. In at least one embodiment, some tasks may be assigned to the core complex 3110, while other tasks may be assigned to the graphics complex 3140. In at least one embodiment, the core complex 3110 is configured to execute main control software associated with the APU 3100, such as an operating system. In at least one embodiment, the core complex 3110 is the main processor of the APU 3100, which controls and coordinates the operations of the other processors. In at least one embodiment, the core complex 3110 issues commands that control the operations of the graphics complex 3140. In at least one embodiment, the core complex 3110 may be configured to execute host executable code derived from CUDA source code, and the graphics complex 3140 may be configured to execute device executable code derived from CUDA source code.
[0280] In at least one embodiment, core complex 3110 includes, but is not limited to, cores 3120(1)-3120(4) and L3 cache 3130. In at least one embodiment, core complex 3110 may include, but is not limited to, any number of cores 3120 and any combination of any number and type of caches. In at least one embodiment, cores 3120 are configured to execute instructions of a particular instruction set architecture ("ISA"). In at least one embodiment, each core 3120 is a CPU core.
[0281] In at least one embodiment, each core 3120 includes, but is not limited to, a fetch / decode unit 3122, an integer execution engine 3124, a floating point execution engine 3126, and an L2 cache 3128. In at least one embodiment, the fetch / decode unit 3122 fetches instructions, decodes these instructions, generates micro-operations, and dispatches separate micro-instructions to the integer execution engine 3124 and the floating point execution engine 3126. In at least one embodiment, the fetch / decode unit 3122 can dispatch one micro-instruction to the integer execution engine 3124 and another micro-instruction to the floating point execution engine 3126 at the same time. In at least one embodiment, the integer execution engine 3124 performs, but is not limited to, integer and memory operations. In at least one embodiment, the floating point engine 3126 performs, but is not limited to, floating point and vector operations. In at least one embodiment, the fetch-decode unit 3122 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 3124 and the floating point execution engine 3126.
[0282] In at least one embodiment, each core 3120(i) may access an L2 cache 3128(i) included in the core 3120(i), where i is an integer representing a specific instance of the core 3120. In at least one embodiment, each core 3120 included in the core complex 3110(j) is connected to other cores 3120 included in the core complex 3110(j) via an L3 cache 3130(j) included in the core complex 3110(j), where j is an integer representing a specific instance of the core complex 3110. In at least one embodiment, a core 3120 included in a core complex 3110(j) may access all L3 caches 3130(j) included in the core complex 3110(j), where j is an integer representing a specific instance of the core complex 3110. In at least one embodiment, the L3 cache 3130 may include, but is not limited to, any number of slices.
[0283] In at least one embodiment, graphics complex 3140 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 3140 is configured to perform graphics pipeline operations, such as drawing commands, pixel operations, geometry calculations, and other operations associated with rendering an image to a display. In at least one embodiment, graphics complex 3140 is configured to perform operations that are not related to graphics. In at least one embodiment, graphics complex 3140 is configured to perform both graphics-related operations and graphics-independent operations.
[0284] In at least one embodiment, graphics complex 3140 includes, but is not limited to, any number of compute units 3150 and L2 cache 3142. In at least one embodiment, compute units 3150 share L2 cache 3142. In at least one embodiment, L2 cache 3142 is partitioned. In at least one embodiment, graphics complex 3140 includes, but is not limited to, any number of compute units 3150 and any number (including zero) and type of caches. In at least one embodiment, graphics complex 3140 includes, but is not limited to, any number of dedicated graphics hardware.
[0285] In at least one embodiment, each computing unit 3150 includes, but is not limited to, any number of SIMD units 3152 and shared memory 3154. In at least one embodiment, each SIMD unit 3152 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, each computing unit 3150 can execute any number of thread blocks, but each thread block is executed on a single computing unit 3150. In at least one embodiment, thread blocks include, but are not limited to, any number of execution threads. In at least one embodiment, a work group is a thread block. In at least one embodiment, each SIMD unit 3152 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), wherein each thread in a warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, a prediction can be used to disable one or more threads in a warp. In at least one embodiment, a channel is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block can be synchronized together and communicated via a shared memory 3154.
[0286] In at least one embodiment, fabric 3160 is a system interconnect that facilitates data and control transfers across core complex 3110, graphics complex 3140, I / O interface 3170, memory controller 3180, display controller 3192, and multimedia engine 3194. In at least one embodiment, APU 3100 may include, but is not limited to, any number and type of system interconnects in addition to or in lieu of fabric 3160 that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to APU 3100. In at least one embodiment, I / O interface 3170 represents any number and type of I / O interfaces (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various types of peripherals are coupled to I / O interface 3170. In at least one embodiment, peripheral devices coupled to the I / O interface 3170 may include, but are not limited to, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.
[0287] In at least one embodiment, the display controller AMD92 displays images on one or more display devices, such as a liquid crystal display (LCD) device. In at least one embodiment, the multimedia engine 240 includes, but is not limited to, any number and type of multimedia-related circuits, such as a video decoder, a video encoder, an image signal processor, etc. In at least one embodiment, the memory controller 3180 facilitates data transfer between the APU 3100 and the unified system memory 3190. In at least one embodiment, the core complex 3110 and the graphics complex 3140 share the unified system memory 3190.
[0288] In at least one embodiment, the APU 3100 implements a memory subsystem including, but not limited to, any number and type of memory controllers 3180 and memory devices (e.g., shared memory 3154) that may be dedicated to a component or shared among multiple components. Components. In at least one embodiment, the APU 3100 implements a cache subsystem including, but not limited to, one or more cache memories (e.g., L2 cache 2728, L3 cache 3130, and L2 cache 3142), each of which may be component-private or shared among any number of components (e.g., core 3120, core complex 3110, SIMD unit 3152, compute unit 3150, and graphics complex 3140).
[0289] Fig.32A CPU 3200 according to at least one embodiment is shown. In at least one embodiment, the CPU 3200 is developed by AMD, Inc. of Santa Clara, California. In at least one embodiment, the CPU 3200 may be configured to execute an application program. In at least one embodiment, the CPU 3200 is configured to execute a main control software, such as an operating system. In at least one embodiment, the CPU 3200 issues commands to control the operation of an external GPU (not shown). In at least one embodiment, the CPU 3200 may be configured to execute a host executable code derived from a CUDA source code, and the external GPU may be configured to execute a device executable code derived from such a CUDA source code. In at least one embodiment, the CPU 3200 includes, but is not limited to, any number of core complexes 3210, structures 3260, I / O interfaces 3270, and memory controllers 3280.
[0290] In at least one embodiment, core complex 3210 includes, but is not limited to, cores 3220(1)-3220(4) and L3 cache 3230. In at least one embodiment, core complex 3210 may include, but is not limited to, any number of cores 3220 and any combination of any number and type of caches. In at least one embodiment, core 3220 is configured to execute instructions of a specific ISA. In at least one embodiment, each core 3220 is a CPU core.
[0291] In at least one embodiment, each core 3220 includes, but is not limited to, a fetch / decode unit 3222, an integer execution engine 3224, a floating point execution engine 3226, and an L2 cache 3228. In at least one embodiment, the fetch / decode unit 3222 fetches instructions, decodes these instructions, generates micro-operations, and dispatches separate micro-instructions to the integer execution engine 3224 and the floating point execution engine 3226. In at least one embodiment, the fetch / decode unit 3222 can dispatch one micro-instruction to the integer execution engine 3224 and another micro-instruction to the floating point execution engine 3226 at the same time. In at least one embodiment, the integer execution engine 3224 performs, but is not limited to, integer and memory operations. In at least one embodiment, the floating point engine 3226 performs, but is not limited to, floating point and vector operations. In at least one embodiment, the fetch-decode unit 3222 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 3224 and the floating point execution engine 3226.
[0292] In at least one embodiment, each core 3220(i) may access an L2 cache 3228(i) included in the core 3220(i), where i is an integer representing a specific instance of the core 3220. In at least one embodiment, each core 3220 included in a core complex 3210(j) is connected to other cores 3220 in the core complex 3210(j) via an L3 cache 3230(j) included in the core complex 3210(j), where j is an integer representing a specific instance of the core complex 3210. In at least one embodiment, a core 3220 included in a core complex 3210(j) may access all L3 caches 3230(j) included in the core complex 3210(j), where j is an integer representing a specific instance of the core complex 3210. In at least one embodiment, the L3 cache 3230 may include, but is not limited to, any number of slices.
[0293] In at least one embodiment, fabric 3260 is a system interconnect that facilitates data and control transfers across core complexes 3210(1)-3210(N) (where N is an integer greater than zero), I / O interface 3270, and memory controller 3280. In at least one embodiment, CPU 3200 may include, in addition to or in lieu of fabric 3260, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to CPU 3200. In at least one embodiment, I / O interface 3270 represents any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripherals are coupled to I / O interface 3270. In at least one embodiment, peripherals coupled to I / O interface 3270 may include, but are not limited to, a display, keyboard, mouse, printer, scanner, joystick or other type of game controller, media recording device, external storage device, network interface card, etc.
[0294] In at least one embodiment, memory controller 3280 facilitates data transfer between CPU 3200 and system memory 3290. In at least one embodiment, core complex 3210 and graphics complex 3240 share system memory 3290. In at least one embodiment, CPU 3200 implements a memory subsystem, which includes, but is not limited to, any number and type of memory controllers 3280 and memory devices that can be dedicated to one component or shared between multiple components. In at least one embodiment, CPU 3200 implements a cache subsystem, which includes, but is not limited to, one or more cache memories (e.g., L2 cache 3228 and L3 cache 3230), each of which can be private to a component or shared between any number of components (e.g., core 3220 and core complex 3210).
[0295] Fig.33 An exemplary accelerator integrated slice 3390 according to at least one embodiment is shown. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, environment management, and interrupt management services on behalf of multiple graphics processing engines in multiple graphics acceleration modules. The graphics processing engines may each include a separate GPU. Optionally, the graphics processing engine may include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module may be a GPU with multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a general package, line card, or chip.
[0296] The application effective address space 3382 within the system memory 3314 stores process elements 3383. In one embodiment, the process element 3383 is stored in response to a GPU call 3381 from an application 3380 executing on the processor 3307. The process element 3383 contains the processing state of the corresponding application 3380. The work descriptor (WD) 3384 contained in the process element 3383 can be a single job requested by the application or may contain a pointer to a job queue. In at least one embodiment, the WD 3384 is a pointer to a job request queue in the application effective address space 3382.
[0297] Graphics acceleration module 3346 and / or each graphics processing engine can be shared by all or part of the processes in the system. In at least one embodiment, an infrastructure for establishing a processing state and sending WD 3384 to graphics acceleration module 3346 to start a job in a virtualized environment can be included.
[0298] In at least one embodiment, a dedicated process programming model is implemented for. In this model, a single process owns a graphics acceleration module 3346 or an individual graphics processing engine. Since the graphics acceleration module 3346 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and the operating system initializes the accelerator integrated circuit for the owned partition when the graphics acceleration module 3346 is allocated.
[0299] In operation, the WD acquisition unit 3391 in the accelerator integrated slice 3390 acquires the next WD 3384, which includes an indication of the work to be completed by one or more graphics processing engines of the graphics acceleration module 3346. Data from the WD 3384 can be stored in registers 3345 for use by a memory management unit (MMU) 3339, an interrupt management circuit 3347, and / or an environment management circuit 3348, as shown. At least one embodiment of the MMU 3339 includes a segment / page roaming circuit for accessing a segment / page table 3386 within an OS virtual address space 3385. The interrupt management circuit 3347 can process an interrupt event (INT) 3392 received from the graphics acceleration module 3346. When performing a graph operation, an effective address 3393 generated by the graphics processing engine is converted to an actual address by the MMU 3339.
[0300] In one embodiment, the same register set 3345 is replicated for each graphics processing engine and / or graphics acceleration module 3346 and can be initialized by a system hypervisor or operating system. Each of these replicated registers can be included in an accelerator integrated slice 3390. An exemplary register that can be initialized by a hypervisor is shown in Table 1.
[0301] Table 1 – Registers initialized by the hypervisor
[0302]
[0303]
[0304] Example registers that may be initialized by the operating system are shown in Table 2.
[0305] Table 2 Operating system initialization registers
[0306]
[0307] In one embodiment, each WD 3384 is specific to a particular graphics acceleration module 3346 and / or a particular graphics processing engine. It contains all the information needed for the graphics processing engine to do its job or work, or it can be a pointer to a memory location where the application has established a command queue for the work to be done.
[0308] Figures 34A-34B An exemplary graphics processor according to at least one embodiment of the present invention is shown. In at least one embodiment, any exemplary graphics processor can be manufactured using one or more IP cores. In addition to the illustrations, other logic and circuits can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the exemplary graphics processor is used in a SoC.
[0309] Fig.34A An exemplary graphics processor 3410 of a SoC integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Fig.34B An additional exemplary graphics processor 3440 of a SoC integrated circuit is shown, which may be manufactured using one or more IP cores, according to at least one embodiment. In at least one embodiment, Fig.34A The graphics processor 3410 is a low power graphics processor core. In at least one embodiment, Fig.34B The graphics processor 3440 of FIG. 5 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 3410, 3440 may be a variation of the graphics processor 510 of FIG. 5.
[0310] In at least one embodiment, the graphics processor 3410 includes a vertex processor 3405 and one or more fragment processors 3415A-3415N (e.g., 3415A, 3415B, 3415C, 3415D to 3415N-1 and 3415N). In at least one embodiment, the graphics processor 3410 can execute different shader programs via separate logic, so that the vertex processor 3405 is optimized to perform operations for the vertex shader program, while one or more fragment processors 3415A-3415N perform fragment (e.g., pixel) shading operations for fragments or pixels or shader programs. In at least one embodiment, the vertex processor 3405 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 3415A-3415N use the primitives and vertex data generated by the vertex processor 3405 to generate a frame buffer displayed on a display device. In at least one embodiment, the fragment processors 3415A-3415N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.
[0311] In at least one embodiment, the graphics processor 3410 additionally includes one or more MMUs 3420A-3420B, caches 3425A-3425B, and circuit interconnects 3430A-3430B. In at least one embodiment, the one or more MMUs 3420A-3420B provide a mapping of virtual to physical addresses for the graphics processor 3410, including for the vertex processor 3405 and / or the fragment processor 3415A-3415N, which can reference vertices or image / texture data stored in memory, in addition to the vertex or image / texture data stored in the one or more caches 3425A-3425B. In at least one embodiment, the one or more MMUs 3420A-3420B can be synchronized with other MMUs within the system, including one or more MMUs associated with one or more application processors 505, image processor 515, and / or video processor 520 of FIG. 5, so that each processor 505-520 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 3430A-3430B enable graphics processor 3410 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.
[0312] In at least one embodiment, graphics processor 3440 includes Fig.34AOne or more MMUs 3420A-3420B, caches 3425A-3425B, and circuit interconnects 3430A-3430B of a graphics processor 3410. In at least one embodiment, the graphics processor 3440 includes one or more shader cores 3455A-3455N (e.g., 3455A, 3455B, 3455C, 3455D, 3455E, 3455F, to 3455N-1 and 3455N) that provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the plurality of shader cores may vary. In at least one embodiment, the graphics processor 3440 includes an inter-core task manager 3445 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 3455A-3455N and a tiling unit 3458 to accelerate tile operations for tile-based rendering, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize use of internal caches.
[0313] Fig.35A 3500 according to at least one embodiment. In at least one embodiment, the graphics core 3500 may include Fig.24 In at least one embodiment, the graphics core 3500 may be Fig.34B 3455A-3455N. In at least one embodiment, the graphics core 3500 includes a shared instruction cache 3502, a texture unit 3518, and a cache / shared memory 3520, which are common to the execution resources within the graphics core 3500. In at least one embodiment, the graphics core 3500 may include multiple slices 3501A-3501N or partitions of each core, and the graphics processor may include multiple instances of the graphics core 3500. The slices 3501A-3501N may include support logic including a local instruction cache 3504A-3504N, a thread scheduler 3506A-3506N, a thread dispatcher 3508A-3508N, and a set of registers 3510A-3510N. In at least one embodiment, the slices 3501A-3501N may include a set of additional function units (AFUs) 3512A-3512N, floating point units (FPUs) 3514A-3514N, integer arithmetic logic units (ALUs) 3516A-3516N, address calculation units (ACUs) 3513A-3513N, double precision floating point units (DPFPUs) 3515A-3515N, and matrix processing units (MPUs) 3517A-3517N.
[0314] In one embodiment, the FPU 3514A-3514N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 3515A-3515N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 3516A-3516N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 3517A-3517N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 3517A-3517N can perform various matrix operations to accelerate CUDA programs, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFU 3512A-3512N can perform additional logical operations that are not supported by the floating-point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).
[0315] Fig.35B A general purpose graphics processing unit (GPGPU) 3530 in at least one embodiment is shown. In at least one embodiment, GPGPU 3530 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, GPGPU 3530 can be configured to enable highly parallel computing operations to be performed by a GPU array. In at least one embodiment, GPGPU 3530 can be directly linked to other instances of GPGPU 3530 to create a multi-GPU cluster to improve the execution time for CUDA programs. In at least one embodiment, GPGPU 3530 includes a host interface 3532 to enable connection with a host processor. In at least one embodiment, host interface 3532 is a PCIe interface. In at least one embodiment, host interface 3532 can be a manufacturer-specific communication interface or communication structure. In at least one embodiment, GPGPU 3530 receives commands from a host processor and dispatches execution threads associated with those commands to a group of computing clusters 3536A-3536H using a global scheduler 3534. In at least one embodiment, the computing clusters 3536A-3536H share a cache memory 3538. In at least one embodiment, the cache memory 3538 can be used as a high-level cache for the cache memories within the computing clusters 3536A-3536H.
[0316] In at least one embodiment, GPGPU 3530 includes memory 3544A-3544B coupled to compute cluster 3536A-3536H via a set of memory controllers 3542A-3542B. In at least one embodiment, memory 3544A-3544B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.
[0317] In at least one embodiment, computing clusters 3536A-3536H each include a set of graphics cores, such as Fig.35A The graphics core 3500, which may include multiple types of integer and floating point logic units, may perform computational operations at various precisions, including computations suitable for use with CUDA programs. In at least one embodiment, at least a subset of the floating point units in each compute cluster 3536A-3536H may be configured to perform 16-bit or 32-bit floating point operations, while a different subset of the floating point units may be configured to perform 64-bit floating point operations.
[0318] In at least one embodiment, multiple instances of GPGPU 3530 may be configured to operate as a computing cluster. In at least one embodiment, computing clusters 3536A-3536H may implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of GPGPU 3530 communicate through host interface 3532. In at least one embodiment, GPGPU 3530 includes an I / O hub 3539 that couples GPGPU 3530 with GPU link 3540, enabling direct connection to other instances of GPGPU 3530. In at least one embodiment, GPU link 3540 is coupled to a dedicated GPU to GPU bridge that enables communication and synchronization between multiple instances of GPGPU 3530. In at least one embodiment, GPU link 3540 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 3530 are located in separate data processing systems and communicate via network devices accessible via host interface 3532. In at least one embodiment, GPU link 3540 may be configured to enable connection to a host processor, in addition to or in lieu of host interface 3532. In at least one embodiment, GPGPU 3530 may be configured to execute CUDA programs.
[0319] Fig.36AA parallel processor 3600 is shown in accordance with at least one embodiment. In at least one embodiment, various components of parallel processor 3600 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or an FPGA.
[0320] In at least one embodiment, parallel processor 3600 includes parallel processing unit 3602. In at least one embodiment, parallel processing unit 3602 includes I / O unit 3604, which enables communication with other devices, including other instances of parallel processing unit 3602. In at least one embodiment, I / O unit 3604 can be directly connected to other devices. In at least one embodiment, I / O unit 3604 connects with other devices by using a hub or switch interface (e.g., memory hub 605). In at least one embodiment, the connection between memory hub 605 and I / O unit 3604 forms a communication link. In at least one embodiment, I / O unit 3604 is connected to host interface 3606 and memory crossbar switch 3616, wherein host interface 3606 receives commands for performing processing operations and memory crossbar switch 3616 receives commands for performing memory operations.
[0321] In at least one embodiment, when the host interface 3606 receives the command buffer via the I / O unit 3604, the host interface 3606 can direct work operations to execute those commands to the front end 3608. In at least one embodiment, the front end 3608 is coupled with a scheduler 3610, which is configured to distribute commands or other work items to the processing array 3612. In at least one embodiment, the scheduler 3610 ensures that the processing array 3612 is properly configured and in a valid state before assigning tasks to the processing array 3612 in the processing array 3612. In at least one embodiment, the scheduler 3610 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 3610 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularity, thereby achieving fast preemption and context switching of threads executing on the processing array 3612. In at least one embodiment, the host software can prove the workload for scheduling on the processing array 3612 through one of a plurality of graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across the processing array 3612 by scheduler 3610 logic within a microcontroller that includes scheduler 3610.
[0322] In at least one embodiment, the processing array 3612 may include up to "N" processing clusters (e.g., cluster 3614A, cluster 3614B, through cluster 3614N). In at least one embodiment, each cluster 3614A-3614N of the processing array 3612 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 3610 may allocate work to the clusters 3614A-3614N of the processing array 3612 using various scheduling and / or work allocation algorithms, which may vary depending on the workload generated by each program or type of computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 3610, or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by the processing array 3612. In at least one embodiment, different clusters 3614A-3614N of the processing array 3612 may be allocated to process different types of programs or to perform different types of computations.
[0323] In at least one embodiment, the processing array 3612 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing array 3612 is configured to perform general-purpose parallel computing operations. In at least one embodiment, the processing array 3612 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.
[0324] In at least one embodiment, the processing array 3612 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing array 3612 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing array 3612 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 3602 may transfer data from the system memory via the I / O unit 3604 for processing. In at least one embodiment, during processing, the transferred data may be stored to an on-chip memory (e.g., parallel processor memory 3622) during processing and then written back to the system memory.
[0325] In at least one embodiment, when parallel processing unit 3602 is used to perform graph processing, scheduler 3610 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 3614A-3614N of processing array 3612. In at least one embodiment, portions of processing array 3612 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of clusters 3614A-3614N can be stored in a buffer to allow the intermediate data to be transferred between clusters 3614A-3614N for further processing.
[0326] In at least one embodiment, the processing array 3612 can receive processing tasks to be performed via the scheduler 3610, which receives commands defining the processing tasks from the front end 3608. In at least one embodiment, the processing tasks can include an index of the data to be processed, which can include surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 3610 can be configured to obtain the index corresponding to the task, or can receive the index from the front end 3608. In at least one embodiment, the front end 3608 can be configured to ensure that the processing array 3612 is configured to a valid state before starting the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.).
[0327] In at least one embodiment, each of the one or more instances of parallel processing unit 3602 can be coupled to parallel processor memory 3622. In at least one embodiment, parallel processor memory 3622 can be accessed via memory crossbar switch 3616, which can receive memory requests from processing array 3612 and I / O unit 3604. In at least one embodiment, memory crossbar switch 3616 can access parallel processor memory 3622 via memory interface 3618. In at least one embodiment, memory interface 3618 can include multiple partition units (e.g., partition unit 3620A, partition unit 3620B to partition unit 3620N), which can each be coupled to a portion of parallel processor memory 3622 (e.g., memory units). In at least one embodiment, the plurality of partition units 3620A-3620N are configured to be equal to the number of memory cells, such that the first partition unit 3620A has a corresponding first memory cell 3624A, the second partition unit 3620B has a corresponding memory cell 3624B, and the Nth partition unit 3620N has a corresponding Nth memory cell 3624N. In at least one embodiment, the number of partition units 3620A-3620N may not be equal to the number of memory devices.
[0328] In at least one embodiment, memory units 3624A-3624N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 3624A-3624N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory units 3624A-3624N, allowing partition units 3620A-3620N to write portions of each rendering target in parallel to efficiently use the available bandwidth of parallel processor memory 3622. In at least one embodiment, local instances of parallel processor memory 3622 may be excluded to facilitate a unified memory design that utilizes system memory in combination with local cache memory.
[0329] In at least one embodiment, any of the clusters 3614A-3614N of the processing array 3612 can process data to be written to any memory unit 3624A-3624N within the parallel processor memory 3622. In at least one embodiment, the memory crossbar 3616 can be configured to transmit the output of each cluster 3614A-3614N to any partition unit 3620A-3620N or another cluster 3614A-3614N, and the cluster 3614A-3614N can perform other processing operations on the output. In at least one embodiment, each cluster 3614A-3614N can communicate with the memory interface 3618 through the memory crossbar 3616 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 3616 has connections to memory interface 3618 to communicate with I / O unit 3604, and connections to local instances of parallel processor memory 3622, thereby enabling processing units within different processing clusters 3614A-3614N to communicate with system memory or other memory that is not local to parallel processing unit 3602. In at least one embodiment, memory crossbar switch 3616 may use virtual channels to separate traffic flows between clusters 3614A-3614N and partition units 3620A-3620N.
[0330] In at least one embodiment, multiple instances of parallel processing unit 3602 may be provided on a single plug-in card, or multiple plug-in cards may be interconnected. In at least one embodiment, different instances of parallel processing unit 3602 may be configured to interoperate, even if different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 3602 may include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 3602 or parallel processor 3600 may be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0331] Fig.36B 3694 is shown in accordance with at least one embodiment. In at least one embodiment, processing cluster 3694 is included within a parallel processing unit. In at least one embodiment, processing cluster 3694 is Fig.36AIn at least one embodiment, the processing cluster 3694 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a specific program executed on a specific set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of generally synchronized threads, which uses a common instruction unit that is configured to issue instructions to a group of processing engines within each processing cluster 3694.
[0332] In at least one embodiment, the operation of the processing cluster 3694 can be controlled by a pipeline manager 3632 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 3632 Fig.36A The scheduler 3610 receives instructions and manages the execution of these instructions through the graphics multiprocessor 3634 and / or the texture unit 3636. In at least one embodiment, the graphics multiprocessor 3634 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of different architectures may be included in the processing cluster 3694. In at least one embodiment, one or more instances of the graphics multiprocessor 3634 may be included in the processing cluster 3694. In at least one embodiment, the graphics multiprocessor 3634 may process data, and the data crossbar 3640 may be used to distribute the processed data to one of a plurality of possible destinations (including other shader units). In at least one embodiment, the pipeline manager 3632 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossbar 3640.
[0333] In at least one embodiment, each graphics multiprocessor 3634 within a processing cluster 3694 may include the same set of function execution logic (e.g., arithmetic logic unit, load store unit (LSU), etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions are completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating point arithmetic, comparison operations, Boolean operations, shifts, and calculation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.
[0334] In at least one embodiment, the instructions transmitted to the processing cluster 3694 constitute threads. In at least one embodiment, a group of threads executed across a group of parallel processing engines is a thread group. In at least one embodiment, the thread group executes the program on different input data. In at least one embodiment, each thread within a thread group may be assigned to a different processing engine within the graphics multiprocessor 3634. In at least one embodiment, a thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 3634. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the cycle that is processing the thread group. In at least one embodiment, the thread group may also include more threads than the number of processing engines within the graphics multiprocessor 3634. In at least one embodiment, when the thread group includes more threads than the number of processing engines within the graphics multiprocessor 3634, processing may be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups may be executed simultaneously on the graphics multiprocessor 3634.
[0335] In at least one embodiment, graphics multiprocessor 3634 includes internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 3634 can abandon the internal cache and use cache memory (e.g., L1 cache 3648) within processing cluster 3694. In at least one embodiment, each graphics multiprocessor 3634 can also access partition units (e.g., Fig.36A L2 caches within the partition units 3620A-3620N) of the graphics multiprocessor 3634 are shared between all processing clusters 3694 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 3634 can also access off-chip global memory, which can include one or more of the local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 3602 can be used as global memory. In at least one embodiment, the processing cluster 3694 includes multiple instances of the graphics multiprocessor 3634, which can share common instructions and data that can be stored in the L1 cache 3648.
[0336] In at least one embodiment, each processing cluster 3694 may include an MMU 3645 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 3645 may reside in Fig.36A3618. In at least one embodiment, the MMU 3645 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles (talk more about tiles) and optionally to cache line indices. In at least one embodiment, the MMU 3645 may include an address translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 3634 or L1 cache 3648 or processing cluster 3694. In at least one embodiment, the physical addresses are processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or miss.
[0337] In at least one embodiment, the processing clusters 3694 can be configured such that each graphics multiprocessor 3634 is coupled to a texture unit 3636 to perform texture mapping operations, which may involve, for example, determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 3634, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 3634 outputs processed tasks to a data crossbar 3640 to provide the processed tasks to another processing cluster 3694 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 3616. In at least one embodiment, a pre-raster operations unit (preROP) 3642 is configured to receive data from the graphics multiprocessor 3634, direct the data to a ROP unit, which may communicate with a partition unit (e.g., Fig.36A In at least one embodiment, the PreROP 3642 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.
[0338] Fig.36C A graphics multiprocessor 3696 is shown according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 3696 is Fig.36B3634. In at least one embodiment, the graphics multiprocessor 3696 is coupled to the pipeline manager 3632 of the processing cluster 3694. In at least one embodiment, the graphics multiprocessor 3696 has an execution pipeline that includes, but is not limited to, an instruction cache 3652, an instruction unit 3654, an address mapping unit 3656, a register file 3658, one or more GPGPU cores 3662, and one or more LSUs 3666. The GPGPU cores 3662 and the LSUs 3666 are coupled to the cache memory 3672 and the shared memory 3670 via a memory and cache interconnect 3668.
[0339] In at least one embodiment, the instruction cache 3652 receives a stream of instructions to be executed from the pipeline manager 3632. In at least one embodiment, the instructions are cached in the instruction cache 3652 and dispatched for execution by the instruction unit 3654. In one embodiment, the instruction unit 3654 can dispatch instructions as thread groups (e.g., warps), assigning each thread of the thread group to a different execution unit within the GPGPU core 3662. In at least one embodiment, the instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 3656 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by the LSU 3666.
[0340] In at least one embodiment, register file 3658 provides a set of registers for the functional units of graphics multiprocessor 3696. In at least one embodiment, register file 3658 provides temporary storage for operands for data paths of functional units (e.g., GPGPU core 3662, LSU 3666) connected to graphics multiprocessor 3696. In at least one embodiment, register file 3658 is divided between each functional unit such that a dedicated portion of register file 3658 is allocated to each functional unit. In at least one embodiment, register file 3658 is divided between different thread groups being executed by graphics multiprocessor 3696.
[0341] In at least one embodiment, the GPGPU cores 3662 may each include an FPU and / or ALU for executing instructions of the graphics multiprocessor 3696. The GPGPU cores 3662 may be similar in architecture or the architecture may be different. In at least one embodiment, the first portion of the GPGPU core 3662 includes a single-precision FPU and an integer ALU, while the second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-3608 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 3696 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores 3662 may also include fixed or special-function logic.
[0342] In at least one embodiment, the GPGPU core 3662 includes SIMD logic capable of executing a single instruction to multiple sets of data. In at least one embodiment, the GPGPU core 3662 can physically execute SIMD4, SIMD8 and SIMD16 instructions, and logically execute SIMD1, SIMD2 and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. In at least one embodiment, eight SIMT threads that perform the same or similar operations can be executed in parallel by a single SIMD8 logic unit.
[0343] In at least one embodiment, the memory and cache interconnect 3668 is an interconnect network that connects each functional unit of the graphics multiprocessor 3696 to the register file 3658 and the shared memory 3670. In at least one embodiment, the memory and cache interconnect 3668 is a crossbar interconnect that allows the LSU 3666 to implement load and store operations between the shared memory 3670 and the register file 3658. In at least one embodiment, the register file 3658 can operate at the same frequency as the GPGPU core 3662, so that the latency of data transfer between the GPGPU core 3662 and the register file 3658 is very low. In at least one embodiment, the shared memory 3670 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 3696. In at least one embodiment, in at least one embodiment, the cache memory 3672 can be used as a data cache to cache texture data communicated between the functional units and the texture unit 3636. In at least one embodiment, the shared memory 3670 can also be used as a program managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 3672, threads executing on GPGPU core 3662 may programmatically store data in shared memory.
[0344] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated with the core on the same package or chip and communicatively coupled to the core via an internal processor bus / interconnect (i.e., inside the package or chip). In at least one embodiment, regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in the WD. In at least one embodiment, the GPU then uses dedicated circuits / logic to efficiently process these commands / instructions.
[0345] General Computing
[0346] The following figures illustrate, but are not limited to, exemplary software configurations for implementing at least one embodiment in general-purpose computing.
[0347] Fig.37A software stack for a programming platform according to at least one embodiment is shown. In at least one embodiment, a programming platform is a platform for utilizing hardware on a computing system to accelerate computing tasks. In at least one embodiment, a software developer can access the programming platform through libraries, compiler instructions, and / or extensions to a programming language. In at least one embodiment, the programming platform can be, but is not limited to, CUDA, Radeon Open Compute Platform ("ROCm"), OpenCL (OpenCL developed by Khronosgroup), TM ), SYCL, or Intel One API.
[0348] In at least one embodiment, the software stack 3700 of the programming platform provides an execution environment for the application 3701. In at least one embodiment, the application 3701 may include any computer software that can be launched on the software stack 3700. In at least one embodiment, the application 3701 may include, but is not limited to, artificial intelligence ("AI") / machine learning ("ML") applications, high performance computing ("HPC") applications, virtual desktop infrastructure ("VDI"), or data center workloads.
[0349] In at least one embodiment, the application 3701 and the software stack 3700 run on the hardware 3707. In at least one embodiment, the hardware 3707 may include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support programming platforms. In at least one embodiment, for example, using CUDA, the software stack 3700 may be vendor-specific and only compatible with devices from a specific vendor. In at least one embodiment, for example, in using OpenCL, the software stack 3700 can be used with devices from different vendors. In at least one embodiment, the hardware 3707 includes a host connected to one or more devices, which can be accessed via an application programming interface (API) call to perform computing tasks. In at least one embodiment, compared to the host in the hardware 3707, which may include but is not limited to a CPU (but may also include a computing device) and its memory, the devices in the hardware 3707 may include but are not limited to a GPU, FPGA, AI engine or other computing device (but may also include a CPU) and its memory.
[0350] In at least one embodiment, the software stack 3700 of the programming platform includes, but is not limited to, multiple libraries 3703, runtime 3705, and device kernel drivers 3706. In at least one embodiment, each library in the library 3703 may include data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, the library 3703 may include, but is not limited to, pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. In at least one embodiment, the library 3703 includes functions optimized for execution on one or more types of devices. In at least one embodiment, the library 3703 may include, but is not limited to, functions for performing math, deep learning, and / or other types of operations on the device. In at least one embodiment, the library 3803 is associated with a corresponding API 3802, which may include one or more APIs that expose functions implemented in the library 3803.
[0351] In at least one embodiment, the application 3701 is written as source code that is compiled into executable code as follows: Fig.42 3701. In at least one embodiment, the executable code of the application 3701 can be run at least in part on an execution environment provided by the software stack 3700. In at least one embodiment, during the execution of the application 3701, code that needs to be run on the device (as opposed to the host) can be obtained. In this case, in at least one embodiment, the runtime 3705 can be called to load and start the necessary code on the device. In at least one embodiment, the runtime 3705 can include any technically feasible runtime system capable of supporting the execution of the application 3701.
[0352] In at least one embodiment, runtime 3705 is implemented as one or more runtime libraries associated with a corresponding API (which is shown as API 3704). In at least one embodiment, one or more such runtime libraries may include, but are not limited to, functions for memory management, execution control, device management, error handling, and / or synchronization, etc. In at least one embodiment, memory management functions may include, but are not limited to, functions for allocating, deallocating, and copying device memory and transferring data between host memory and device memory. In at least one embodiment, execution control functions may include, but are not limited to, functions for launching functions on a device (sometimes referred to as "kernels" when the function is a global function callable from the host), and functions for setting attribute values in buffers maintained by the runtime library for a given function to be executed on the device.
[0353] In at least one embodiment, the runtime library and corresponding API 3704 may be implemented in any technically feasible manner. In at least one embodiment, one (or any number of) APIs may expose a low-level set of functions for fine-grained control of a device, while another (or any number of) APIs may expose such a higher-level set of functions. In at least one embodiment, a high-level runtime API may be built on top of the low-level APIs. In at least one embodiment, one or more runtime APIs may be language-specific APIs layered on top of a language-independent runtime API.
[0354] In at least one embodiment, the device kernel driver 3706 is configured to facilitate communication with the underlying device. In at least one embodiment, the device kernel driver 3706 can provide APIs such as API3704 and / or low-level functions that other software relies on. In at least one embodiment, the device kernel driver 3706 can be configured to compile intermediate representation ("IR") code into binary code at runtime. In at least one embodiment, for CUDA, the device kernel driver 3706 can compile non-hardware-specific parallel thread execution ("PTX") IR code into binary code (cache compiled binary code) for a specific target device at runtime, which is sometimes also referred to as "final" code. In at least one embodiment, doing so can allow the final code to run on a target device that may not exist when the source code is initially compiled into PTX code. Alternatively, in at least one embodiment, the device source code can be compiled into binary code offline without the need for the device kernel driver 3706 to compile IR code at runtime.
[0355] Fig.38 According to at least one embodiment, Fig.37 3801. In at least one embodiment, the CUDA software stack 3800 on which the application 3801 can be launched includes a CUDA library 3803, a CUDA runtime 3805, a CUDA driver 3807, and a device kernel driver 3808. In at least one embodiment, the CUDA software stack 3800 is executed on hardware 3809, which may include a CUDA-enabled GPU developed by NVIDIA Corporation of Santa Clara, California.
[0356] In at least one embodiment, the application 3801, the CUDA runtime 3805, and the device kernel driver 3808 can respectively perform functions similar to the application 3701, the runtime 3705, and the device kernel driver 3706, and the above combination Fig.37It is described. In at least one embodiment, the CUDA driver 3807 includes a library (libcuda.so) that implements the CUDA driver API 3806. In at least one embodiment, similar to the CUDA runtime API 3804 implemented by the CUDA runtime library (cudart), the CUDA driver API 3806 can disclose but is not limited to functions for memory management, execution control, device management, error handling, synchronization and / or graphics interoperability, etc. In at least one embodiment, the CUDA driver API 3806 is different from the CUDA runtime API 3804 in that the CUDA runtime API 3804 simplifies device code management by providing implicit initialization, context (similar to process) management, and module (similar to dynamically loaded library) management. In contrast to the high-level CUDA runtime API 3804, in at least one embodiment, the CUDA driver API 3806 is a low-level API that provides finer-grained control over the device, particularly with respect to context and module loading. In at least one embodiment, the CUDA driver API 3806 can disclose functions for context management that are not disclosed by the CUDA runtime API 3804. In at least one embodiment, the CUDA driver API 3806 is also language-independent and supports, for example, OpenCL in addition to the CUDA runtime API 3804. Furthermore, in at least one embodiment, the development libraries including the CUDA runtime 3805 may be considered separate from the driver components, including the user-mode CUDA driver 3807 and the kernel-mode device driver 3808 (sometimes also referred to as a "display" driver).
[0357] In at least one embodiment, the CUDA library 3803 may include, but is not limited to, a math library, a deep learning library, a parallel algorithm library, and / or a signal / image / video processing library, which may be utilized by parallel computing applications (e.g., application 3801). In at least one embodiment, the CUDA library 3803 may include a math library, such as the cuBLAS library, which is an implementation of the Basic Linear Algebra Subroutines ("BLAS") for performing linear algebra operations; the cuFFT library for computing fast Fourier transforms ("FFTs"), and the cuRAND library for generating random numbers, etc. In at least one embodiment, the CUDA library 3803 may include a deep learning library, such as the cuDNN library for primitives for deep neural networks and the TensorRT platform for high-performance deep learning inference, etc.
[0358] Fig.39 According to at least one embodiment, Fig.37In at least one embodiment, the ROCm software stack 3900 on which the application 3901 can be launched includes a language runtime 3903, a system runtime 3905, a thunk 3907, a ROCm kernel driver 3908, and a device kernel driver 3909. In at least one embodiment, the ROCm software stack 3900 is executed on hardware 3909, which may include a ROCm-enabled GPU developed by AMD, Inc. of Santa Clara, California.
[0359] In at least one embodiment, application 3901 may execute in conjunction with the above Fig.37 In addition, in at least one embodiment, the language runtime 3903 and the system runtime 3905 can perform functions similar to those of the application 3701 discussed above. Fig.37 In at least one embodiment, the language runtime 3903 and the system runtime 3905 are different in that the system runtime 3905 is a language-independent runtime that implements the ROCr system runtime API 3904 and utilizes the heterogeneous system architecture ("HAS") runtime API. In at least one embodiment, the HAS runtime API is a thin user mode API that exposes interfaces for accessing and interacting with the AMD GPU, including functions for memory management, execution control of kernels dispatched by the architecture, error handling, system and agent information, and runtime initialization and shutdown. In at least one embodiment, compared to the system runtime 3905, the language runtime 3903 is an implementation of a language-specific runtime API 3902 layered on top of the ROCr system runtime API 3904. In at least one embodiment, the language runtime API may include, but is not limited to, a portable heterogeneous computing interface ("HIP") language runtime API, a heterogeneous computing compiler ("HCC") language runtime API, or an OpenCL API, among others. In particular, the HIP language is an extension of the C++ programming language with a functionally similar version of the CUDA mechanism, and in at least one embodiment, the HIP language runtime API includes a combination of the above Fig.38 Functions similar to the CUDA runtime API 3804 discussed above, such as functions for memory management, execution control, device management, error handling, synchronization, etc.
[0360] In at least one embodiment, thunk (ROCt) 3907 is an interface that can be used to interact with the underlying ROCm driver 3908. In at least one embodiment, the ROCm driver 3908 is a ROCk driver, which is a combination of the AMDGPU driver and the HAS kernel driver (amdkfd). In at least one embodiment, the AMDGPU driver is a device kernel driver for GPUs developed by AMD that performs the above combined Fig.37 The HAS kernel driver 3706 is a device kernel driver that allows different types of processors to share system resources more efficiently via hardware features.
[0361] In at least one embodiment, various libraries (not shown) may be included in the ROCm software stack 3900 above the language runtime 3903 and provide for integration with the above. Fig.38 The various libraries may include, but are not limited to, math, deep learning, and / or other libraries, such as a hipBLAS library that implements functions similar to CUDA cuBLAS, a rocFFT library similar to CUDA cuFFT for computing FFT, and the like.
[0362] Fig.40 According to at least one embodiment, Fig.37 4001. In at least one embodiment, the OpenCL software stack 4000 on which the application 4001 can be launched includes an OpenCL framework 4005, an OpenCL runtime 4006, and a driver 4007. In at least one embodiment, the OpenCL software stack 4000 is executed on hardware 3809 that is not vendor-specific. In at least one embodiment, because devices developed by different vendors support OpenCL, specific OpenCL drivers may be required to interoperate with hardware from such vendors.
[0363] In at least one embodiment, the application 4001, the OpenCL runtime 4006, the device kernel driver 4007 and the hardware 4008 can respectively execute the above combined Fig.37 The discussed application 3701, runtime 3705, device kernel driver 3706, and hardware 3707 have similar functionality. In at least one embodiment, the application 4001 also includes an OpenCL kernel 4002 having code to be executed on the device.
[0364] In at least one embodiment, OpenCL defines a "platform" that allows a host to control devices connected to the host. In at least one embodiment, the OpenCL framework provides a platform layer API and a runtime API, shown as platform API 4003 and runtime API 4005. In at least one embodiment, the runtime API 4005 uses contexts to manage the execution of kernels on devices. In at least one embodiment, each identified device can be associated with a respective context, and the runtime API 4005 can use the context to manage the command queue, program objects and kernel objects, shared memory objects, etc. of the device. In at least one embodiment, the platform API 4003 exposes functions that allow device contexts to be used to select and initialize devices, submit work to devices via command queues, and enable data transfers from and to devices, etc. In addition, in at least one embodiment, the OpenCL framework provides various built-in functions (not shown), including mathematical functions, relational functions, image processing functions, etc.
[0365] In at least one embodiment, compiler 4004 is also included in OpenCL framework 4005. In at least one embodiment, source code can be compiled offline before executing the application or compiled online during execution of the application. In contrast to CUDA and ROCm, OpenCL applications in at least one embodiment can be compiled online by compiler 4004, which is included to represent any number of compilers that can be used to compile source code and / or IR code (e.g., Standard Portable Intermediate Representation ("SPIR-V") code) into binary code. Alternatively, in at least one embodiment, OpenCL applications can be compiled offline before executing such applications.
[0366] Fig.41 Software supported by a programming platform according to at least one embodiment is shown. In at least one embodiment, programming platform 4104 is configured to support various programming models 4103, middleware and / or libraries 4102, and frameworks 4101 that applications 4100 can rely on. In at least one embodiment, application 4100 can be an AI / ML application implemented using, for example, a deep learning framework (in at least one embodiment, MXNet, PyTorch, or TensorFlow), which can rely on libraries such as cuDNN, NVIDIA Collective Communications Library ("NCCL"), and / or NVIDIA Developer Data Loading Library ("DALI") CUDA libraries to provide accelerated computation on the underlying hardware.
[0367] In at least one embodiment, programming platform 4104 can be a combination of the above Fig.38 , Fig.39 and Fig.40 One of the CUDA, ROCm, or OpenCL platforms described. In at least one embodiment, programming platform 4104 supports multiple programming models 4103, which are abstractions of the underlying computing system that allow the expression of algorithms and data structures. In at least one embodiment, programming model 4103 can expose features of the underlying hardware to improve performance. In at least one embodiment, programming model 4103 can include, but is not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism ("C++AMP"), Open Multiprocessing ("OpenMP"), Open Accelerators ("OpenACC"), and / or Vulcan Compute.
[0368] In at least one embodiment, the library and / or middleware 4102 provides an abstract implementation of the programming model 4104. In at least one embodiment, such a library includes data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, in addition to those that can be obtained from the programming platform 4104, such middleware also includes software that provides services to the application. In at least one embodiment, the library and / or middleware 4102 may include, but is not limited to, cuBLAS, cuFFT, cuRAND and other CUDA libraries, or rocBLAS, rocFFT, rocRAND and other ROCm libraries. In addition, in at least one embodiment, the library and / or middleware 4102 may include NCCL and ROCm communication collection libraries ("RCCL") libraries that provide communication routines for GPUs, MIOpen libraries for deep learning acceleration, and / or intrinsic libraries for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.
[0369] In at least one embodiment, application framework 4101 relies on library and / or middleware 4102. In at least one embodiment, each application framework 4101 is a software framework for implementing a standard structure of application software. In at least one embodiment, AI / ML applications can be implemented using a framework such as Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning framework.
[0370] Fig.42 Compiled code according to at least one embodiment is shown to Figure 37-40In at least one embodiment, compiler 4201 receives source code 4200, which includes both host code and device code. In at least one embodiment, compiler 4201 is configured to convert source code 4200 into host executable code 4202 for execution on the host and device executable code 4203 for execution on the device. In at least one embodiment, source code 4200 can be compiled offline before executing the application, or compiled online during execution of the application.
[0371] In at least one embodiment, source code 4200 may include code in any programming language supported by compiler 4201, such as C++, C, Fortran, etc. In at least one embodiment, source code 4200 may be included in a single-source file having a mixture of host code and device code, and indicating the location of the device code therein. In at least one embodiment, the single-source file may be a .cu file including CUDA code or a .hip.cpp file including HIP code. Alternatively, in at least one embodiment, source code 4200 may include multiple source code files, rather than a single source file, in which host code and device code are separated.
[0372] In at least one embodiment, compiler 4201 is configured to compile source code 4200 into host executable code 4202 for execution on a host and device executable code 4203 for execution on a device. In at least one embodiment, compiler 4201 performs operations including parsing source code 4200 into an abstract system tree (AST), performing optimizations, and generating executable code. In at least one embodiment where source code 4200 includes a single source file, compiler 4201 can separate device code from host code in such a single source file, compile the device code and host code into device executable code 4203 and host executable code 4202, respectively, and link device executable code 4203 and host executable code 4202 together in a single file, as described below with respect to Fig.26 discussed in more detail.
[0373] In at least one embodiment, host executable code 4202 and device executable code 4203 may be in any suitable format, such as binary code and / or IR code. In the case of CUDA, in at least one embodiment, host executable code 4202 may include native object code, while device executable code 4203 may include code in a PTX intermediate representation. In at least one embodiment, in the case of ROCm, both host executable code 4202 and device executable code 4203 may include target binary code.
[0374] Other variations are within the spirit of the present disclosure. Therefore, while the disclosed technology is susceptible to various modifications and alternative configurations, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, it is intended to cover all modifications, alternative configurations, and equivalents that fall within the spirit and scope of the present disclosure as defined by the appended claims.
[0375] Unless otherwise noted or clearly contradictory to the context, in the context of describing the disclosed embodiments (particularly in the context of the appended claims), the use of the terms "one" and "an" and "the" and similar references should be interpreted as covering the singular and plural, rather than as a definition of the term. Unless otherwise noted, the terms "include", "have", "include" and "contain" should be interpreted as open terms (meaning "including but not limited to") unless otherwise noted. The term "connected" (which refers to a physical connection when unmodified) should be interpreted as partially or completely included, attached to or connected together, even if there is some intervention. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, respectively, and each individual value is incorporated into the specification as if it were individually narrated herein. In at least one embodiment, unless otherwise noted or contradictory to the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set including one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.
[0376] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B and C" are understood in context to be generally used to indicate an item, clause, or the like, which may be A or B or C, or any non-empty subset of the set of A and B and C. In at least one embodiment of a set having three members, the conjunction phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunction language is generally not intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise indicated or contradicted by context, the term "plurality" refers to a plural state (e.g., "plurality of items" means a plurality of items). In at least one embodiment, the number of items in the plurality of items is at least two, but may be more if expressly indicated or indicated by context. Further, the phrase "based on" means "based at least in part on" rather than "based solely on" unless otherwise specified or clear from context.
[0377] Unless otherwise indicated herein or clearly contradictory to the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are executed together on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program, which in at least one embodiment includes a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all of the code, but rather a plurality of non-transitory computer-readable storage media stores all of the code together. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, and in at least one embodiment, the non-transitory computer-readable storage medium stores instructions, and the main central processing unit ("CPU") executes some instructions, while the graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.
[0378] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. In addition, a computer system that implements at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices that operate in different ways, so that the distributed computer system performs the operations described herein, and so that a single device does not perform all operations.
[0379] The use of any and all at least one example or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential to practicing the disclosure.
[0380] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0381] In the specification and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. On the contrary, in at least one embodiment, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0382] Unless explicitly stated otherwise, it is to be understood that throughout the specification, terms such as “processing”, “computing”, “calculating”, “determining” and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers or other such information storage, transmission or display devices of the computing system.
[0383] In a similar manner, the term "processor" may refer to a portion of any device or memory that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As one of the non-limiting at least one embodiment, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, in at least one embodiment, a "software" process may include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.
[0384] In this document, reference may be made to obtaining, acquiring, receiving or inputting analog or digital data into a subsystem, a computer system or a computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving or inputting analog and digital data may be accomplished in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving or inputting analog or digital data may be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving or inputting analog or digital data may be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending or presenting analog or digital data may be accomplished by transmitting data as an input or output parameter of a function call, an application programming interface or an interprocess communication mechanism.
[0385] Although the above discussion sets forth one implementation in at least one embodiment of the described technology, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. In addition, although specific responsibilities are defined above for discussion purposes, various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.
[0386] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A processor, include: One or more circuits for using a plurality of neural networks, each neural network being associated with a corresponding flow control valve in a plurality of flow control valves and controlling coolant flow to one of a plurality of different components in a data center, to perform at least one of: a) predicting coolant conditions at the plurality of different components, or b) regulating the coolant to independently cool the plurality of different components based at least in part on sensor data and one or more cooling requirements of the plurality of different components.
2. The processor according to claim 1, in, The coolant flows through the plurality of flow control valves of the liquid cooling system of the data center to control variations in coolant flow rates across the data center, and wherein the liquid cooling system includes a liquid cooling loop having a reverse return portion that provides similar path lengths for the plurality of different components cooled using the liquid cooling loop.
3. The processor of claim 1, wherein each of the plurality of flow control valves controls a flow of the coolant into or out of each of the plurality of different components.
4. The processor of claim 1, wherein the plurality of different components comprises at least one of a server rack, a server, a computer component, or a cold plate.
5. The processor of claim 1 , wherein the one or more circuits are further configured to use one or more additional neural networks to adjust the coolant to cool the multiple different components and maintain a change in coolant flow rate across the data center based at least in part on the predicted coolant conditions at the multiple different components output by the multiple neural networks.
6. The processor of claim 1 , wherein the one or more circuits are further configured to receive the sensor data to provide as input to one or more of the plurality of neural networks to determine one or more adjustments to be made to the plurality of flow control valves, the sensor data comprising at least one of flow rate, pressure, temperature, fluid velocity, power consumption, or workload at one or more locations in the data center.
7. A liquid flow distribution system, include: A control system for controlling coolant flow to one of a plurality of different components of a data center using a plurality of neural networks, each neural network being associated with a corresponding flow control valve in a plurality of flow control valves, to perform at least one of: a) predicting coolant conditions at the plurality of different components, or b) regulating the coolant to independently cool the plurality of different components based at least in part on sensor data and one or more cooling requirements of the plurality of different components.
8. The system according to claim 7, further comprising: include: A liquid cooling loop includes a reverse return portion that provides similar path lengths for the plurality of different components of the data center.
9. The system of claim 7, wherein the sensor data comprises at least one of flow rate, pressure, temperature, fluid velocity, power consumption, or workload of one or more of the plurality of different components.
10. The system of claim 7, wherein each flow control valve of the plurality of control valves controls flow of the coolant into or out of each of the plurality of different components.
11. The system of claim 8, wherein the plurality of different components include at least one of a server rack, a server, a server component, or a cold plate.
12. The system of claim 8, wherein the control system is further configured to adjust the coolant to cool the multiple different components and maintain a change in coolant flow rate across the data center using one or more additional neural networks based at least in part on the predicted coolant conditions at the multiple different components output by the multiple neural networks.
13. A liquid flow distribution system, include: One or more processors for using a plurality of neural networks, each neural network being associated with a corresponding flow control valve in a plurality of flow control valves and controlling coolant flow to one of a plurality of different components in a data center, to perform at least one of: a) predicting coolant conditions at the plurality of different components, or b) regulating the coolant to independently cool the plurality of different components based at least in part on sensor data and one or more cooling requirements of the plurality of different components.
14. The system of claim 13, wherein the data center include: A liquid cooling system includes a liquid cooling loop having a reverse return portion that provides an equivalent path length for the plurality of different components cooled using the liquid cooling loop.
15. The system of claim 13, wherein each of the plurality of flow control valves controls flow of the coolant into or out of each of the plurality of different components.
16. The system of claim 13, wherein the plurality of different components include at least one of a server rack, a server, a server component, or a cold plate.
17. The system of claim 13, wherein the one or more processors are further configured to maintain a change in coolant flow rate below a maximum change threshold.
18. The system of claim 13, wherein the one or more processors are further configured to receive the sensor data comprising at least one of flow rate, pressure, temperature, fluid velocity, power consumption, or workload at one or more locations in the data center.
19. A method for distributing a liquid flow, include: Using multiple neural networks, each neural network is associated with a corresponding flow control valve in a plurality of flow control valves and controls the flow of coolant to one of a plurality of different components in a data center to perform at least one of the following: a) predict coolant conditions at the plurality of different components, or b) regulate the coolant to independently cool the plurality of different components based at least in part on sensor data and one or more cooling requirements of the plurality of different components.
20. The method according to claim 19, in, The coolant flows through the plurality of flow control valves of the liquid cooling system of the data center to control variations in coolant flow rates across the data center, and wherein the liquid cooling system includes a liquid cooling loop having a reverse return portion that provides similar path lengths for the plurality of different components cooled using the liquid cooling loop.
21. The method of claim 19, wherein each of the plurality of flow control valves controls flow of the coolant into or out of each of the plurality of different components.
22. The method of claim 19, wherein the plurality of different components comprises at least one of a server rack, a server, a server component, or a cold plate.
23. The method according to claim 19, further comprising: include: Based at least in part on the predicted coolant conditions at the plurality of different components output by the plurality of neural networks, one or more additional neural networks are used to adjust the coolant to cool the plurality of different components and maintain a variation in coolant flow rate across the data center.
24. The method according to claim 19, further comprising: include: The sensor data is received and includes at least one of flow rate, pressure, temperature, fluid velocity, power consumption, or workload at one or more locations in the data center.
25. A machine-readable medium having stored thereon a set of instructions which, if executed by one or more processors, cause the one or more processors to at least: Using multiple neural networks, each neural network is associated with a corresponding flow control valve in a plurality of flow control valves and controls the flow of coolant to one of a plurality of different components in a data center to perform at least one of the following: a) predict coolant conditions at the plurality of different components, or b) regulate the coolant to independently cool the plurality of different components based at least in part on sensor data and one or more cooling requirements of the plurality of different components.
26. The machine-readable medium of claim 25, in, The coolant flows through the plurality of flow control valves of the liquid cooling system of the data center to control variations in coolant flow rates across the data center, and wherein the liquid cooling system includes a liquid cooling loop having a reverse return portion that provides equal path lengths for the plurality of different components cooled using the liquid cooling loop.
27. The machine-readable medium of claim 25, wherein each of the plurality of flow control valves controls flow of coolant into or out of each of the plurality of different components.
28. The machine-readable medium of claim 25, wherein the plurality of different components comprises at least one of a server rack, a server, a server component, or a cold plate.
29. The machine-readable medium of claim 25, wherein the one or more processors are further configured to use one or more additional neural networks to adjust the coolant to cool the multiple different components and maintain a change in coolant flow rate across the data center based at least in part on the predicted coolant states at the multiple different components output by the multiple neural networks.
30. The machine-readable medium of claim 25, wherein the one or more processors are further configured to receive the sensor data comprising at least one of flow rate, pressure, temperature, fluid velocity, power consumption, or workload at one or more locations in the data center.
Citation Information
Patent Citations
Provisioning cooling elements for chillerless data centers
CN104220949A
Adjustment of a pump speed based on a valve position
CN108693939A
Data center thermal performance optimization using distributed cooling systems
US20100076607A1
Cooling panel system
US20200208854A1