Power optimization technology based on closed-loop machine learning
By performing closed-loop machine learning inference locally on the workload server in the distributed cloud platform, and using local machine learning accelerators to predict and attribute adjustment, the processing delay problem caused by centralized closed-loop inference is solved, and effective optimization of delay-sensitive workloads is achieved.
Patent Information
- Application Number
- CN202480004881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-18
- Filing Date
- 2024-08-15
- Publication Date
- 2025-06-27
AI Technical Summary
Machine learning-based closed-loop inference often results in processing delays in a centralized way, which fails to meet delay-sensitive workloads such as radio resource management (RRM), which takes sub-milliseconds to tens of milliseconds to decision-making time.
Closed-loop machine learning inference is performed locally on a workload server running in a distributed cloud platform, using a local machine learning accelerator for predictions, and automatically adjusting the server's properties (such as power usage) to optimize the network.
By performing closed-loop machine learning inference locally, response time is significantly reduced and latency is reduced, especially for delay-sensitive workloads, avoiding the latency problem caused by centralized closed-loop inference.
Smart Images

Figure CN120225993A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is a continuation of U.S. Patent Application No. 18 / 381,367, filed on October 18, 2023, the disclosure of which is hereby incorporated herein by reference. Background Art
[0003] Machine - learning - based closed - loop inference refers to a continuous feedback mechanism based on one or more machine - learning models. The machine - learning model continuously monitors one or more metrics, performs inferences based on the one or more metrics, and automatically performs updates or adjustments in response to the inferences. The feedback loop of monitoring, inferring, and adjusting can occur automatically to self - regulate, maintain stability, and / or achieve a desired result. However, machine - learning - based closed - loop inference typically occurs in a centralized manner using dedicated servers that are located separately from the servers executing the workloads, which can significantly slow down processing. For example, centralized closed - loop inference may not be able to meet latency - sensitive workloads, such as radio resource management (RRM), which require decision times from sub - milliseconds to tens of milliseconds, e.g., to schedule cell resources based on predicted user handovers from neighbor cells to the serving cell. Summary of the Invention
[0004] Aspects of the present disclosure relate to network optimization of workload servers through locally executed closed - loop machine - learning inference on various workload servers running in a distributed cloud platform. The workload servers can each be equipped with one or more machine - learning accelerators to perform local predictions for the workload servers respectively. In response to the local predictions, attributes of the workload servers, such as power usage, can be automatically adjusted for network optimization. The implementation of the local machine - learning accelerators in each workload server of the distributed cloud platform can reduce the response time for adjusting the attributes, resulting in significant savings in latency, especially for latency - sensitive workloads that may not tolerate the latency from centralized closed - loop inference.
[0005] One aspect of the present disclosure provides a method for managing one or more local processing units in a server computing device of a distributed cloud platform, the method comprising: receiving, by one or more processors, one or more metrics associated with the one or more local processing units for executing a workload; generating, by the one or more processors, one or more predictions for one or more states of the one or more local processing units based on the one or more metrics using a machine - learning model deployed on one or more accelerators in the server computing device; and adjusting, by the one or more processing units, the one or more states of the one or more local processing units based on the predictions.
[0006] In an example, the one or more metrics include at least one of the following: power utilization per processing unit core, power consumption per application running on a processing unit core, number of enabled processing unit core C-states, number of enabled processing unit core P-states, or number of instructions per cycle that the processing unit is processing. In another example, the one or more local processing units include at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). In yet another example, the one or more accelerators include at least one of a tensor processing unit (TPU) or a wafer-scale engine (WSE).
[0007] In yet another example, the method further includes locally training, by the one or more processors using the one or more metrics, a machine learning model in the server computing device. In yet another example, the method further includes receiving, by the one or more processors, a machine learning model that is externally pre-trained on a decentralized service management and orchestration (SMO) platform.
[0008] In yet another example, adjusting the one or more states includes adjusting at least one of the frequency, voltage, power, C-state, P-state, or sleep state of the one or more local processing units. In another example, adjusting the one or more states includes adjusting the one or more states of a group of the one or more local processing units. In yet another example, the workload includes at least one of a radio access network (RAN) function, an access and mobility management function (AMF), a user plane function (UPF), or a session management function (SMF).
[0009] Another aspect of the present disclosure provides a system including: one or more processors; and one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for managing one or more local processing units in a server computing device of a distributed cloud platform, the operations including: receiving one or more metrics associated with the one or more local processing units for performing a workload; generating, based on the one or more metrics using a machine learning model deployed on one or more accelerators in the server computing device, one or more predictions for one or more states of the one or more local processing units; and adjusting the one or more states of the one or more local processing units based on the predictions.
[0010] In an example, the one or more metrics include at least one of the following: power utilization per processing unit core, power consumption per application running on a processing unit core, number of enabled processing unit core C-states, number of enabled processing unit core P-states, or number of instructions per cycle that the processing unit is processing. In another example, the one or more local processing units include at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field programmable gate array (FPGA). In yet another example, the one or more accelerators include at least one of a tensor processing unit (TPU) or a wafer scale engine (WSE).
[0011] In yet another example, the operations further include locally training a machine learning model in the server computing device using the one or more metrics. In yet another example, the operations further include receiving a machine learning model that has been externally pre-trained on a decentralized service management and orchestration (SMO) platform.
[0012] In yet another example, adjusting the one or more states includes adjusting at least one of the frequency, voltage, power, C-state, P-state, or sleep state of the one or more local processing units. In yet another example, adjusting the one or more states includes adjusting the one or more states of a group of the one or more local processing units. In yet another example, the workload includes at least one of a radio access network (RAN) function, an access and mobility management function (AMF), a user plane function (UPF), or a session management function (SMF).
[0013] Another aspect of the present disclosure provides a non-transitory computer-readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for managing one or more local processing units in a server computing device of a distributed cloud platform, the operations including: receiving one or more metrics associated with the one or more local processing units for performing a workload; generating, based on the one or more metrics, one or more predictions for one or more states of the one or more local processing units using a machine learning model deployed on one or more accelerators in the server computing device; and adjusting the one or more states of the one or more local processing units based on the predictions.
[0014] In an example, the operations further include locally training a machine learning model in the server computing device using the one or more metrics. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1A A block diagram depicting an example server computing device that implements local closed-loop machine learning inference in accordance with aspects of the present disclosure is shown.
[0016] Figure 1B A block diagram depicting an example environment for implementing local closed-loop machine learning inference for a distributed cloud platform in accordance with aspects of the present disclosure.
[0017] Figure 2 A block diagram depicting one or more machine learning model architectures in accordance with aspects of the present disclosure.
[0018] Figure 3 A block diagram depicting an example server management system that can perform local closed-loop machine learning inference on one or more processors using a locally trained machine learning model in accordance with aspects of the present disclosure.
[0019] Figure 4 A block diagram depicting an example server management system that can perform local closed-loop machine learning inference on one or more processors using an externally trained machine learning model in accordance with aspects of the present disclosure.
[0020] Figure 5 A flowchart depicting an example process for performing local closed-loop machine learning inference using a local accelerator to manage one or more local processors in accordance with aspects of the present disclosure. Detailed Description
[0021] The present technology generally relates to closed-loop machine learning-based power optimization techniques for a distributed cloud platform. Power settings for the distributed cloud platform can be predicted based on one or more local machine learning accelerator chips in each distributed cloud platform server.
[0022] Cloud platform servers can each include a power manager and a metric collector. The power manager can include a P manager and a C manager to perform power control tasks by changing the P state and / or C state of individual processors respectively. The metric collector can derive one or more metrics such as per-processor metrics from the processors based on the power control tasks, and stream the metrics upstream to other servers in the distributed cloud platform. Example metrics can include power utilization per processor core, power consumption per application running on a processor core, the number of enabled processor core C states, and / or the number of instructions per cycle that the processor is processing. The power manager can statically adjust power consumption by attaching a power profile to individual processors or a group of processors selected by the scheduler to run an application. The power profile can be derived by a prediction engine based on one or more metrics collected by the metric collector. The power profile can include one or more states of the individual processor or group of processors, such as frequency, voltage, and / or sleep state. The power manager can also dynamically adjust the power consumption of an individual processor or group of processors by using the prediction engine to schedule shared processors based on one or more metrics. Scheduling of the shared processors can include scheduling processors from a reserved pool of processors and / or a partial pool of processors. The power manager can also dynamically adjust power consumption by deriving insights and patterns based on one or more metrics to dynamically change the power and sleep settings of an individual processor or group of processors.
[0023] Based on insights and predictions, the power manager can also dynamically adjust the power controls of local processors, thereby allowing power settings to be predicted based on the local processors of each cloud platform server. This will remove the need for a centralized control loop in a large-scale network, which can delay decision-making and leverage distributed decision-making regardless of the particular server. The large-scale network can be, for example, a distributed cloud platform running one or more of RAN, AMF, UPF, SMF, and other services, which are part of the infrastructure and implementation of a mobile telecommunications network. This will also allow power settings to be predicted in cases where a centralized control loop is unavailable, such as due to resource constraints.
[0024] The cloud platform server of the distributed cloud platform includes one or more local power manager processors, such as TPUs, GPUs, and / or CPUs, which are on-board and are used to utilize local closed-loop machine learning inference without involving a control loop external to the cloud platform server. The local power manager processors can be associated with a controller and each is associated with a manager instance. The local power manager processors can access local processor metrics via a metrics collector. The controller can control the inference via each local power manager processor based on trends from a model trained on the local processor metrics. The manager instance can act on the inference made by the controller to control the power settings of the cloud platform server via the frequency and sleep state of the local processors. The local power manager processors can apply millisecond to microsecond-level granularity when controlling the power settings.
[0025] The distributed cloud platform can include decentralized service management and orchestration (SMO) for power management of the cloud platform server. The decentralized SMO can include an AI / ML platform for performing machine learning model management, training, and / or lifecycle management (LCM). The AI / ML platform can collect metrics data from the cloud platform server and use the metrics data to perform model training. Machine learning models can be trained by workload, such as by network function. Example network functions can include radio access network functions, such as virtualized distributed units (VDUs) and / or virtualized centralized units (VCUs), access and mobility management functions (AMFs), user plane functions (UPFs), and / or session management functions (SMFs), but any workload can be used to train machine learning models.
[0026] For example, the VDU function can be implemented as a virtualized network function (VNF) or a containerized network function (CNF), which are decoupled from the underlying hardware and operate on a server. The PHY layer and MAC layer may require high computational complexity, such as for channel estimation and detection, forward error correction (FEC), and / or scheduling algorithms, resulting in a high load on the server's computing power and degrading the performance of the VDU. Some computationally intensive tasks with repetitive structures, such as FEC, can be offloaded to an alternative hardware chip installed on the server for acceleration. The VDU can include multiple pods, including one or more containers conforming to a microservices-type architecture. The server resources such as processor cores and memory occupied by each pod can vary significantly. Additionally, the pods can be scaled based on capacity requirements. This allows the VDU to configure the appropriate processor core and memory sizes using a trained machine learning model according to the capacity and performance requirements in a specific network deployment. The management and orchestration of VDU containers can be supported by a system such as one that automatically distributes, scales, and manages containerized applications using a trained machine learning model.
[0027] The AI / ML platform may include a directory for publishing trained machine learning models for downloading to each cloud platform server. Alternatively or additionally, each cloud platform server may include a training platform for using the metric data of the cloud platform server for model training. Each cloud platform server may also include an inference platform that works in coordination with a manager instance. Each cloud platform server may include one or more workloads bound to a specific local processor. The machine learning model trained by the workload may predict one or more states of the local processor, such as frequency, voltage, and / or sleep state. Using the machine learning model prediction, the manager instance and the controller may provide instructions to a power manager, such as a C manager and / or a P manager, to configure one or more states of the local processor.
[0028] Figure 1A A block diagram of an example server computing device 10 implementing local closed-loop machine learning inference is depicted. For example, the server computing device 10 may be an edge device that provides an entry point into a distributed cloud platform. Example edge devices may include routers, switches, or other access devices. The server computing device 10 may include one or more local processing units 12, one or more local accelerators 14, and a server management system 16. The server management system 16 may receive one or more metrics 18 associated with the local processing unit 12. Based on the metrics 18, the server management system 16 may deploy one or more machine learning models to use the local accelerators 14 to generate inferences 20 about the performance of the local processing unit 12. In response to the inferences 20, the server management system 16 may adjust one or more states 20 of the local processing unit 12. The server management system 16 may receive the metrics 18, generate the inferences 20, and iteratively adjust the states 20, which represents a closed loop within the server computing device 10 for optimizing the performance of the local processing unit 12.
[0029] Figure 1B A block diagram of an example environment 100 for implementing local closed-loop machine learning inference in a distributed cloud platform is depicted. The distributed cloud platform may provide services that allow the provisioning or maintenance of computing resources and / or applications such as for data centers, cloud environments, and / or container frameworks. For example, a cloud-based platform may be used as a service to provide software applications such as communication services, accounting, word processing, inventory tracking, etc. As another example, the infrastructure of the platform may be partitioned in the form of virtual machines or containers on which software applications run.
[0030] A distributed cloud platform may be implemented on one or more devices having one or more processors in one or more locations, such as being implemented in a plurality of server computing devices 102A-N and one or more client computing devices 104. Any of the server computing devices among the plurality of server computing devices 102 may correspond to the server computing device 10 depicted as in Figure 1A The plurality of server computing devices 102 and the client computing devices 104 may be communicatively coupled to one or more storage devices 106 via a network 108. The storage device 106 may be a combination of volatile and non-volatile memories and may be located at the same or different physical locations from the computing devices 102, 104. For example, the storage device 106 may include any type of non-transitory computer-readable medium capable of storing information, such as a hard disk drive, a solid-state drive, a tape drive, an optical storage device, a memory card, ROM, RAM, DVD, CD-ROM, writable memory, and read-only memory.
[0031] Each of the server computing devices 102 may include one or more processors 110, a memory 112, and a hardware accelerator 114. The processor 110 may be designated to execute one or more workloads of the services provided by the distributed cloud platform. The accelerator 114 may be designated to deploy one or more machine learning models, such as for predicting processor states, such as frequency, voltage, and / or sleep state. Example processors 110 may include a central processing unit (CPU), a graphics processing unit (GPU), and / or a field-programmable gate array (FPGA). Example accelerators 114 may also include a GPU and / or an FPGA, as well as an application-specific integrated circuit (ASIC), such as a tensor processing unit (TPU) or a wafer-scale engine (WSE).
[0032] The memory 112 may store information accessible by the processor 110 and / or the accelerator 114, including instructions 116 executable by the processor 110 and / or the accelerator 114. The memory 112 may also include data 118 retrievable, manipulable, or storable by the processor 110 and / or the accelerator 114. The memory 112 may be any type of transitory or non-transitory computer-readable medium capable of storing information accessible by the processor 110 and / or the accelerator 114, such as volatile or non-volatile memory. Example memories 112 may include high-bandwidth memory (HBM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), flash memory, and / or read-only memory (ROM).
[0033] Instruction 116 may include one or more instructions that, when executed by processor 110 and / or accelerator 114, cause one or more processors 110 and / or accelerators 114 to perform the actions defined by instruction 116. Instruction 116 may be stored in a target code format for direct processing by processor 110 and / or accelerator 114, or in other formats including interpreted-on-demand or pre-compiled interpretable scripts or collections of stand-alone source code modules. Instruction 116 may include instructions for implementing server management system 120, which will be described further below. Server management system 120 may be executed using processor 110 and / or accelerator 114 and / or using other processors and / or accelerators remotely located on other server computing devices.
[0034] Data 118 may be retrieved, stored, or modified by processor 110 and / or accelerator 114 according to instruction 116. Data 118 may be stored in computer registers, stored in a relational or non-relational database as a table with multiple different fields and records, or stored as a JSON, YAML, proto, or XML document. Data 118 may also be formatted in a computer-readable format such as, but not limited to, binary values, ASCII, or Unicode. Additionally, data 118 may include information sufficient to identify related information such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories - including other network locations - or information used by a function to compute related data.
[0035] Client computing device 104 may be similarly configured to server computing device 102, having one or more processors 122, memory 124, instructions 126, and data 128. Client computing device 104 may also include user input 130 and user output 132. User input 130 may include any suitable mechanism or technique for receiving input from a user, such as a keyboard, mouse, mechanical actuator, soft actuator, touch screen, microphone, and sensors. User output 132 may include any suitable mechanism or technique for providing information to a platform user of client computing device 104. For example, user output 132 may include a display for displaying at least a portion of data received from one or more of server computing devices 102. As another example, user output 132 may include an interface between client computing device 104 and one or more of server computing devices 102. As yet another example, user output 132 may include one or more speakers, transducers, or other audio outputs, or a haptic interface or other haptic feedback that provides non-visual and non-auditory information to a platform user of client computing device 104.
[0036] Although Figure 1B processors 110, 122, memories 112, 124, and accelerators 114 are shown as being within respective computing devices 102, 104, the components described herein can include multiple processors, memories, and accelerators that can be in different physical locations and operate in different computing devices. For example, some of instructions 116, 126 and data 118, 128 can be stored on a removable SD card, and others can be stored within a read-only computer chip. Some or all of instructions 116, 126 and data 118, 128 can be stored in a location physically remote from processors 110, 122 and / or accelerators 114 but still accessible by the processor and / or the accelerator. Similarly, processors 110, 122 and / or accelerators 114 can include a collection of processors and / or accelerators that can perform concurrent and / or sequential operations. Computing devices 102, 104 can each include one or more internal clocks that provide timing information that can be used for the time measurement of operations and programs run by computing devices 102, 104.
[0037] One or more server computing devices among server computing devices 102 can be configured to receive a request to process data, such as a portion of a query for a specific task, from a client computing device 104. Server computing device 102 can receive the query, process the query, and, in response, generate output data, such as a response to the query for a specific task. While server computing device 102 is processing and responding to requests, server management system 120 can monitor various metrics of processor 110. Server management system 120 can continuously or periodically input those metrics into a deployed machine learning model using accelerator 114. The machine learning model can output one or more predictions associated with the usage of processor 110. Based on the one or more predictions, server management system 120 can adjust various states of processor 110, such as for optimizing power usage.
[0038] Figure 2 Block diagram 200 is depicted, which shows one or more machine learning model architectures 202 for deployment in server computing device 204, more specifically, for each architecture 202A through 202N, the server computing device accommodating one or more hardware accelerators 206 on which the deployed machine learning models 202 will execute. Server computing device 204 can correspond to any of the server computing devices among server computing devices 102 as depicted in Figure 1B Hardware accelerator 206 can be any type of processor, such as a CPU, GPU, FPGA, and / or ASIC, such as a TPU or WSE.
[0039] The architecture 202 of a machine learning model may refer to defining the characteristics of the model, such as the characteristics of the layers of the model, how the layers process the input, or how the layers interact with each other. The architecture 202 of a machine learning model may also define the types of operations performed within each layer. One or more machine learning model architectures 202 that can output results may be generated, such as for network optimization in a distributed cloud platform. For example, network optimization aims to improve the operational efficiency of components that provide and execute scalable telecommunications services in a distributed cloud platform, such components as individual server devices. Example model architectures 202 may correspond to predictive models, such as classification models, clustering models, forecasting models, outlier models, time series models, neural networks, decision trees, generalized linear models, and / or gradient boosting models.
[0040] Return reference Figure 1B , devices 102, 104 may be able to communicate directly and indirectly via network 108. For example, using network sockets, client computing device 104 may connect via an Internet protocol to a service operating using server computing device 102. Devices 102, 104 may establish listening sockets that can accept incoming connections for sending and receiving information. Network 108 may include various configurations and protocols, including the Internet, World Wide Web, intranet, virtual private network, wide area network, local area network, and private networks using one or more company - proprietary communication protocols. Network 108 may support various short - range and long - range connections. Short - range and long - range connections may operate at different bandwidths, such as 2.402 Ghz to 2.480 GHz typically associated with standards, 2.4 Ghz and 5 GHz typically associated with communication protocols; or may operate using various communication standards, such as or 5G standards for wireless broadband communication. Additionally or alternatively, network 108 may also support wired connections between devices 102, 104, including via various types of Ethernet connections.
[0041] Although three server computing devices 102, client computing device 104, and storage device 106 are shown in Figure 1B , it should be understood that example environment 100 may be implemented according to any number of server computing devices 102, client computing devices 104, and storage devices 106. Example environment 100 may be implemented according to various different configurations and numbers of computing devices, including being implemented in a paradigm for sequential or parallel processing, being implemented via a distributed network of multiple devices.
[0042] Figure 3Depicts a block diagram of an example server management system 300 that can perform local closed-loop machine learning inference on one or more processors, such as for network optimization. The server management system 300 can be implemented on each of a plurality of server computing devices in a distributed cloud platform, such as the server computing device 102 depicted in Figure 1B as shown.
[0043] The server management system 300 can be configured to receive metric data 302 from one or more local processors 304. Local processors can refer to one or more processors located on the same server computing device as the server management system 300. The server management system 300 can also be configured to receive metric data 306 from one or more external processors (not shown). External processors can refer to one or more processors located on other server computing devices other than the server computing device containing the server management system 300, such as upstream or downstream server computing devices in a distributed cloud platform. Example metrics can include power utilization per processor core, power consumption per application running on a processor core, the number of enabled processor core C states, the number of enabled processor core P states, and / or the number of instructions per cycle that the processor is processing. The metric data 302, 306 can be per-processor metric data or metric data for a group of processors.
[0044] The server management system 300 can receive the metric data 302, 306 as part of a call to an application programming interface (API), thereby exposing the server management system 300 to one or more server computing devices. Example APIs can include remote procedure call (RPC) and / or representational state transfer (REST). The server management system 300 can also receive the metric data 302, 306 via a storage medium, such as remote storage connected to the server computing device containing the server management system 300. The server management system 300 can also receive the metric data 302, 306 via a user interface on a client computing device that is coupled to the server computing device via a network.
[0045] Based on the metric data 302, 306, the server management system 300 can be configured to output processor adjustment data 308 for adjusting one or more states of the local processor 104. The server management system 300 can also be configured to output external processor adjustment data 310 for adjusting one or more states of one or more external processors (not shown). Example states of the processor can include a frequency amount, a voltage amount, a power amount, and / or whether the processor should be placed in a sleep or active state. The processor adjustment data 308, 310 can be per-processor processor adjustment data or processor adjustment data for a group of processors. In some examples, adjusting one or more states of an individual processor includes changing the P state and / or the C state.
[0046] The server management system 300 can be configured to provide the processor adjustment data 308, 310 as a set of computer-readable instructions, such as one or more computer programs. The computer programs can be written in any type of programming language and according to any programming paradigm—for example, declarative, procedural, assembly, object-oriented, data-oriented, functional, or imperative. The computer programs can be written to perform one or more different functions and operate within a computing environment, such as on a physical device, on a virtual machine, or across multiple devices. The computer programs can also implement the functionality described herein, such as the functionality performed by a system, an engine, a module, or a model. The server management system 300 can also forward the processor adjustment data 308, 310 to one or more other devices configured to convert the output data into an executable program written in a computer programming language. The server management system 300 can also be configured to send the processor adjustment data 308, 310 for display on a client device. The server management system 300 can also be configured to send the processor adjustment data 308, 310 to a storage device for storage and later retrieval.
[0047] The server management system 300 can include a metric collector 312, a model trainer 314, an accelerator manager and / or controller 316, and a processor manager and / or controller 318. The metric collector 312, the model trainer 314, the accelerator manager and / or controller 316, and the processor manager and / or controller 318 can be implemented as one or more computer programs, specially configured electronic circuitry, or any combination thereof.
[0048] The metric collector 312 can be configured to continuously or periodically derive metrics regarding the local processor 304 and / or an external processor. The metrics can be per-processor metrics or metrics for a group of processors. Example metrics can include power utilization per processor core, power consumption per application running on a processor core, the number of enabled processor core C-states, the number of enabled processor core P-states, and / or the number of instructions per cycle that the processor is processing. As an example, the metric collector 312 can organize the metrics into a tabular format where rows can represent each local processor in the local processor 104 and columns can represent the metrics that describe the local processor 104. The metrics in the columns can be features on which a machine learning model is trained. The metric collector 312 can also be configured to continuously or periodically output the metrics as training data 320 for training one or more machine learning models, and as inference data 322 for performing inference by the deployed machine learning models. The metric output as training data 320 and the metric output as inference data 322 can be the same metric data or different metric data from the same or different local processors 304. The metric collector 312 can also continuously or periodically output the metrics for storage in a database.
[0049] The model trainer 314 can be configured to train one or more machine learning models 324 for performing local closed-loop machine learning inference on the local processor 304, such as for network optimization. The machine learning models 324 can be trained per workload, such as per network function or per service provided by a distributed cloud platform. Example network functions can include radio access network functions such as virtualized distributed unit (VDU) and / or virtualized centralized unit (VCU), access and mobility management function (AMF), user plane function (UPF), and / or session management function (SMF), but any workload or service can be used as a scenario for training the machine learning models. The model trainer 314 can use the metrics received by the metric collector 312 to train the machine learning models 324. The model trainer 314 can provide the trained machine learning models 324 to the accelerator manager and / or controller 316.
[0050] The machine learning model 324 can be trained according to one of various different learning techniques. The learning techniques for training the model can include supervised learning, unsupervised learning, semi-supervised learning, and / or reinforcement learning techniques. The training data can include a plurality of training examples that can be received as input by the machine learning model 324. The training examples can be labeled with the expected output of the machine learning model 324 when processing the labeled training examples. For example, the training examples can include one or more of the metrics from the metric collector 312 and the expected output regarding the throughput, energy consumption, and / or latency of the local processor 304. The labels and the model output can be evaluated by a loss function to determine the error, and the error can be backpropagated through the machine learning model 324 to update the weights of the machine learning model. For example, supervised learning techniques can be applied to calculate the error between the outputs using the ground truth labels of the training examples processed by the machine learning model 324. Any of various loss or error functions can be utilized, such as cross-entropy loss for classification tasks or mean squared error for regression tasks. The error gradients with respect to the different weights of the candidate model on the candidate hardware can be calculated, for example, using the backpropagation algorithm, and the weights of the machine learning model can be updated. The machine learning model 324 can be trained until a stopping criterion is met, such as the number of training iterations, the maximum time period, convergence, or when a minimum accuracy threshold is satisfied. Regularization, data augmentation, and / or early stopping can also be utilized to train the machine learning model 324 to prevent overfitting and improve the generalization ability of the machine learning model 324.
[0051] The accelerator manager and / or controller 316 can be configured to deploy the machine learning model 324 on the local accelerator 326 in the server computing device. The accelerator manager and / or controller 316 can configure the machine learning model 324 to perform model inference 328 in terms of the performance of the local processor 304 using the metrics 322 from the metric collector 312, e.g., to determine predictions, patterns, and / or trends. For example, the accelerator manager and / or controller 316 can be configured to deploy one or more machine learning models 324 on the local accelerator 326 to determine the C-state, P-state, frequency, and / or voltage settings per local processor 304 or per group of local processors 304. The machine learning model 324 can output these settings along with a confidence score that indicates the degree of confidence of the machine learning model 324 in the accuracy of these settings. The model inference 328 determined by the machine learning model 324 using the local accelerator 326 can be output back to the accelerator manager and / or controller 316 to then be sent to the processor manager and / or controller 318. Alternatively or additionally, the model inference 328 determined by the machine learning model 324 can be directly output to the processor manager and / or controller 318.
[0052] The processor manager and / or controller 318 may be configured to adjust one or more settings or states of one or more of the local processors 304 based on a model inference 328 determined by a machine learning model 324. The processor manager and / or controller 318 may include one or more processor-setting-specific sub-controllers for controlling individual settings or states of the local processors 304, such as a P-state controller, a C-state controller, and / or a power controller. The processor manager and / or controller 318 may change the settings or states per processor or per group of processors. The processor manager and / or controller 318 may also provide instructions to change the settings or states of an external processor based on the model inference 328. Example settings or states that may be changed include the frequency, voltage, power, C-state, P-state, and / or sleep state of the local processors 304.
[0053] Figure 4 A block diagram of an example server management system 400 is depicted, which may perform local closed-loop machine learning inference on one or more processors, such as for network optimization. Figure 4 The server management system 400 in Figure 3 may be configured similarly to the server management system 300 depicted in
[0054] The server management system 400 may be configured to receive metric data 402 from one or more local processors 404. Based on the metric data 402, the server management system 400 may be configured to output processor adjustment data 408 for adjusting one or more states of the local processors 404. The metric collector 412 may be configured to continuously or periodically derive metrics regarding the local processors 404 and output the metrics 422 for model inference to the accelerator manager and / or controller 416. The accelerator manager and / or controller 416 may be configured to deploy a machine learning model 424 on the local accelerator 426 to perform a model inference 428 on one or more states or settings of the local processors 404. The accelerator manager and / or controller 416 may output the model inference 428 to the processor manager and / or controller 418 to perform a processor adjustment 408 on one or more states or settings of the local processors 404.
[0055] The machine learning model 424 deployed by the server management system 400 can be externally trained on one or more other server computing devices other than the server computing device including the server management system 400. For example, a distributed cloud platform can include decentralized service management and orchestration (SMO) 410 for the management of server computing devices. The SMO 410 can include an external model trainer 414, which serves as a dedicated platform for performing machine learning management, training, and / or lifecycle management. The external model trainer 414 can train one or more machine learning models 424 for deployment on the server management system 400 and other machine learning models 430 for deployment to other server management systems on other server computing devices. The external model trainer 414 can use the metrics 420 from the metric collector 412 and the metrics 406 from the metric collectors of other server management systems to train the machine learning models. The machine learning model 424 can be trained according to one of a variety of different learning techniques, such as supervised learning, unsupervised learning, semi-supervised learning, and / or reinforcement learning techniques.
[0056] Implementing the external model trainer 414 can improve the overall accuracy of the machine learning model 424 being deployed because the external model trainer 414 can use more data from many metric collectors of various server computing devices in the distributed cloud platform to train the machine learning model 424. Implementing the external model trainer 414 can also accelerate training because a larger number of accelerators are available for training the machine learning model 424.
[0057] Figure 5 A flowchart depicting an example process 500 for performing local closed-loop machine learning inference using a local accelerator to manage one or more local processors is shown. The example process 500 can be executed on a system of one or more processors and / or accelerators on a server computing device of a distributed cloud platform - such as the server management system 300 or the server management system 400 depicted respectively in Figure 3 and Figure 4 as depicted.
[0058] As shown in block 510, the server management system 300 or 400 can receive one or more metrics associated with one or more local processors. The metrics can be per-workload metrics, such as per-network function metrics. Example workloads can include radio access network (RAN) functions, access and mobility management functions (AMF), user plane functions (UPF), and / or session management functions (SMF). As an example, one or more metrics can include power utilization per processor core, power consumption per application running on a processor core, number of enabled processor core C-states, number of enabled processor core P-states, and / or number of instructions per cycle that the processor is processing. As an example, one or more local processors can include a CPU, a GPU, and / or an FPGA.
[0059] As shown in block 520, the server management system 300 or 400 can generate one or more predictions for one or more states of one or more local processors based on one or more metrics. The server management system 300 or 400 can use a machine learning model deployed on one or more accelerators in the server computing device to generate the predictions. As an example, one or more accelerators can include a TPU and / or a WSE. The machine learning model can be locally trained in the server computing device or externally pre-trained in one or more other server computing devices, such as via an SMO platform. The one or more metrics and / or one or more metrics from external processors on other server computing devices can be used to train the machine learning model.
[0060] As shown in block 530, the server management system 300 or 400 can adjust one or more states of one or more local processors based on the predictions. Adjusting one or more states can also include adjusting the frequency, voltage, power, C-state, P-state, or sleep state of one or more local processors. Adjusting one or more states can also include adjusting the states of a group of one or more local processors or a single local processor.
[0061] Aspects of the present disclosure may be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, and / or in computer hardware, such as the structures disclosed herein, structural equivalents thereof, or combinations thereof. Aspects of the present disclosure may also be implemented as one or more computer programs, such as one or more modules of computer program instructions encoded on a tangible non-transitory computer storage medium for execution by, or to control the operation of, one or more data processing apparatuses. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination thereof. The computer program instructions may be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to the appropriate receiver device for execution by the data processing apparatus.
[0062] The term "configured" is used herein in connection with systems and computer program components. For a system of one or more computers that is configured to perform particular operations or actions, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that causes the system to perform those operations or actions. For one or more computer programs that are configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by one or more data processing apparatuses, cause the apparatus to perform those operations or actions.
[0063] The term "data processing apparatus" or "data processing system" refers to data processing hardware and encompasses a variety of devices, apparatuses, and machines for processing data, including programmable processors, computers, or combinations thereof. The data processing apparatus may include dedicated logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The data processing apparatus may include code that creates an execution environment for a computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof.
[0064] The term "computer program" refers to a program, software, software application, app, module, software module, script, or code. A computer program can be written in any form of programming language, including compiled, interpreted, declarative, or procedural languages, or a combination thereof. A computer program can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program can correspond to a file in a file system and can be stored as part of a file that holds other programs or data—such as one or more scripts stored in a markup language document, stored in a single file dedicated to the program in question, or stored in multiple coordinated files—such as files that hold one or more modules, subroutines, or portions of code. A computer program can be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0065] The term "database" refers to any collection of data. The data can be unstructured or structured in any way. The data can be stored on one or more storage devices in one or more locations. For example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.
[0066] The term "engine" refers to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. An engine can be implemented as one or more software modules or components or can be installed on one or more computers located at one or more locations. A particular engine can have one or more computers dedicated to it, or multiple engines can be installed and run on the same one or more computers.
[0067] The processes and logical flows described herein can be performed by one or more computers that execute one or more computer programs to perform functions by operating on input data and generating output data. The processes and logical flows can also be performed by, or by a combination of, special-purpose logic circuitry and one or more computers.
[0068] A computer or a dedicated logic circuit system that executes one or more computer programs may include a central processing unit including a general-purpose or a dedicated microprocessor for running or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit may receive instructions and data from one or more memory devices, such as read-only memory, random access memory, or a combination thereof, and may run or execute these instructions. The computer or the dedicated logic circuit system may also include or be operatively coupled to one or more storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, to receive data from or transfer data to the storage device. The computer or the dedicated logic circuit system may be embedded in another device, such as, by way of example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS), or a portable storage device, such as a universal serial bus (USB) flash drive.
[0069] A computer-readable medium suitable for storing one or more computer programs may include any form of volatile or non-volatile memory, medium, or memory device. Examples include: semiconductor memory devices, such as EPROM, EEPROM, or flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; CD-ROM disks; DVD-ROM disks; or a combination thereof.
[0070] Aspects of the present disclosure may be implemented in a computing system that includes: backend components, such as, by way of example, a data server; middleware components, such as, by way of example, an application server; or frontend components, such as, by way of example, a client computer having a graphical user interface, a web browser, or an app, or any combination thereof. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0071] The computing system may include a client and a server. The client and the server may be remote from each other and interact via a communication network. The relationship between the client and the server arises from computer programs that run on the respective computers and have a client-server relationship with each other. For example, the server may transfer data, such as HTML pages, to the client device, for example, for the purpose of displaying data to and receiving user input from a user who interacts with the client device. Data generated at the client device, such as, by way of example, the result of a user interaction, may be received at the server from the client device.
[0072] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive and can be implemented in various combinations to achieve unique advantages. Since these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be construed in an illustrative rather than a limiting sense with respect to the subject matter defined by the claims. Additionally, the provision of the examples described herein and clauses expressed as "such as", "including", etc. should not be construed as limiting the subject matter of the claims to the specific examples; rather, these examples are intended to illustrate only one of many possible embodiments. Further, the same reference numerals in different figures may identify the same or similar elements.
Claims
1. A method for managing one or more local processing units in a server computing device of a distributed cloud platform, the method comprising: receiving, by one or more processors, one or more metrics associated with the one or more local processing units for executing a workload; generating, by the one or more processors, one or more predictions for one or more states of the one or more local processing units based on the one or more metrics using a machine learning model deployed on one or more accelerators in the server computing device; as well as The one or more states of the one or more local processing units are adjusted, by the one or more processing units, based on the prediction.
2. The method of claim 1, wherein the one or more metrics include at least one of: power utilization per processing unit core, power consumption per application running on a processing unit core, a number of enabled processing unit core C-states, a number of enabled processing unit core P-states, or a number of instructions per cycle being processed by a processing unit.
3. The method according to claim 1 or 2, wherein the one or more local processing units include at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field programmable gate array (FPGA). 4 . The method according to claim 1 , wherein the one or more accelerators include at least one of a tensor processing unit (TPU) or a wafer-scale engine (WSE).
5. The method according to one of claims 1 to 4 also includes training the machine learning model locally in the server computing device using the one or more metrics by the one or more processors.
6. The method according to one of claims 1 to 5 also includes receiving the machine learning model by the one or more processors, wherein the machine learning model is externally pre-trained on a decentralized service management and orchestration (SMO) platform.
7. The method of one of claims 1 to 6, wherein adjusting the one or more states comprises adjusting at least one of a frequency, a voltage, a power, a C-state, a P-state, or a sleep state of the one or more local processing units.
8. The method of one of claims 1 to 7, wherein adjusting the one or more states comprises adjusting the one or more states of a group of the one or more local processing units.
9. The method according to one of claims 1 to 8, wherein the workload comprises at least one of a Radio Access Network RAN function, an Access and Mobility Management Function AMF, a User Plane Function UPF or a Session Management Function SMF.
10. A system comprising: one or more processors; as well as One or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for managing one or more local processing units in a server computing device of a distributed cloud platform, the operations comprising: receiving one or more metrics associated with the one or more local processing units for executing a workload; generating one or more predictions for one or more states of the one or more local processing units based on the one or more metrics using a machine learning model deployed on one or more accelerators in the server computing device; and The one or more states of the one or more local processing units are adjusted based on the prediction.
11. The system of claim 10, wherein the one or more metrics include at least one of: power utilization per processing unit core, power consumption per application running on a processing unit core, a number of enabled processing unit core C-states, a number of enabled processing unit core P-states, or a number of instructions per cycle being processed by a processing unit.
12. The system according to claim 10 or 11, wherein the one or more local processing units include at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field programmable gate array (FPGA).
13. The system of one of claims 10 to 12, wherein the one or more accelerators include at least one of a tensor processing unit (TPU) or a wafer-scale engine (WSE).
14. The system of one of claims 10 to 13, wherein the operations further comprise training the machine learning model locally in the server computing device using the one or more metrics.
15. The system according to one of claims 10 to 13, wherein the operation further comprises receiving the machine learning model, wherein the machine learning model is externally pre-trained on a decentralized service management and orchestration (SMO) platform.
16. The system of one of claims 10 to 15, wherein adjusting the one or more states comprises adjusting at least one of a frequency, a voltage, a power, a C-state, a P-state, or a sleep state of the one or more local processing units.
17. The system of one of claims 10 to 16, wherein adjusting the one or more states comprises adjusting the one or more states of a group of the one or more local processing units.
18. The system according to one of claims 10 to 17, wherein the workload comprises at least one of a Radio Access Network RAN function, an Access and Mobility Management Function AMF, a User Plane Function UPF or a Session Management Function SMF.
19. A non-transitory computer-readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for managing one or more local processing units in a server computing device of a distributed cloud platform, the operations comprising: receiving one or more metrics associated with the one or more local processing units for executing a workload; generating one or more predictions for one or more states of the one or more local processing units based on the one or more metrics using a machine learning model deployed on one or more accelerators in the server computing device; as well as The one or more states of the one or more local processing units are adjusted based on the prediction.
20. The non-transitory computer-readable medium of claim 19, wherein the operations further comprise training the machine learning model locally in the server computing device using the one or more metrics.