Energy savings through flexible Kubernetes cabin capacity selection during horizontal cabin automatic expansion and contraction (HPA)
By using an AI/ML-driven horizontal cabin auto-expansioner in a Kubernetes environment to dynamically select cabin capacity and expansion/shrinkage strategies, the energy waste caused by fixed cabin capacity is solved, achieving energy savings and performance optimization.
Patent Information
- Application Number
- CN202380098643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-10-30
- Publication Date
- 2025-12-30
AI Technical Summary
In a Kubernetes environment, horizontal bay autoscaling (HPA) results in a system capacity exceeding the actual traffic load due to the fixed bay capacity, leading to energy waste.
The system employs an AI/ML-based Horizontal Cabin Automatic Expander (HPA) that measures current flow demand and predicts future demand by allocating resources and measuring capacity performance of the receiving cabin. It then selects the optimal cabin capacity and expansion/contraction strategy to provide fine-grained expansion/contraction commands and optimize energy consumption.
Significantly reduce energy consumption and operating costs of cloud data centers, lower carbon dioxide emissions, improve resource utilization, and meet the performance requirements of current and future traffic demands.
Smart Images

Figure CN121241333A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present description relates to energy saving by flexible Kubernetes pod capacity selection during horizontal pod autoscaling (HPA), and methods of using the same. BACKGROUND
[0002] Open-Radio Access Network (O-RAN) advocates a virtualized RAN where decomposed components are connected via open interfaces and optimized by intelligent controllers. O-RAN networks can be built with multi-vendor interoperable components and can be programmatically optimized through centralized abstraction layers and data-driven closed-loop control. The O-RAN architecture includes Service Management and Orchestration (SMO), which handles the orchestration, management, and automation aspects of O-RAN. SMO supports and manages a multi-vendor RAN environment.
[0003] Various elements of the O-RAN architecture, such as O-Radio Unit (O-RU), O-Distributed Unit (O-DU), O-Centralized Unit (O-CU), and Near-Real-Time RAN Intelligent Controller (Near-RT RIC), can be deployed in the cloud and physical environments. Components of the O-RAN architecture can also be deployed on Kubernetes clusters. In Kubernetes-based cloud applications, autoscaling allows for optimal allocation of resources to applications based on their current resource consumption. Vertical scaling (VS) and horizontal scaling (HS) of virtual Radio Access Networks (RAN slices), including dynamic instantiation and termination of on-demand RAN slices, enable resource allocation to be adjusted based on demand changes. In Kubernetes, workload resources, such as Deployments or StatefulSets, can be updated, scaling the workload to match demand. Horizontal scaling refers to deploying more pods. For Kubernetes, vertical scaling refers to assigning more resources (e.g., memory or CPU) to a running pod for the workload. Workload resources can be automatically scaled down in response to a drop in load when the number of pods is above a configured minimum. Horizontal pod autoscaling is not applicable to objects that cannot be scaled, such as DaemonSets.
[0004] However, due to fixed pod capacity, system capacity can exceed actual traffic load. This can result in energy waste. SUMMARY
[0005] In at least one embodiment, a method for conserving energy through flexible Kubernetes pod capacity selection during horizontal pod autoscaling (HPA) includes implementing an artificial intelligence / machine learning (AI / ML) based horizontal pod autoscaler (HPA). Performance metrics regarding resource allocation and capacity of a pod are received at the HPA. Current traffic demand is measured and future traffic demand is predicted relative to current system capacity. Pod capacity and scaling is selected in terms of number of pods and pod capacity version based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity to provide optimal performance for current traffic demand and future traffic demand according to pod capacity class. Scaling commands are generated for the selected pod capacity and the selected scaling to provide fine-grained scaling according to pod capacity class for optimizing energy consumption. The pod is scaled according to pod capacity class to meet current traffic demand and future traffic demand based on the scaling commands.
[0006] In at least one embodiment, an apparatus includes a KPI predictor configured to receive performance metrics regarding resource allocation and capacity of a pod, measure current traffic demand, and predict future traffic demand relative to current system capacity. A scaling decision is configured to select pod capacity and scaling in terms of number of pods and pod capacity version based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity to provide optimal performance for current traffic demand and future traffic demand according to pod capacity class. The scaling decision generates scaling commands for the selected pod capacity and the selected scaling to provide fine-grained scaling according to pod capacity class for optimizing energy consumption. A container manager is configured to receive the scaling commands and scale the pod according to pod capacity class to meet current traffic demand and future traffic demand.
[0007] In at least one embodiment, a non-transitory computer-readable medium having stored thereon computer-readable instructions for performing operations comprising: implementing an artificial intelligence / machine learning (AI / ML) based horizontal pod autoscaler (HPA). Receiving, at the HPA, performance metrics regarding resource allocation and capacity of a pod. Measuring current traffic demand and predicting future traffic demand relative to current system capacity. Selecting, based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity, a pod capacity and autoscaling in terms of a number of pods and a pod capacity version to provide optimal performance for current traffic demand and future traffic demand according to a pod capacity class. Generating an autoscaling command for the selected pod capacity and the selected autoscaling to provide fine-grained autoscaling according to the pod capacity class for optimizing energy consumption. Autoscaling, based on the autoscaling command, the pod according to the pod capacity class to meet the current traffic demand and the future traffic demand. BRIEF DESCRIPTION OF DRAWINGS
[0008] Aspects of the disclosure can be best understood from the following detailed description when read with the accompanying drawings. It is noted that, in accordance with the spirit of the industry customs, the various features of the drawings are not drawn to scale. In fact, the size of each feature can be exaggerated or reduced for the purpose of facilitating discussion.
[0009] Figure 1 Horizontal pod autoscaling (HPA) in Kubernetes is shown in accordance with at least one embodiment.
[0010] Figure 2 A traffic curve demonstrating energy waste cases of current HPA is illustrated.
[0011] Figure 3 Horizontal pod autoscaling in Kubernetes applied to open central unit (O-CU) 300 is illustrated in accordance with at least one embodiment.
[0012] Figure 4 A workflow for O-CU-CP or O-CU-UP horizontal autoscaling is illustrated in accordance with at least one embodiment.
[0013] Figure 5 Cloud resource energy saving via Non-RT RIC is illustrated in accordance with at least one embodiment.
[0014] Figure 6 Horizontal scaling of network functions (NFs) is illustrated in accordance with at least one embodiment.
[0015] Figure 7 Horizontal scaling of network function (NF) orchestration is illustrated in accordance with at least one embodiment.
[0016] Figure 8FIGURE illustrates horizontal pod auto-scaling (HPA) according to at least one embodiment.
[0017] Figure 9 FIGURE illustrates AI / ML based horizontal pod auto-scaling with flexible pod selection according to at least one embodiment.
[0018] Figure 10 FIGURE illustrates an architecture for providing an AI / ML based architecture for horizontal auto-scaling with flexible pod selection according to at least one embodiment.
[0019] Figure 11 FIGURE illustrates AI / ML based HPA with flexible pod selection for O-CU according to at least one embodiment.
[0020] Figure 12 FIGURE illustrates AI / ML based HPA with flexible pod selection for O-Cloud according to at least one embodiment.
[0021] Figure 13 is a flow diagram of a method to save energy through flexible Kubernetes pod capacity selection during horizontal pod auto-scaling (HPA) according to at least one embodiment.
[0022] Figure 14 is a high-level functional block diagram of a processor-based system according to at least one embodiment. DETAILED DESCRIPTION
[0023] Embodiments described herein describe examples for implementing different features of the subject matter provided. To simplify the present disclosure, examples of components, values, operations, materials, arrangements, and the like will be described below. Of course, these are merely examples and are not intended to limit the present disclosure. Other components, values, operations, materials, arrangements, and the like are also contemplated by the present disclosure. For example, in the description below, where a first feature is formed over or on a second feature includes embodiments where the first and second features are in direct contact, as well as embodiments where additional features are formed between the first and second features such that the first and second features are not in direct contact. Further, the present disclosure repeatedly references reference numerals and / or letters in various examples. Such repetition is for the sake of brevity and does not indicate a relationship or commonality between the various embodiments and / or configurations being discussed.
[0024] Moreover, spatial or directional terms, such as "below," "above," "lower," "upper," and the like can be used in this text and are intended to be used for ease of description to describe the orientations of elements or features in the figures. The spatial or directional terms are intended to encompass different orientations of the device in addition to the orientation depicted in the figures. The spatial or directional terms are used in this text for ease of description to describe the orientations of elements or features in the figures. It is not intended to limit the descriptions to the precise configurations or arrangements shown in the figures. Therefore, terms such as "below," "above," "lower," "upper," and the like can be understood as not necessitating a strict directional or spatial relationship.
[0025] As used in the specification and related drawings, the terms "user equipment," "mobile station," "mobile stations," "mobile device," "subscriber station," "subscriber device," "access terminal," "terminal," "handset," and similar terminology can refer to a wireless device used by a user or subscriber of a wireless communication service to receive or convey data, control, voice, video, audio, gaming, data streams, or signaling streams. The above terms can be used interchangeably in the present specification and related drawings. The terms "access point," "base station," "Node B," "Evolved Node B (eNode B)," "Next Generation Node B (gNB)," "Enhanced gNB (en-gNB)," "Home Node B (HNB)," "Home Access Point (HAP)," and similar terminology can refer to a wireless network component or device that serves and receives data, control, voice, video, audio, gaming, data streams, or signaling streams from a UE.
[0026] In at least one embodiment, a method for conserving energy through flexible Kubernetes pod capacity selection during horizontal pod autoscaling (HPA) includes implementing an artificial intelligence / machine learning (AI / ML)-based horizontal pod autoscaler (HPA); receiving, at the HPA, performance metrics regarding resource allocation and capacity of a pod; measuring current traffic demand and predicting future traffic demand relative to current system capacity; selecting a pod capacity in terms of a number of pods and a pod capacity version and a scale based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity, thereby providing optimal performance for current traffic demand and future traffic demand according to a pod capacity class; generating scale commands for the selected pod capacity and the selected scale to provide fine-grained scaling for optimizing energy consumption according to the pod capacity class; and scaling the pod to meet the current traffic demand and the future traffic demand according to the pod capacity class based on the scale commands.
[0027] Embodiments described herein provide a method that provides one or more advantages. For example, energy is saved in a cloud data center, thereby significantly reducing Operating Expense (OPEX) and reducing carbon dioxide emissions. In a 5G cellular network, 20-40% of OPEX comes from energy bills, and 20-30% of energy consumption occurs in a cloud data center that deploys O-DU, O-CU, core network, applications, and management services. In addition, according to statistics, 2% of global carbon dioxide emissions come from mobile networks, which is a considerable number. Therefore, reducing carbon dioxide emissions is an important goal.
[0028] Figure 1 FIG. illustrates horizontal pod autoscaling (HPA) in Kubernetes 100, according to at least one embodiment.
[0029] In Figure 1 In an embodiment, metrics 110 are provided to a horizontal pod autoscaler (HPA) 120. Metrics 110 indicate the busy degree of the system, for example, indicating traffic. For example, metrics 110 can indicate the number of RRC connections, the number of active / inactive UEs, the number of DRBs, the average throughput, and so on. Horizontal pod autoscaler (HPA) 120 determines the scaling of the number of pods in pod deployment 130 bound to deployment / replication controller (RC) 140 based on metrics 110. HPA 120 checks metrics 110 collected from the system against preconfigured thresholds. In response to the current value being higher than the specified threshold, HPA 120 attempts to increase the number of pods. HPA 120 is also able to scale down the number of pods in pod deployment 130. HPA 120 communicates with deployment / RC 140 to achieve automatic scaling in and scaling out of Kubernetes to meet system requirements, for example, the configuration in pod deployment 130 and the number of pods are scaled in and scaled out.
[0030] By default, Kubernetes supports CPU-based and memory-based pod autoscaling. However, users are able to configure the HPA 120 to scale based on metrics, such as custom metrics or external metrics. During horizontal autoscaling, the CPU and memory allocated to the pods in the pod deployment 130 is fixed (e.g., same application and same capacity, such as 6 gigabytes per second (GB / s)), and thus, the capacity of the system is scaled up and scaled down in fixed steps to follow the trend of actual traffic demand. However, according to at least one embodiment, for the pod deployment 130, pods with different configurations are pre-generated and pre-compiled. For example, there are pod configuration 1, pod configuration 2, pod configuration 3, etc., where pod 1 132 is 2 GB / s, pod 2 134 is 4 GB / s, pod 3 136 is 6 GB / s, and so on, including pod n 138 as well.
[0031] Figure 2 A traffic curve 200 is illustrated that demonstrates the energy waste situation of current HPA.
[0032] In Figure 2 , the site traffic level 210 is plotted from "no traffic" to "maximum traffic" over time 220. The curve 230 indicates real-time traffic over time. The scale threshold 240 is shown at the point in the data packet transmission graph. Due to the fixed step, most of the time, the actual capacity 250 provided by the system is greater than the actual traffic demand 260, i.e., the allocated CPU and memory resources are greater than the specified resources for use, resulting in energy waste of the cloud data center.
[0033] The system capacity 250 is shown for data packet transmission at different levels. Thus, due to the fixed pod capacity, the system capacity 250 increases to be greater than the actual traffic load 260. This results in energy waste. In response to the capacity increase, the HPA continuously increases the number of pods to scale the capacity 250 horizontally to meet the demand from traffic. The curve 230 indicates real-time traffic over time. For example, the traffic is low 270 in the morning, but gradually increases. Thus, the system initially uses one pod, providing a throughput of 6 GB per second. As the traffic increases over time, the number of parallel pods also increases.
[0034] In Figure 2 , the capacity 250 of the system is higher than the actual traffic demand 260, which results in energy waste. For example, even though the traffic only uses half of the capacity from the server, the server is still running. Thus, as shown in Figure 2 , half of the energy from the server is wasted.
[0035] Figure 3Figure illustrates horizontal pod scaling in Kubernetes 300 applied to an Open Centralized Unit (O-CU) 300, according to at least one embodiment.
[0036] In Figure 3 In the O-CU 310 captures performance metrics 320 in a performance matrix. The performance metrics 320 include the number of radio resource control (RRC) connections, the number of active / inactive user equipment (UE), the number of data radio bearers (DRB), average throughput, and so on. In Figure 3 In the pods to scale are the O-CU-CP pod 330 and the O-CU-UP pod 332. The O-CU-CP pod 330 and the O-CU-UP pod 332 are allocated fixed CPU and memory resources and support a fixed capacity, for example, 1250 UEs per O-CU-CP pod 330 and 6 Gbps throughput per O-CU-UP pod 332. The capacity of the O-CU-CP pod 330 or the O-CU-UP pod 332 is increased / decreased in fixed steps, which results in energy waste when the system capacity is higher than the actual traffic demand of the network. External performance metrics 320 from the O-CU 310 are collected by the HPA 340 for making scaling decisions based on configured custom thresholds.
[0037] Figure 350 illustrates data packet transmission traffic over time for 24 hours. Figure 350 illustrates scaling thresholds 352 and capacity 354 of the O-CU. Due to the fixed pod capacity (e.g., 6 Gbps steps), the capacity 354 of the O-CU is greater than the actual traffic load 356, which results in energy waste.
[0038] Figure 4 Figure illustrates a workflow 400 for O-CU-CP or O-CU-UP horizontal autoscaling, according to at least one embodiment.
[0039] In Figure 4 In the capacity related metrics are scraped (410) by the CU 412 (e.g., one or more of a CU Control Plane (CU-CP) or a CU-User Plane (CU-UP)) for subscriber count, throughput, DRB count, and so on.
[0040] The event monitor 420 (Prometheus) receives capacity related metrics (410) and provides queries (422) via a metrics server 430 (e.g., Prometheus adapter custom metrics server). The metrics are output to a custom metrics API extension 450 (440) that is controlled by a Kubernetes API server 452. A horizontal pod autoscaler (HPA) 460 queries the CM API 450 for capacity count data (462), e.g., subscriber count, throughput, DRB count, etc. The HPA 460 references a scaling policy 464 and performs horizontal scaling (470) based on the capacity count data and the scaling policy 464.
[0041] Figure 5 Cloud resource energy saving via Non-RT RIC 500 is illustrated in accordance with at least one embodiment.
[0042] In Figure 5 In
[0043] The O-Cloud 520 includes an interface management service (IMS) 522 that is coupled to the SMO 510 via an O2 interface 530. A deployment management service (DMS) 524 is coupled to the SMO 510 via an O2 interface 532. A Near-RT RIC 526 and an E2 node 528 are coupled to the SMO 510 via an O1 interface 534. The Near-RT RIC 526 is coupled to the Non-RT RIC 516 via an Al interface 550. The E2 node 528 is coupled to the Near-RT RIC 526 via an E2 interface 552. An O-RU 560 is coupled to the E2 node 528 of the O-Cloud 520 via an open fronthaul (FH) management (M) plane interface 554. The SMO 510 is coupled to the O-RU 560 via an open fronthaul (FH) M plane interface 556.
[0044] Currently, the control of horizontal scaling / expansion of O-Cloud 520 involves SMO 510 (FOCOM 512 / NFO514), which receives guidance for optimizing the energy consumption of various resources of O-Cloud 520 and generates actions for energy saving towards O-Cloud 520. However, in O-RAN, a fixed-step horizontal cabin autoscaling scheme is used. NFO 514 receives KPI metrics from E2 node 538 and O-RU 560 through O1 interface 534 and O2 interfaces 530, 532 for making horizontal scaling / expansion decisions.
[0045] The Non-RT RIC 516 collects O-Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data via O2 interfaces 530 and 532, and data from E2 node 528 via O1 interface 534. It trains and deploys an artificial intelligence / machine learning (AI / ML) model to generate guidance based on the data provided via O2 interfaces 530, 532, and O1 interface 534. Guidance for O-Cloud 520 or E2 node 528 is generated based on priority, load, and energy consumption. The Non-RT RIC 516 applies scaling commands to the pods deployed in O-Cloud 520.
[0046] O-Cloud 520 exposes O2 data (IMS 533 / DMS 534) to SMO 510 / Non-RT RIC 516, which performs changes recommended by SMO 510 / Non-RT RIC 516.
[0047] Horizontal scaling commands issued by SMO 510 are sent to DMS 524 via O2 interface 532 in O-Cloud 520. DMS 524 then controls the number of bays used for horizontal scaling / scaling via the Kubernetes API. New NF deployment units are created on new bays (horizontal scaling), or NF deployment units are removed. Bay resource allocation and capacity are fixed, leading to the same energy waste problem discussed earlier.
[0048] Figure 6 The illustration shows a horizontal extension of a network function (NF) 600 according to at least one embodiment.
[0049] exist Figure 6 This describes horizontal expansion orchestration use cases in O-RAN.WG6.ORCH-USE-CASES-v05.00. It covers orchestration use cases and specific standards for O-RAN / virtualized RAN.
[0050] The horizontal scaling network functionality includes building block identifiers and legend 610.
[0051] At Figure 6 , the SMO framework 620 includes a Network Function Orchestrator (NFO) 622 and an Operations and Maintenance (OAM) function 624. The O-Cloud 630 includes Deployment Management Services (DMS) 632, while the O-RAN 640 includes an O-RAN Managed Element (ME) 642.
[0052] At the build block identification and legend 610, the network demand exceeds the current NF capacity threshold, which will trigger a SMO capacity level scaling for the NF 612.
[0053] At the start build block 650, for the horizontal scaling of the NF, the SMO 620 determines what NF deployment scaling to use to increase the NF capacity based on the NF descriptor. The SMO 620 determines the new deployment criteria and selects a resource pool for the new Network Function (NF) deployment unit 652.
[0054] The NFO 622 sends a request message 654 to the DMS 632 via the O2 interface to create a new NF deployment unit for the NF horizontal scaling.
[0055] The DMS 632 creates the new NF deployment unit 656. The DMS 632 sends a message 658 to the NFO 622 via the O2 interface indicating that the NF deployment unit creation is complete (via the O2 interface). The OAM function 624 is able to send a configure NF message 660 to the O-RAN Managed Element 642 via the Ol interface for the O-RAN Managed Element 642 to configure the NF when changes are made. As such, the OAM function 624 in the SMO framework 620 communicates with the ME 642 via the Ol interface to configure the NF. The new NF deployment unit is put into traffic 662.
[0056] The end build block 670 indicates that the process is complete.
[0057] Figure 7 FIG. illustrates scaling in a Network Function (NF) orchestration 700, according to at least one embodiment.
[0058] At Figure 7 , the SMO framework 720 includes a Network Function Orchestrator (NFO) 722 and an Operations and Maintenance (OAM) function 724. The O-Cloud 730 includes Deployment Management Services (DMS) 732, while the O-RAN 740 includes an O-RAN Managed Element (ME) 742.
[0059] At building block identification with legend 710, network demand falls below the current NF capacity threshold, which will trigger a horizontal shrink of the SMO capacity level of 712 for the NF.
[0060] At start building block 750, for a horizontal shrink of the NF, SMO 720 determines the new NF deployment specification and selects the NF deployment unit to terminate 752. OAM function 724 sends a configuration (NF) message (with deployment ID) 754 to ME 742 via Ol interface to remove the NF deployment unit.
[0061] ME 742 responds with an acknowledge NF deployment unit removal message 756 (with deployment ID) via Ol interface. ME 742 removes the NF deployment unit from traffic 758.
[0062] OAM function 724 waits for ME 742 to remove the NF deployment unit from traffic 760.
[0063] ME 742 sends a notification message 762 to OAM function 724 that the traffic in that NF has been exhausted. NFO 722 sends a delete message 764 to DMS 732 via 02 interface to delete the NF deployment unit (identified by deployment ID). DMS 732 sends an acknowledge message 766 to NFO 722 via 02 interface that the deletion of the NF deployment unit is complete.
[0064] End building block 770 indicates the end of the process.
[0065] Figure 8 Fig. 8 illustrates a horizontal pod auto-scaling (HPA) 800, according to at least one embodiment.
[0066] In Figure 8 Instead of vertically scaling up / down the system with fixed pod resource allocation and capacity with fixed capacity step size, multiple pod versions with different resource allocation and capacity are pre-configured for HPA to select during horizontal auto-scaling. As Figure 8As shown, multiple cabin versions are pre-generated and compiled, such as cabin resource configuration 1 810, cabin resource configuration 2 812, cabin resource configuration 3 814, and so on. These cabins have specific configurations, enabling each cabin to support different capacities. For example, cabin resource configuration 1 812 can provide a throughput of 3 GB per second. The resource configuration tool can generate cabins with slightly higher capacities, such as 60 GB per second. Therefore, during scaling up and down, the controller does not need to increase the number of cabins of the same version (e.g., the same configuration or the same capacity), but can flexibly select the best version to meet the current needs. The cabin version is selected progressively based on demand, allowing the flow curve to be used to meet demand with finer granularity. This reduces wasted time and, consequently, wasted energy.
[0067] Figure 8 This illustrates three cabin versions implemented simultaneously during pre-deployment, such as cabin resource configuration 1 810, cabin resource configuration 2 812, cabin resource configuration 3 814, and so on. These versions are provided to the cloud 820, and the HPA 830 is able to select the optimal cabin version based on metrics 840. This is similar to vertical scaling, where allocating more or less CPU resources changes capacity. However, vertical scaling is a highly challenging problem across the industry. The embodiments described herein achieve dynamic capacity changes based on horizontal scaling, where resources are not dynamically allocated to CPUs, but rather different versions of cabins built previously are selected. The advantage is that the vendor generates cabins for a given capacity and pre-tests them. Cabins are then allocated from the HPA based on metrics. The optimal cabin version is selected to conform to the traffic curve, thus avoiding wasted energy usage.
[0068] The HPA 830 receives performance metrics 840 from Kubernetes or external sources of cloud applications and measures scaling conditions, such as current and future traffic demands relative to current system capacity. The HPA 830 then selects the optimal pod version with the best capacity match for horizontal scaling down / up. As a result, the system is able to track actual traffic demands at a finer granularity, rather than in fixed steps, where most of the time the system capacity exceeds actual traffic demand, leading to wasted energy—that is, allocated resources are not fully utilized.
[0069] Metric 840 is provided to HPA 830, and HPA 830 adds or removes cabins based on cabin capacity category. Cabins with different resource allocations can be configured. Deployment / Resource Controller 850 performs automatic horizontal shrinking / expansion. Figure 8The diagram illustrates two-module resource configuration 1 810. Two-module resource configuration 2 812 is also shown. A single-module resource configuration 3 814 is illustrated. For example, a module with resource configuration 1 810 can be equipped with the minimum CPU and memory resources, a module with resource configuration 2 812 can be equipped with more CPU and memory resources than resource configuration 1 810, and a module with resource configuration 3 814 can be equipped with the maximum CPU and memory resources.
[0070] An orchestrator (such as NFO) requests an O-Cloud entity (such as DMS) to horizontally scale new NF deployment units based on one or more of cabin configurations 810, 812, or 814, for example, by creating additional cabins(s). The number of cabins and their corresponding configuration indices are specified in the request.
[0071] Figure 860 shows data packet transmission versus time. The scaling threshold is shown 862. The AI / ML in Non-RT RIC predicts traffic patterns based on historical data and actively selects cabin capacity and scaling modes. AI / ML provides finer tracking of the traffic curve 864, thus avoiding wasted energy (866).
[0072] According to at least one embodiment, the Horizontal Pod Autoscaling (HPA) 830 differs from previous Vertical Autoscaling or combinations of horizontal and vertical autoscaling. In previous vertical scaling, Vertical Pod Autoscaling (VPA) automatically allocated CPU and memory resources based on traffic demand, which is impractical in most real-world applications. Due to various implementation reasons, most cloud applications cannot vertically scale up / down in capacity as resources dynamically increase / decrease at runtime.
[0073] In contrast, the Horizontal Cabin Auto-Scaling (HPA) 830 according to at least one embodiment allows different cabin versions 810, 812, and 814 with different resource allocations and capacities to be pre-configured in advance during the compilation or system initialization phases, and performance can be manually fine-tuned by application developers. HPA 830 simply selects the correct version with the most suitable capacity for runtime during auto-scaling, rather than relying on VPA to automatically allocate resources from the platform level.
[0074] During automatic scaling, the appropriate cabin capacity is selected to minimize CPU and memory resource consumption, thereby saving energy and ensuring network performance meets specifications without degradation. To achieve optimal performance, the HPA830 measures current scaling-related KPIs based on performance metrics collected at the current time (840) and predicts future changes in these KPIs. Furthermore, horizontal scaling / horizontal expansion decisions also consider the Quality of Service (QoS) specifications of different types of application traffic currently occurring and expected in the network to ensure compliance. Therefore, AI / ML-based solutions can deliver optimal scaling performance.
[0075] Figure 9 The illustration shows an AI / ML-based automatic expansion and contraction of a horizontal cabin with flexible cabin selection according to at least one embodiment.
[0076] exist Figure 9 In this process, metrics 910 are provided to AI / ML-based Horizontal Cabin Auto-Scaling (HPA) 920. Based on these metrics, AI / ML-based HPA 920 measures and predicts current and future scaling-related KPIs. AI / ML-based HPA 920 proactively selects cabin capacity and horizontal scaling / scaling in terms of cabin number and capacity version selection to achieve optimal performance and Quality of Service (QoS) guarantees. The request is then provided to a container manager with API 930, such as the Kubernetes API. The container manager with API 930 instructs deployment / RC 940 to scale cabins according to cabin capacity categories (e.g., add / remove cabins). Deployment / RC 940 performs horizontal scaling / scaling using cabin resource configurations 1950, 1952, 1954, etc.
[0077] Figure 10 The illustration shows an AI / ML-based architecture 1000 for providing horizontal autoscaling with flexible cabin selection, according to at least one embodiment.
[0078] exist Figure 10In the AI / ML-based HPA 1010, a KPI predictor 1012 includes an interface for receiving metrics 1002. Metric 1002 includes performance metrics 1003, which include resource allocation metrics 1004 and cabin capacity metrics 1005, such as one or more of the following: the number of Radio Resource Control (RRC) connections, the number of active / inactive User Equipment (UEs), the number of Data Radio Bearers (DRBs), or average throughput. Metric 1002 also includes measured current traffic demand 1006 and current system capacity 1007. The KPI predictor 1002 predicts future traffic demand 1013 based on performance metrics 1003, measured current traffic demand 1006, and current system capacity 1007.
[0079] The AI / ML-based HPA 1010 is also coupled to an AI / ML model 1020, which is used by the AI / ML HPA 1010's KPI predictor 1012 to predict traffic trends (e.g., future traffic demand 1013), and scaling decisions 1016 generate scaling commands 1018 in a proactive manner.
[0080] The AI / ML HPA 1010 interacts with the container manager 1030 (such as the Kubernetes API). The KPI predictor 1012 receives metrics 1002 and uses the AI / ML model 1020 to predict future traffic demands 1013 for horizontal scaling / scaling decisions, providing the prediction results 1014 to the scaling decision 1016. The scaling decision 1016 receives input from application type and Quality of Experience (QoE) specifications 1040, Quality of Service (QoS) related parameters and specifications 1042, scaling policies 1044, and hardware configuration 1046, where the scaling policy 1044 is used to improve performance, increase energy savings, reduce hardware resource utilization, maximize throughput, minimize latency, and meet latency budgets. The scaling decision 1016 then generates scaling commands 1018 and sends them to the container manager 1030 to change scaling, change policy versions, or add or delete pods via the container manager 1030. The AI / ML-based Architecture 1000 uses CU-CP and CU-UP to provide scaling.
[0081] The AI / ML HPA 1010 receives metrics 1002 from the CU. CUs have different configurations, and the container manager 1030 automatically adds or removes pods based on scaling commands 1018 to provide the optimal pod version that offers the best capacity. Future KPIs are predicted by a KPI predictor 1012, or based on a flexible choice using the AI / ML model 1020. The KPI predictor 1012 predicts which pod configuration will provide better performance, such as reduced energy consumption and reduced operating costs, based on future traffic estimates. The performance of multiple possibilities is evaluated. The predictions generated based on the evaluation 1013 are used by the scaling decision 1016 to determine the final scaling decision, such as adding pod configuration 1, pod configuration 2, pod configuration 3, etc.
[0082] AI / ML model selection 1020 can include linear regression models 1022, feedforward neural network (FNN) / convolutional neural network (CNN) models 1024, long short-term memory (LSTM) models 1026, and so on. The AI / ML model selected from AI / ML model selection 1020 is based on the performance and capacity conditions to be achieved. KPI predictor 1012 generates prediction results 1014 based on the impact on KPIs, such as how adding a new cabin with a specific version / configuration differs from adding different cabins with different configurations. Scaling decision 1016 generates scaling commands 1018 to provide optimal performance for current traffic demand 1006 and future traffic demand 1013 according to cabin capacity category, wherein scaling commands 1018 are based on the selection of cabin capacity and scaling in terms of cabin quantity and cabin capacity version, based on the current traffic demand 1006 measured relative to the current system capacity 1007 and the predicted future traffic demand 1003. Scaling decision 1016 generates a scaling command 1018 for the selected cabin capacity and the selected scaling to provide fine-grained scaling based on cabin capacity category for energy consumption optimization. This effect includes the performance provided by the container manager 1030 via the prediction result 1014 achieved by the scaling command 1018 from scaling decision 1016.
[0083] The chosen AI / ML model 1014 depends on the flow curve. If the flow curve is linear, the flow is easy to predict, and a linear regression model 1022 can be used to predict it. In some scenarios, the flow is not perfectly linear; for example, the curve fluctuates or includes significant variations, making prediction more difficult. In these cases, the flow is not easy to predict, and a more advanced module, such as LSTM 1026, is required. The choice of AI / ML model 1020 also depends on the availability of data for training the AI / ML model 1020. If there is less data for training the model, a simpler model, such as linear regression 1022, is used. If there is more data for training the AI / ML model 1020, a more advanced model can be used. For example, LSTM 1026 is significantly more complex than linear regression 1022.
[0084] One method for selecting the AI / ML model 1020 is through manual selection. Operators can use manual selection to provide KPI predictors 1012 for certain cell sites. The system trains the AI / ML model 1020, and then the operator manually selects the best model. However, given that there are 50,000 cells in the network, and at least some cells have slightly different conditions, manual selection is not feasible. In this case, software or the system automatically selects the AI / ML model 1020.
[0085] The AI / ML-based HPA 1010 includes an AI / ML scaling decision 1016, which receives current metrics 1015 and prediction results 1014 (such as prediction metrics / KPIs) to make a final automatic scaling decision on cabin capacity and scaling mode based on various factors. The scaling decision 1016 also uses application type and Quality of Experience (QoE) specifications 1040, application Quality of Service (QoS) specifications 1042, operator-provided scaling strategies 1044 (e.g., optimal performance, optimal energy savings, minimum hardware resource utilization, maximum throughput, minimum latency, meeting latency budgets, etc.), and hardware configuration 1046 to provide input to the scaling decision 1016. The scaling strategy 1044 is based on conditions or parameters received from the operator, such as whether the scaling strategy is more aggressive or more conservative in order to provide optimal performance and minimize the risk of insufficient resources to meet traffic demands. However, while aggressive scaling can achieve optimal performance, it can lead to resource waste; while conservative scaling can provide optimal energy consumption, it risks sacrificing some performance. Cloud operators can adjust service strategies that influence scaling decisions.
[0086] Scaling decision 1016 determines which pod selection provides optimal performance. Scaling decision 1016 provides scaling commands 1018 (e.g., horizontal shrink / horizontal expand commands) to container manager 1030 to proactively add / remove pods with capacity version selections based on pod capacity categories, thereby meeting current and future traffic demands. Previously, pods with fixed configurations (e.g., 6 GHz) were implemented, and pods were simply added or removed. According to at least one embodiment, pods with different configurations are pre-generated and pre-compiled. For example, there are pod configuration 1, pod configuration 2, pod configuration 3, etc., where pod 1 is 2 GB / s, pod 2 is 4 GB / s, pod 3 is 6 GB / s, and so on.
[0087] According to at least one embodiment, a horizontal cabin autoscaler (HPA) based on artificial intelligence / machine learning (AI / ML) is implemented. This AI / ML-based HPA includes a non-real-time radio access network intelligent controller (Non-RT RIC), which uses rApps to apply scaling commands to cabins deployed in an open cloud (O-Cloud) system. Performance metrics include those obtained by the Non-RT RIC through an O2 interface collecting O-Cloud fault, configuration, billing, performance, and security (FCAPS) data and through an O1 interface collecting E2 node data. The Non-RT RIC trains and deploys AI / ML models to generate scaling guidance for O-Cloud or E2 nodes based on priority, load, energy consumption, and quality of service specifications. Training and deploying the AI / ML models by the Non-RT RIC includes training and deploying at least one of the following: a linear regression model, a feedforward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model. Performance metrics related to resource allocation and cabin capacity include receiving one or more of the following: the number of Radio Resource Control (RRC) connections, the number of active / inactive User Equipment (UEs), the number of Data Radio Bearers (DRBs), or average throughput. Performance metrics related to resource allocation and cabin capacity are received at the HPA. Current traffic demand is measured relative to the current system capacity, and future traffic demand is predicted. Cabin capacity and scaling up / down are selected in terms of cabin quantity and capacity version based on the measured current traffic demand and predicted future traffic demand relative to the current system capacity, thereby providing optimal performance for current and future traffic demand according to cabin capacity category. Cabin capacity and scaling up / down are selected in terms of cabin quantity and capacity version selection based on the measured current traffic demand. Predicted future traffic demand includes using fine-grained tracking of traffic demand to match current system capacity with actual traffic demand, thereby matching resource utilization with resource demand. Scaling up / down commands are generated for the selected cabin capacity and selected scaling up / down to provide fine-grained scaling up / down according to cabin capacity category for optimizing energy consumption. Based on data received from application type and Quality of Experience (QoE) application specifications, Quality of Service (QoS) related configurations and specifications, scaling strategies for improving performance, increasing energy savings, minimizing hardware resource utilization, maximizing throughput, minimizing latency and meeting latency budgets, and hardware configuration, scaling decisions generate scaling commands. These scaling commands are sent to the Kubernetes API to scale the pod according to pod capacity categories to meet current and future traffic demands. Sending scaling commands to the Kubernetes API instructs the deployment / RC to scale the pod according to pod capacity categories.
[0088] Figure 11The illustration shows an AI / ML-based HPA with flexible cabin selection for O-CU 1100 according to at least one embodiment.
[0089] exist Figure 11 As described above, the AI / ML HPA 1110 collects performance metrics 1122 from the CU 1120, such as from the O-CU-CP compartment 1124 and the O-CU-UP compartment 1126. Performance metrics 1122 include the number of Radio Resource Control (RRC) connections, the number of active / inactive User Equipment (UE), the number of Data Radio Bearers (DRB), average throughput, and so on.
[0090] The AI / ML HPA 1110 measures and predicts current and future scaling-related conditions, and makes optimal scaling decisions by taking into account the various factors discussed above.
[0091] The AI / ML HPA 1110 provides scaling commands 1130 to the container manager 1140 (e.g., the Kubernetes API) to add / remove cabins based on cabin capacity categories. The container manager 1140 controls the horizontal scaling / expansion of the O-CU-CP cabin 1124 and O-CU-UP cabin 1126 in the O-CU 1120.
[0092] Figure 12 The illustration shows an AI / ML-based HPA with flexible cabin selection for O-Cloud 1200 according to at least one embodiment.
[0093] exist Figure 12 In this context, the auto-scaling solution is applied to O-Cloud 1220, where the same AI / ML HPA functionality is implemented as rApps in Non-RT RIC 1216.
[0094] The Non-RT RIC 1216 collects O-Cloud FCAPS data through the O2 interface and collects data from E2 node 1228 through O1 1234 to train and deploy AI / ML models. Based on priority, load, energy consumption, and quality of service specifications, it generates scaling guidelines for O-Cloud 1220 or E2 node 1228.
[0095] SMO 1210 is equipped with a Non-RT RIC 1216, a Joint O-Cloud Orchestration and Management (FOCOM) 1212, and a Network Functions Orchestrator (NFO) 1214. O-Cloud 1220 includes an Interface Management Service (IMS) 1222, which is coupled to SMO 1210 via O2 interface 1230. Deployment Management Service (DMS) 1224 is coupled to SMO 1210 via O2 interface 1232. Near-RT RIC 1226 and E2 node 1228 are coupled to SMO 1210 via O1 interface 1234. Near-RT RIC 1226 is coupled to Non-RT RIC 1226 via A1 interface 1250. E2 node 1228 is coupled to Near-RT RIC 1226 via E2 interface 1252. The O-RU 1260 is coupled to the E2 node 1228 of the O-Cloud 1220 via the Open Fronthaul (FH) Management (M) plane interface 1254. The SMO 1210 is coupled to the O-RU 1260 via the Open Fronthaul (FH) M plane interface 1256.
[0096] The SMO 1210 (FOCOM 1212 / NFO 1214) receives guidance from the Non-RT RIC 1216 to optimize the energy consumption of various resources of the O-Cloud 1220, and generates actions for energy saving of the O-Cloud 1220 based on commands to select add / remove bays by capacity category. Commands to select add / remove bays using capacity category are sent from the SMO 1210 to the O-Cloud 1220 via O2 interfaces 1230 and 1232.
[0097] The Non-RT RIC 1216 collects O-Cloud FCAPS data via O2 interfaces 1230 and 1232, and data from E2 node 1228 via O1 interface 1234. AI / ML models are trained and deployed to generate guidance based on O1 and O2 data. Guidance for O-Cloud 1220 or E2 node 1228 is generated based on priority, load, and energy consumption. O-Cloud 1220 exposes O2 data (IMS 1222 / DMS 1224) to SMO 1210 / Non-RT RIC 1216, which then executes the changes suggested by SMO 1210 / Non-RT RIC 1216.
[0098] Figure 13This is a flowchart 1300 of a method for saving energy during horizontal cabin automatic expansion and contraction (HPA) by means of flexible Kubernetes cabin capacity selection, according to at least one embodiment.
[0099] exist Figure 13 In step S1310, the method begins (S1302) with the implementation of an AI / ML-based Horizontal Cabin Autoscaler (HPA). Implementing the AI / ML-based HPA involves implementing a Non-Real-Time Radio Access Network Intelligent Controller (Non-RT RIC). The Non-RT RIC uses rApps to apply scaling commands to cabins deployed in an Open Cloud (O-Cloud) system. Performance metrics received by the Non-RT RIC include O-Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data collected via the O2 interface, and E2 node data collected via the O1 interface. The Non-RT RIC trains and deploys an AI / ML model to generate scaling guidance for O-Cloud or E2 nodes based on priority, load, energy consumption, and quality of service specifications.
[0100] The HPA receives performance metrics regarding cabin resource allocation and capacity (S1314). Receiving performance metrics includes: performance metrics obtained by the Non-RT RIC through the O2 interface by collecting O-Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data and through the O1 interface by collecting E2 node data. The Non-RT RIC trains and deploys AI / ML models to generate scaling guidelines for O-Cloud or E2 nodes based on priority, load, energy consumption, and quality of service specifications. Training and deploying AI / ML models by the Non-RT RIC includes training and deploying at least one of the following: a linear regression model, a feedforward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory (LSTM) model. Performance metrics regarding resource allocation and cabin capacity include one or more of the following: the number of Radio Resource Control (RRC) connections, the number of active / inactive User Equipment (UE), the number of Data Radio Bearers (DRBs), or average throughput.
[0101] The system measures current traffic demand relative to current system capacity and predicts future traffic demand (S1318). Based on the measured current traffic demand and predicted future traffic demand relative to current system capacity, the system selects cabin capacity and scales up or down in terms of cabin quantity and capacity version, thereby providing optimal performance for current and future traffic demands according to cabin capacity category. Selecting cabin capacity and scaling up or down based on the measured current traffic demand and predicted future traffic demand includes: using fine-grained tracking of traffic demand to match system capacity with actual traffic demand, thereby matching resource utilization with resource demand.
[0102] Based on the measured current flow demand and the predicted future flow demand according to the current system capacity, the cabin capacity and scaling are selected in terms of cabin quantity and cabin capacity version, thereby providing optimal performance for current and future flow demands according to the cabin capacity category (S1322). Cabin capacity and scaling are used to track flow demand with fine granularity so that the current system capacity matches the actual flow demand, thereby matching resource utilization with resource demand.
[0103] Generate scaling commands for the selected cabin capacity and selected scaling to provide fine-grained scaling based on cabin capacity category for energy consumption optimization (S1326). The scaling commands are generated by scaling decisions based on data received from application type and Quality of Experience (QoE) application specifications, Quality of Service (QoS) related configurations and specifications, scaling strategies for improving performance, improving energy savings, minimizing hardware resource utilization, maximizing throughput, minimizing latency and meeting latency budgets, and hardware configuration.
[0104] Based on the cabin capacity category, use the scaling command to scale the cabin to meet current and future traffic demands (S1330). The scaling command is sent to the Kubernetes API to instruct the deployment / RC to scale the cabin according to the cabin capacity category.
[0105] Then, the process terminates (S1340). At least one embodiment provides a method for saving energy during horizontal cabin autoscaling (HPA) through flexible Kubernetes cabin capacity selection, the method comprising: implementing an artificial intelligence / machine learning (AI / ML) based horizontal cabin autoscaling (HPA); receiving performance metrics at the HPA regarding cabin resource allocation and capacity; measuring current traffic demand relative to current system capacity and predicting future traffic demand; selecting cabin capacity and cabin capacity version in terms of cabin number and cabin capacity version based on the measured current traffic demand and predicted future traffic demand relative to current system capacity, to provide optimal performance for current and future traffic demand according to cabin capacity category; generating scaling commands for the selected cabin capacity and selected scaling to provide fine-grained scaling according to cabin capacity category for optimizing energy consumption; and scaling cabins according to cabin capacity category based on scaling commands to meet current and future traffic demand.
[0106] Figure 14 This is a high-level functional block diagram of a processor-based system 1400 according to at least one embodiment.
[0107] In at least one embodiment, the processing circuitry provides energy savings through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). The processing circuitry uses processor 1402 to implement energy savings through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). The processing circuitry also includes a non-transitory computer-readable storage medium 1404 for energy savings through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). The non-transitory computer-readable storage medium 1404, among other functions, is encoded (i.e., stores) with instructions (i.e., computer program code) that are executed by processor 1402, causing processor 1402 to perform operations to save energy through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). Processor 1402 executes instructions 1406 (at least partially) representing an application that implements at least a portion of the methods described herein (hereinafter referred to as "the process and / or method") according to one or more embodiments.
[0108] Processor 1402 is electrically connected to non-transitory computer-readable storage medium 1404 via bus 1408. Processor 1402 is also electrically connected to input / output (I / O) interface 1410 via bus 1408. Network interface 1412 is also electrically connected to processor 1402 via bus 1408. Network interface 1412 is connected to a network, allowing processor 1402 and non-transitory computer-readable storage medium 1404 to be connected to external components via the network. Processor 1402 is configured to execute instructions 1406 encoded in non-transitory computer-readable storage medium 1404 to make processing circuitry available for performing at least a portion of a process and / or method. In one or more embodiments, processor 1402 may be a central processing unit (CPU), a multiprocessor, a distributed processing system, an application-specific integrated circuit (ASIC), and / or a suitable processing unit.
[0109] The processing circuitry includes an I / O interface. The I / O interface 1410 is connected to external circuitry. In one or more embodiments, the I / O interface 1410 includes a keyboard, keypad, mouse, trackball, touchpad, touchscreen, and / or arrow keys for transmitting information and commands to the processor 1402.
[0110] The processing circuitry also includes a network interface 1412 coupled to the processor 1402. The network interface 1412 allows the processing circuitry to communicate with a network 1414, to which one or more other computer systems are connected. The network interface 1412 may include a wireless network interface such as Bluetooth, Wi-Fi, Global Microwave Access Interoperability (WiMAX), General Packet Radio Service (GPRS), or Wideband Code Division Multiple Access (WCDMA); or a wired network interface such as Ethernet, Universal Serial Bus (USB), or IEEE 864.
[0111] The processing circuitry is configured to receive information via I / O interface 1410. The information received via I / O interface 1410 includes one or more instructions, data, design rules, cell libraries, and / or other parameters for processing by processor 1402. This information is transmitted to processor 1402 via bus 1408. The processing circuitry is also configured to receive information related to user interface (UI) 1420 via the I / O interface. This information is stored as UI 1420 in non-transitory computer-readable storage medium 1404 for use in network data / cabin expansion 1422.
[0112] In one or more embodiments, one or more non-transitory computer-readable storage media 1404 store instructions 1406 (in compressed or uncompressed form) that can be used to program a computer, processor, or other electronic device to perform the processes or methods described herein. The one or more non-transitory computer-readable storage media 1404 include electronic storage media, magnetic storage media, optical storage media, quantum storage media, and the like.
[0113] For example, non-transitory computer-readable storage media 1404 includes, but is not limited to, hard disk drives, floppy disks, optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic cards or optical cards, solid-state storage devices, or other physical media suitable for storing electronic instructions. In one or more embodiments using optical disks, one or more non-transitory computer-readable storage media 1404 include read-only compressed disc storage (CD-ROM), read-write compressed disc (CD-R / W), and / or digital video optical disc (DVD).
[0114] In one or more embodiments, non-transitory computer-readable storage medium 1404 stores at least a portion of a process or method configured to cause processor 1402 to perform energy-saving operations through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). In one or more embodiments, non-transitory computer-readable storage medium 1404 also stores information such as algorithms that facilitate the performance of at least a portion of the process and / or method for saving energy through flexible Kubernetes bay capacity selection during horizontal bay autoscaling (HPA). Thus, in at least one embodiment, processor 1402 executes instructions 1406 stored on one or more non-transitory computer-readable storage media 1404 to implement an artificial intelligence / machine learning (AI / ML) based horizontal bay autoscaling (HPA) 1438, thereby saving energy through flexible Kubernetes bay capacity selection during autoscaling. Radio access network (RAN) 1430 includes a Serving, Management and Orchestration (SMO) platform 1432. SMO platform 1432 includes a FOCOM / NFO 1434 and a non-real-time RAN intelligent controller (Non-RT RIC) 1436. HPA 1438 is implemented using the Non-RT RIC1436. HPA 1438 includes a KPI predictor 1440 that receives metrics 1460 regarding cabin resource allocation and capacity. HPA 1438 measures current traffic demand relative to current system capacity and predicts future traffic demand. KPI predictor 1440 predicts both current and future demand for the cabin. The Non-RT RIC1436 uses rApps to apply scaling commands to cabins deployed in the Open Cloud (O-Cloud) system. HPA 1438 receives metrics 1460, which include O-Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data collected via the O2 interface and data collected from E2 nodes 1450 via the O1 interface. The HPA1438 also includes scaling decision 1442, which generates scaling commands for selected cabin capacity and selected scaling based on inputs from KPI predictor 1440 and metrics 1460, thereby providing fine-grained scaling for energy consumption optimization based on cabin capacity category. Scaling decision 1442 selects cabin capacity and scaling based on measured current traffic demand and predicted future traffic demand relative to current system capacity, thus providing optimal performance for current and future traffic demand based on cabin capacity category. Metrics 1460 also includes data such as the number of RRC connections, the number of active / inactive UEs, the number of DRBs, average throughput, etc. Scaling decision 1442 also receives data on application type and QoE specifications, QoS-related parameters and specifications, scaling strategies, and hardware configuration.The Non-RTRIC 1436 trains and deploys AI / ML models to generate scaling guidance for O-Cloud or E2 nodes based on priority, load, energy consumption, and QoS specifications. The KPI predictor 1440 uses the AI / ML model 1444 to predict current and future cabin demand. The AI / ML model 1444 includes predictive models such as linear regression, FNN / CNN, LSTM, and more. RAN 1430 also includes Open Cloud (O-Cloud) 1446. O-Cloud 1446 includes Near-RT RIC 1448, one or more E2 nodes 1450, Interface Management Service (IMS) 1452, and Deployment Management Service (DMS) 1454. The scaling decision 1442 sends scaling commands to a container manager with an API 1456 (e.g., the Kubernetes API) to scale the cabin according to cabin capacity categories to meet current and future traffic demands. The container manager 1456 supports scaling bays based on scaling commands received from the HPA 1438. The deployment / replication controller (RC) 1458 executes scaling commands to add or remove bays. The display 1470 presents the user interface (UI) 1472. The UI 1472 is used to display network data, scaling, metrics 1474, and other such data to enable horizontal bay auto-scaling (HPA), thereby providing flexible Kubernetes bay capacity selection to save energy.
[0115] The embodiments described herein provide a method that offers one or more advantages. For example, it saves energy in cloud data centers, thereby significantly reducing operating costs (OPEX) and carbon dioxide emissions. In 5G cellular networks, 20-40% of operating costs come from energy bills, and 20-30% of energy consumption occurs in cloud data centers where ODUs, OCUs, core networks, applications, and management services are deployed. Furthermore, according to statistics, 2% of global carbon dioxide emissions come from mobile networks, which is a considerable figure. Therefore, reducing carbon dioxide emissions is an important objective.
[0116] One aspect of this description relates to a method [1] for saving energy through flexible Kubernetes cabin capacity selection during horizontal cabin autoscaling (HPA), the method comprising: implementing an artificial intelligence / machine learning (AI / ML) based horizontal cabin autoscaling (HPA); receiving performance metrics at the HPA regarding cabin resource allocation and capacity; measuring current traffic demand relative to current system capacity and predicting future traffic demand; selecting cabin capacity and cabin capacity version based on the measured current traffic demand and predicted future traffic demand relative to current system capacity, and selecting scaling to provide optimal performance for current and future traffic demand according to cabin capacity category; generating scaling commands for the selected cabin capacity and selected scaling to provide fine-grained scaling according to cabin capacity category for optimizing energy consumption; and scaling cabins according to cabin capacity category based on scaling commands to meet current and future traffic demand.
[0117] The method described in [1], wherein implementing AI / ML-based HPA includes implementing a Non-Real-Time Radio Access Network Intelligent Controller (Non-RT RIC), wherein the Non-RT RIC uses rApps to apply scaling commands to cabins deployed in an Open Cloud (O-Cloud) system, and wherein receiving performance metrics includes: the Non-RT RIC collecting O-Cloud fault, configuration, billing, performance, and security (FCAPS) data via the O2 interface and collecting E2 node data via the O1 interface to obtain performance metrics, and wherein the Non-RT RIC trains and deploys an AI / ML model to generate scaling guidance for O-Cloud or E2 nodes based on priority, load and energy consumption, and quality of service specifications.
[0118] The method described in any one of [1] to [2], wherein training and deploying an AI / ML model by a Non-RT RIC includes training and deploying at least one of the following: a linear regression model, a feedforward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
[0119] The method described in any one of [1] to [3], wherein scaling commands are generated by scaling decisions based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies and hardware configurations, and the scaling policies are used to improve performance, improve energy savings, minimize hardware resource utilization, maximize throughput, minimize latency and meet latency budgets.
[0120] The method of any one of [1] to [4], wherein receiving performance metrics of the cabin’s resource allocation and capacity at the HPA includes receiving one or more of the following: the number of radio resource control (RRC) connections, the number of active / inactive user equipment (UE), the number of data radio bearers (DRB), or the average throughput.
[0121] The method described in any one of [1] to [5], wherein sending a scaling command to the container manager includes sending a scaling command to the container manager to instruct the deployment / replication controller (RC) to scale the pod according to the pod capacity category.
[0122] The method described in any one of [1] to [6], wherein selecting cabin capacity and selecting scaling up or down based on the measured current flow demand and the predicted future flow demand includes: using fine-grained tracking of flow demand to match the current system capacity with the actual flow demand, thereby matching resource utilization with resource demand.
[0123] The present invention relates to an apparatus [8] comprising: a KPI predictor configured to receive performance metrics relating to cabin resource allocation and capacity, measure current flow demand relative to current system capacity, and predict future flow demand; a scaling decision configured to select cabin capacity and scaling in terms of cabin number and cabin capacity version based on the measured current flow demand and the predicted future flow demand relative to current system capacity, to provide optimal performance for current and future flow demand according to cabin capacity category, wherein the scaling decision generates scaling commands for the selected cabin capacity and selected scaling to provide fine-grained scaling for optimizing energy consumption according to cabin capacity category; and a container manager configured to receive scaling commands to scale cabins according to cabin capacity category to meet current and future flow demand.
[0124] The device described in [8] also includes a Non-Real-Time Radio Access Network Intelligent Controller (Non-RT RIC), which is configured to collect Open Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data via an O2 interface and E2 node data via an O1 interface; wherein the Non-RT RIC uses rApps to apply scaling commands received from scaling decisions to cabins deployed in the Open Cloud (O-Cloud) system; wherein performance metrics include performance metrics for the Non-RT RIC based on the Open Cloud Fault, Configuration, Billing, Performance, and Security (FCAPS) data collected via the O2 interface and the E2 node data collected via the O1 interface; and wherein the Non-RT RIC trains and deploys AI / ML models to generate scaling guidance for O-Cloud or E2 nodes based on priority, load and energy consumption, and quality of service specifications.
[0125] The device described in any one of [8] to [9], wherein the Non-RT RIC is configured to train and deploy an AI / ML model by training and deploying at least one of the following: a linear regression model, a feedforward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
[0126] The device described in any one of [8] to
[10] , wherein scaling decisions are configured to generate scaling commands based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies, and hardware configurations, the scaling policies being used to improve performance, improve energy savings, minimize hardware resource utilization, maximize throughput, minimize latency, and meet latency budgets.
[0127] The device described in any one of [8] to
[11] , wherein the KPI predictor is configured to receive performance metrics relating to cabin resource allocation and capacity by receiving one or more of the following: the number of radio resource control (RRC) connections, the number of active / inactive user equipment (UE), the number of data radio bearers (DRB), or the average throughput.
[0128] The device described in any one of [8] to
[12] , wherein the processor is configured to instruct the deployment / replication controller (RC) to scale the pod according to the pod capacity category by sending scaling commands to the Kubernetes API.
[0129] The device described in [8] to
[13] , wherein scaling decisions are configured to select cabin capacity and scaling based on measured current flow demand and predicted future flow demand in terms of cabin quantity and cabin capacity version selection: using fine-grained tracking of flow demand to match current system capacity with actual flow demand, thereby matching resource utilization with resource demand.
[0130] The present invention relates to a non-transitory computer-readable medium
[15] having stored thereon computer-readable instructions for performing operations including: implementing an artificial intelligence / machine learning (AI / ML) based horizontal cabin automatic expander (HPA); receiving performance metrics at the HPA regarding cabin resource allocation and capacity; measuring current flow demand relative to current system capacity and predicting future flow demand; selecting cabin capacity and selecting expansion based on the measured current flow demand and predicted future flow demand relative to current system capacity, to provide optimal performance for current and future flow demand according to cabin capacity category; generating expansion commands for the selected cabin capacity and selected expansion to provide fine-grained expansion according to cabin capacity category for optimizing energy consumption; and expanding and shrinking cabins according to cabin capacity category based on expansion commands to meet current and future flow demand.
[0131]
[15] describes a non-transitory computer-readable medium in which implementing an AI / ML-based HPA includes implementing a non-real-time radio access network intelligent controller (Non-RT RIC), wherein the Non-RT RIC uses rApps to apply scaling commands to pods deployed in an open cloud (O-Cloud) system, and wherein receiving performance metrics includes: acquiring performance metrics by the Non-RT RIC through the O2 interface to collect O-Cloud fault, configuration, billing, performance, and security (FCAPS) data and through the O1 interface to collect E2 node data, and wherein the Non-RT RIC trains and deploys an AI / ML model to generate scaling instructions for O-Cloud or E2 nodes based on priority, load and energy consumption, and quality of service specifications; and wherein training and deploying the AI / ML model by the Non-RT RIC includes training and deploying at least one of the following: a linear regression model, a feedforward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
[0132]
[15] to
[16] A non-transitory computer-readable medium wherein scaling commands are generated by scaling decisions based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies and hardware configurations, scaling policies used to improve performance, improve energy savings, minimize hardware resource utilization, maximize throughput, minimize latency and meet latency budgets.
[0133] The non-transitory computer-readable medium described in any one of
[15] to
[17] , wherein receiving performance metrics of resource allocation and capacity of the compartment at the HPA includes receiving one or more of the following: the number of radio resource control (RRC) connections, the number of active / inactive user equipment (UE), the number of data radio bearers (DRB), or the average throughput.
[0134] The non-transitory computer-readable medium described in any one of
[15] to
[18] , wherein sending a scaling command to the Kubernetes API includes sending a scaling command to the Kubernetes API to instruct the deployment / replication controller (RC) to scale the pod according to the pod capacity category.
[0135] The non-transitory computer-readable medium described in any one of
[15] to
[19] , wherein selecting cabin capacity and selecting scaling up or down based on measured current flow demand and predicted future flow demand includes: using fine-grained tracking of flow demand to match current system capacity with actual flow demand, thereby matching resource utilization with resource demand.
[0136] Independent instances of these programs can be executed or distributed across any number of independent computer systems. Therefore, while some steps have been described as being performed by a specific device, software program, process, or entity, this does not imply limitation. Those skilled in the art will understand that various alternative implementations exist.
[0137] Furthermore, those skilled in the art will recognize that the above-described techniques can be used in a variety of devices, environments, and situations. Although the embodiments have been described using specific language for structural features or methodological behavior, the subject matter defined in the appended claims is not necessarily limited to the specific features or behaviors described. Rather, these specific features and behaviors are disclosed only as exemplary forms for implementing the claims.
Claims
1. A method for conserving energy through flexible Kubernetes pod capacity selection during horizontal pod autoscaling (HPA), comprising: implementing an artificial intelligence / machine learning (AI / ML) based horizontal pod autoscaler (HPA); receiving, at the HPA, performance metrics regarding resource allocation and capacity of a pod; measuring current traffic demand and predicting future traffic demand relative to current system capacity; selecting pod capacity in terms of number of pods and pod capacity version and scaling based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity to provide optimal performance for the current traffic demand and the future traffic demand according to pod capacity class; generating scaling commands for the selected pod capacity and the selected scaling to provide fine-grained scaling for optimizing energy consumption according to pod capacity class; and scaling the pod according to pod capacity class to meet the current traffic demand and the future traffic demand based on the scaling commands.
2. The method of claim 1, wherein implementing the AI / ML-based HPA comprises: implementing a non-real-time radio access network intelligent controller (Non-RT RIC), wherein the Non-RT RIC applies the scaling commands to pods deployed in an open cloud (O-Cloud) system using rApps, and wherein receiving the performance metrics comprises: obtaining performance metrics by the Non-RT RIC collecting O-Cloud fault, configuration, charging, performance, security (FCAPS) data through an O2 interface and E2 node data through an O1 interface, and wherein the Non-RT RIC trains and deploys AI / ML models to generate scaling guidance for the O-Cloud or the E2 node based on priority, load, and energy consumption, and quality of service specifications.
3. The method of claim 2, wherein training and deploying an AI / ML model by the Non-RT RIC comprises: training and deploying at least one of: a linear regression model, a feed-forward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
4. The method of claim 1, wherein the scaling commands are generated from scaling decisions based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies for improving performance, improving energy conservation, minimizing hardware resource utilization, maximizing throughput, minimizing latency, and meeting latency budget, and hardware configurations.
5. The method of claim 1, wherein receiving, at the HPA, performance metrics related to resource allocation and capacity of a pod comprises: receiving one or more of: number of radio resource control (RRC) connections, number of active / inactive user equipment (UE), number of data radio bearers (DRBs), or average throughput.
6. The method of claim 1, wherein sending the scale command to a container manager comprises: sending the scaling commands to the container manager to instruct a replication controller (RC) to scale pods according to pod capacity class.
7. The method of claim 1, wherein selecting the pod capacity and selecting the scaling based on the measured current traffic demand and the predicted future traffic demand in the number of pods and pod capacity version selection comprises: using fine-grained tracking of traffic demand to match the current system capacity to actual traffic demand, thereby matching resource utilization to resource demand.
8. An apparatus, comprising: a KPI predictor configured to receive performance metrics regarding resource allocation and capacity of a pod, measure current traffic demand and predict future traffic demand relative to current system capacity; scaling decision configured to select pod capacity and scaling in terms of pod number and pod capacity version based on measured current traffic demand and predicted future traffic demand relative to current system capacity to provide optimal performance for the current traffic demand and the future traffic demand according to pod capacity classes, wherein the scaling decision generates a scaling command for selected pod capacity and selected scaling to provide fine-grained scaling for optimizing energy consumption according to the pod capacity classes; and and a container manager configured to receive the scaling command to scale pods according to pod capacity classes to meet the current traffic demand and the future traffic demand.
9. The apparatus of claim 8, further comprising: a Non-Real-Time Radio Access Network Intelligent Controller (Non-RT RIC) configured to collect Open Cloud Fault, Configuration, Accounting, Performance, and Security (FCAPS) data through an O2 interface and E2 node data through an O1 interface; wherein the Non-RT RIC applies the scaling command received from the scaling decision to pods deployed in an Open Cloud (O-Cloud) system using rApps; wherein the performance metrics include performance metrics for Non-RT RIC based on the Open Cloud Fault, Configuration, Accounting, Performance, and Security (FCAPS) data collected through the O2 interface and E2 node data collected through the O1 interface; and wherein the Non-RT RIC trains and deploys AI / ML models to generate scaling guidance for the O-Cloud or the E2 node based on priority, load, and energy consumption and quality of service specifications.
10. The device of claim 9, wherein the Non-RT RIC is configured to train and deploy AI / ML models by training and deploying at least one of: a linear regression model, a feed-forward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
11. The device of claim 8, wherein the scaling decision is configured to generate the scaling command based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies for improving performance, improving energy saving, minimizing hardware resource utilization, maximizing throughput, minimizing latency, and meeting latency budget, and hardware configurations.
12. The device of claim 8, wherein the KPI predictor is configured to receive performance metrics on resource allocation and capacity of pods by receiving one or more of: number of radio resource control (RRC) connections, number of active / inactive user equipment (UEs), number of data radio bearers (DRBs), or average throughput.
13. The device of claim 8, wherein the scaling decision is configured to send the scaling command to the container manager to instruct a replication controller (RC) to scale pods according to pod capacity classes.
14. The device of claim 8, wherein the scaling decision is configured to select the pod capacity and select the scaling based on the measured current traffic demand and the predicted future traffic demand in terms of the number of pods and pod capacity versions selection by using fine-grained tracking of traffic demand to match current system capacity to actual traffic demand to match resource utilization to resource demand.
15. A non-transitory computer readable medium having stored thereon computer readable instructions for performing operations comprising: implementing an artificial intelligence / machine learning (AI / ML) based horizontal pod autoscaler (HPA); receiving, at the HPA, performance metrics related to resource allocation and capacity of pods; measuring current traffic demand and predicting future traffic demand relative to current system capacity; selecting pod capacity and selecting scaling in terms of number of pods and pod capacity versions based on the measured current traffic demand and the predicted future traffic demand relative to current system capacity to provide optimal performance for the current traffic demand and the future traffic demand according to pod capacity classes; generating scaling commands for the selected pod capacity and the selected scaling to provide fine-grained scaling for optimizing energy consumption according to pod capacity classes; and scaling pods according to pod capacity classes to meet the current traffic demand and the future traffic demand based on the scaling commands. implementing a non-real-time radio access network intelligent controller (Non-RT RIC), wherein the Non-RT RIC applies the scaling commands to pods deployed in an open cloud (O-Cloud) system using rApps, and wherein receiving the performance metrics comprises: obtaining performance metrics by the Non-RT RIC collecting O-Cloud fault, configuration, accounting, performance, security (FCAPS) data through an O2 interface and collecting E2 node data through an O1 interface, and wherein the Non-RT RIC trains and deploys AI / ML models to generate scaling guidance for the O-Cloud or the E2 node based on priority, load, and energy consumption, and quality of service specifications; and 16. The non-transitory computer-readable medium of claim 15, wherein implementing AI / ML-based HPA comprises: wherein training and deploying AI / ML models by the Non-RT RIC comprises training and deploying at least one of: a linear regression model, a feed-forward neural network (FNN), a convolutional neural network (CNN) model, or a long short-term memory model.
17. The non-transitory computer readable medium of claim 16, wherein the scaling commands are generated by a scaling decision based on data received from application type and quality of experience (QoE) application specifications, quality of service (QoS) related configurations and specifications, scaling policies for improving performance, improving energy saving, minimizing hardware resource utilization, maximizing throughput, minimizing latency, and meeting latency budget, and hardware configurations. 18. The non-transitory computer-readable medium of claim 15, wherein receiving, at the HPA, performance metrics related to resource allocation and capacity of a pod comprises: Receiving one or more of the following: a number of radio resource control (RRC) connections, a number of active / inactive user equipments (UEs), a number of data radio bearers (DRBs), or an average throughput.
19. The non-transitory computer-readable medium of claim 15, wherein sending the scale command to a Kubernetes API comprises: Sending the scaling command to the Kubernetes API to instruct a deployment / replication controller (RC) to scale a pod according to the pod capacity category.
20. The non-transitory computer readable medium of claim 15, wherein selecting the pod capacity and selecting the scaling based on the measured current traffic demand and the predicted future traffic demand in the number of pods and pod capacity version selection comprises: Using fine-grained tracking of traffic demand to match the current system capacity to actual traffic demand, thereby matching resource utilization to resource demand.