Distributed monitoring and correlation system for telecom workloads in cloud-native environments
A distributed monitoring and correlation system with AI-driven decision engines optimizes network performance and resource utilization in cloud-native vRAN environments by decentralizing monitoring tasks across edge clusters, addressing latency and scalability issues in centralized systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Centralized monitoring systems in telecom networks introduce latency and scalability issues, hindering efficient fault detection and management in cloud-native distributed vRAN environments, which are critical for maintaining Service Level Agreements (SLAs).
A distributed monitoring and correlation system is implemented, deploying NEOs across edge clusters using Kubernetes operators and AI-driven decision engines for real-time monitoring and fault correlation, optimizing network performance and resource utilization.
The system decentralizes monitoring tasks, reducing latency and overhead, enabling efficient fault detection and performance optimization by leveraging cloud-native tools and AI for dynamic optimization.
Smart Images

Figure KR2025018314_15052026_PF_FP_ABST
Abstract
Description
DISTRIBUTED MONITORING AND CORRELATION SYSTEM FOR TELECOM WORKLOADS IN CLOUD-NATIVE ENVIRONMENTS
[0001] The proposed invention relates to cloud and network function virtualization and more particularly relates to a method and a system for managing and monitoring virtual Radio Access Networks (vRANs) in a communication system.
[0002] The telecommunications industry has recently witnessed a significant shift towards container-based microservice architectures, particularly among next-generation 5G and 6G telco vendors and operators. This paradigm shift addresses many challenges inherent in traditional monolithic architecture applications, offering improved scalability, flexibility, and efficiency. To fully leverage the benefits of microservices, it is essential to deploy these applications using technologies that align with the characteristics of microservices. Consequently, cloud-native container-runtime and container-orchestrator solutions have become the preferred deployment formats for microservice applications within telco products.
[0003] Several commercial 5G telecommunication network products, including network elements management systems (EMS), radio access central units (CU), radio access distributed units (DU), and 5G core network functions, are being redesigned to fit the microservice paradigm as containers. This transition is in alignment with the standards set forth by prominent 5G standardization bodies such as the 3rd Generation Partnership Project (3GPP) and the European Telecommunication Standards Institute (ETSI).
[0004] As telecom networks evolve towards more distributed and cloud-native architectures, traditional centralized monitoring systems are proving inadequate. These legacy systems introduce latency and scalability issues that are incompatible with the dynamic requirements of modern technologies like virtualized Radio Access Network (vRAN) and Open Radio Access Network (oRAN). The centralized approach to monitoring and orchestration results in increased latency, slower fault detection, and scalability challenges, which are not suitable for the low-latency, high-performance demands of cloud-native distributed vRAN environments. Maintaining strict Service Level Agreements (SLAs) in such environments is critical, and the inefficiencies of traditional systems pose significant obstacles.
[0005] In the context of distributed vRAN deployments, there is a substantial challenge in monitoring and correlating workload-specific performance and fault events in real-time. The reliance on centralized orchestration for these tasks exacerbates latency and scalability issues, hindering the ability to detect faults quickly and efficiently. This centralization is particularly problematic for the dynamic and low-latency requirements of cloud-native distributed vRAN environments.
[0006] Thus, it is desired to address the above-mentioned disadvantages, issues or other shortcomings or at least provide a useful alternative.
[0007] The principal object of the embodiments herein is to provide a system and method for distributed monitoring and correlation for telecom workloads in cloud-native environments.
[0008] Another object of the embodiments herein is to provide a distributed orchestration apparatus that creates and migrates a Network Element Orchestrator (NEO) across central Service Management and Orchestration (SMO) entity vRAN clusters or at individual vRANs.
[0009] Yet another object of the embodiments herein is to perform fault correlation and to determine the root cause of network issues in faults correlated with events associated with the cluster orchestrator and forward alarms correlating faults and performance statistics to the distributed orchestration
[0010] In an aspect, the objects are achieved by providing a method for managing and monitoring vRANs in a communication system. The method includes initializing by a distributed orchestration apparatus a plurality of vRANs and determining a plurality of parameters associated with the plurality of vRANs. Further, the method includes grouping by the distributed orchestration apparatus the plurality of vRANs into a plurality of clusters based on a plurality of parameters and creating by the distributed orchestration apparatus a plurality of NEOs at a central SMO entity within a vRAN cluster of a group of a plurality of vRANs or at individual vRANs. Each NEO of the plurality of NEOs manages and monitors the plurality of vRANs. Furthermore, the method includes monitoring by the distributed orchestration apparatus the plurality of parameters associated with the plurality of vRAN clusters and adjusting the number of NEOs in each of the plurality of vRAN clusters based on the plurality of parameters associated with the plurality of vRAN clusters.
[0011] In another aspect, the objects are achieved by providing a method for managing and monitoring vRANs by the cluster orchestrator. The method includes interfacing by a cluster orchestrator at least one NEO with a Container Infrastructure Service Management (CISM), wherein the cluster orchestrator is integrated with the NEO and the cluster orchestrator is at least one of a central SMO entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN cluster. Further, the method includes polling by the cluster orchestrator a central repository with policy configuration and retrieving a plurality of policies from the central repository for deployment and monitoring Managed Container Infrastructure Objects (MCIO) and Container Network Functions (CNFs), correlating fault and performance statistics to ensure optimal performance, and forwarding correlated statistics to a distributed orchestration apparatus for workload distribution. The method further includes performing by the cluster orchestrator a fault correlation using multi-dimensional data analysis techniques and determining a root cause of network issues in faults correlated with events associated with the cluster orchestrator and forwarding the alarms correlating faults and performance statistics to the distributed orchestration apparatus.
[0012] In yet another aspect, the objects are achieved by providing a distributed orchestration apparatus for managing and monitoring vRANs in a communication system. The distributed orchestration apparatus initializes a plurality of vRANs and determines a plurality of parameters associated with the plurality of vRANs. The distributed orchestration apparatus groups the plurality of vRANs into a plurality of clusters based on a plurality of parameters and creates a plurality of NEOs at the central SMO entity within a vRAN cluster of a group of a plurality of vRANs or at individual vRANs. Each NEO of the plurality of NEOs manages and monitors the plurality of vRANs. The distributed orchestration apparatus monitors the plurality of parameters associated with the plurality of vRAN clusters and adjusts the number of NEOs in each of the plurality of vRAN clusters based on the plurality of parameters associated with the plurality of vRAN clusters.
[0013] In yet another aspect, the objects are achieved by providing a cluster orchestrator for managing and monitoring the vRANs in a communication system. The cluster orchestrator interfaces at least one NEO with the CISM. The cluster orchestrator is integrated with the NEO and the cluster orchestrator is at least one of a central SMO entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN. The cluster orchestrator polls the central repository with policy configuration and retrieves a plurality of policies from the central repository for deployment and monitors MCIO and CNFs, correlating fault and performance statistics to ensure optimal performance, and forwards correlated statistics to a distributed orchestration apparatus for workload distribution. Further, the cluster orchestrator performs a fault correlation using multi-dimensional data analysis techniques, determines a root cause of network issues in faults correlated with events associated with the cluster orchestrator, and forwards alarms correlating faults and performance statistics to the distributed orchestration apparatus.
[0014] The aspects of these embodiments will be better understood with the following description and accompanying drawings. The descriptions, while indicating preferred embodiments and specific details, are for illustration and not limitation. Many changes and modifications can be made within the scope of these embodiments without departing from their spirit, and all such modifications are included.
[0015] The features, aspects, and advantages of the present embodiments are illustrated in the accompanying drawings, where like reference letters indicate corresponding parts across various figures. The embodiments will be better understood from the following description and drawings.
[0016] Fig. 1 is a block diagram that illustrates the orchestration tier according to prior art.
[0017] Fig. 2 is a block diagram that illustrates network slicing in telco according to prior art.
[0018] Fig. 3 illustrates Network-Based Management and Orchestration (NB MANO) according to prior art.
[0019] Fig. 4 illustrates the existing systems of central management and orchestration.
[0020] Fig. 5 is a block diagram that illustrates a distributed monitoring and correlation system according to the embodiments as disclosed herein.
[0021] Fig. 6 is a block diagram that illustrates a distributed monitoring and correlation system at different vRAN clusters according to the embodiments as disclosed herein.
[0022] Fig. 7 illustrates the key components included in container infrastructure services (CIS) and management services according to the embodiments as disclosed herein.
[0023] Fig. 8 is a block diagram that illustrates the hardware features associated with the distributed orchestration apparatus according to the embodiments as disclosed herein.
[0024] Fig. 9 is a block diagram that illustrates the hardware features associated with the cluster orchestrator according to the embodiments as disclosed herein.
[0025] Figs. 10 and 11 are sequence diagrams that illustrate the life cycle management through the distributed orchestration according to the embodiments as disclosed herein.
[0026] Fig. 12 is a sequence diagram that illustrates the method of managing the lifecycle of vRAN NE CNFs according to the prior art.
[0027] Fig. 13 is a sequence diagram that illustrates the distributed orchestration according to the embodiments as disclosed herein.
[0028] Fig. 14 is a sequence diagram that illustrates the lifecycle steps after CNF instantiation according to the embodiments as disclosed herein.
[0029] Fig. 15 is a flow diagram that illustrates the flow of the NEO agent placement method as disclosed herein.
[0030] Fig. 16 is a flow diagram that illustrates the flow of enhanced NEO agent placement with AI / ML according to the embodiments.
[0031] Fig. 17 is a flow diagram that illustrates the AI-driven predictive scaling of NEO agents according to the embodiments as disclosed herein.
[0032] Fig. 18 is a flow diagram that illustrates the federated learning for geo-distributed NEO agents according to the embodiments as disclosed herein.
[0033] Fig. 19 is a flow diagram that illustrates the digital twin for NEO agent placement according to the embodiments as disclosed herein.
[0034] Fig. 20A illustrates the central monitoring according to the prior art.
[0035] Fig. 20B illustrates the distributed monitoring according to the embodiments as disclosed herein.
[0036] Fig. 21 illustrates the adaptive TelMon placement and federation method according to the embodiments as disclosed herein.
[0037] Fig. 22 is a flow diagram that illustrates the method of distributed monitoring and correlation for telecom workloads in cloud-native environments according to the embodiments as disclosed herein.
[0038] Figs. 23 and 24 are flow diagrams that illustrate the method for managing and monitoring vRANs according to the embodiments as disclosed herein.
[0039] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and details in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term “or” as used herein, refers to a non-exclusive or, unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples are not be construed as limiting the scope of the embodiments herein.
[0040] As is traditional in the field, embodiments are described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and optionally be driven by firmware and software. The circuits, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments be physically separated into two or more interacting and discrete blocks without departing from the scope of the proposed method. Likewise, the blocks of the embodiments be physically combined into more complex blocks without departing from the scope of the proposed method.
[0041] Traditional monitoring tools such as Prometheus and Velero provide basic metrics collection and backup capabilities but do not offer real-time localized monitoring at the edge. Existing orchestrators manage monitoring centrally, which introduces delays and scaling challenges. The proposed invention distributes these tasks across edge clusters, improving performance and scalability.
[0042] Unlike current systems, which are static and manually configured, the proposed invention uses DRL to dynamically adjust monitoring placements based on real-time network conditions, ensuring optimal performance and resource use. TelMon / ATPFA differs from traditional tools by using a decentralized distributed monitoring approach tailored for edge environments in 5G and vRAN networks, whereas traditional tools rely on centralized monitoring. TelMon / ATPFA integrates with cloud-native tools and utilizes AI for dynamic optimization, offering more flexibility and faster fault detection. It is specifically designed for real-time processing at the network edge.
[0043] TelMon / ATPFA focuses on decentralized real-time monitoring and correlation for distributed vRAN deployments using Kubernetes operator patterns and AI-driven dynamic optimization. In contrast, conventional methods address centralized network fault management with predictive analysis for traditional telecom infrastructures. TelMon / ATPFA is optimized for edge environments, leveraging cloud-native tools like Prometheus, while conventional methods do not emphasize edge processing or the use of cloud-native orchestration tools. Further, TelMon / ATPFA's use of AI for adaptive placement and federation is distinct from the more static centralized approaches of conventional methods.
[0044] Embodiments disclosed herein provide a system and method for distributed monitoring and correlation for telecom workloads in cloud-native environments. The present disclosure presents a distributed monitoring and correlation system for vRAN telecom workloads in cloud-native environments, leveraging Kubernetes operator patterns and AI-based decision engines. The system enables real-time monitoring, fault correlation, and dynamic placement of monitoring components at edge clusters, optimizing network performance and resource utilization while minimizing latency and overhead.
[0045] The proposed system decentralizes monitoring tasks by deploying TelMon instances at edge clusters using Kubernetes operators. It integrates with existing tools like Prometheus for federated metrics collection and employs AI-based decision engines for dynamic placement, optimizing the monitoring process based on real-time network conditions.
[0046] Referring now to the drawings and more particularly to Figs. 1 through 24, where similar reference characters denote corresponding features consistently throughout the figure, these are shown preferred embodiments.
[0047] Fig. 1 is a block diagram illustrating the orchestration tier according to prior art. This orchestration tier is a vendor-agnostic solution for Network Automation across physical, virtual, and container networks. It is capable of creating Network Slice SLA, Network Health monitoring, and other functionalities.
[0048] The orchestrator automatically provisions, deploys, scales, and manages cloud resources and applications based on defined policies and workflows. Tasks such as VM lifecycle management, distributed job scheduling, and parallel processing of computing jobs on diverse cloud resources are handled by the orchestrator.
[0049] Serving as the central intelligence in a network service architecture, the orchestrator tier coordinates end-to-end service delivery across multiple layers. Customer-facing inputs from the OSS / BSS layer (101) are integrated, translated into global service management through the Global Orchestrator (102) and slice managers (103), and distributed to domain orchestrators responsible for their network or cloud segments. These domain orchestrators then coordinate closely with the NFV Management and Orchestration (MANO) layer, which automates lifecycle management of virtual and containerized network functions on the infrastructure.
[0050] Resource optimization, automated deployment, fault management, and cross-domain policy enforcement are achieved through standardized interfaces and layered control by the orchestrator tier. The NFV Management and Orchestration framework supports this deployment, with the NFV Orchestrator (NFVO) (113) coordinating services, the CNF / VNF Manager (114) managing individual functions, and the Cloud / Virtualized Infrastructure Manager (CIM) (115) controlling the underlying compute, storage, and network resources. Together, these components enable dynamic and automated creation of network slices tailored to different customer requirements.
[0051] Fig. 2 is a block diagram illustrating network slicing in telecommunications according to prior art. The figure depicts the architectural framework for managing and orchestrating Network Functions Virtualization (NFV) components. Network Slicing Management functions are layered as follows: Communication Service Management Function (CSMF) (116A), which translates communication service requirements into network slice requirements and communicates with the Network Slice Management Function (NSMF); Network Slice Management Function (NSMF) (116B), which manages and orchestrates End-to-End Network Slice Instances (NSIs) across different network domains (RAN, Core, Transport), derives network slice subnet requirements, and communicates with the Network Slice Subnet Management Function (NSSMF) and CSMF; and Network Slice Subnet Management Function (NSSMF) (116C), which manages and orchestrates Network Slice Subnet Instances (NSSIs) within a network domain and communicates with the NSMF. The NFV Orchestrator (NFVO) manages the lifecycle of Network Services and global NFVI resource requests.
[0052] Interaction between the NSMF (116B) and the ETSI NFV MANO framework (107) facilitates the management of the underlying infrastructure for network slicing. NFV MANO provides an abstraction layer for resource and VNF / NS lifecycle management. 3GPP slicing concepts can be mapped to NFV concepts, with a Network Slice potentially viewed as a Network Service in NFV. While the NFVO manages Network Services, the NSMF manages the lifecycle of Network Slice Instances that can be realized by Network Services handling End-to-End (E2E) aspects across domains. E2E NSMF requires orchestration across multiple domains, which may include interacting with different NFVOs.
[0053] The architecture supports three types of slice isolation: Isolated (120), Independent (121), and Customized (122), depending on service needs. These slices span the Access Domain (123) (with CU / DU / RU functions), the Core Domain (124) (housing functions such as AMF, SMF, PCF, UDM, and NSSF), and extend to the Internet (125). Examples of service-specific slices include an enhanced Mobile Broadband (eMBB) slice offering 10 Mbps for VR / AR applications, an ultra-Reliable Low Latency Communication (uRLLC) slice with 2 ms latency for autonomous or mission-critical services, and a massive Machine Type Communication (mMTC) slice enabling IoT connectivity at 10 kbps. This orchestration hierarchy ensures flexible, SLA-driven, and delivery of diverse 5G services.
[0054] Fig. 3 illustrates the NB MANO. The NB of MANO in ETSI defines the protocol and data model for the following interfaces. A VNF Package includes all of the required files and meta-data descriptors necessary to validate and instantiate a VNF management (128). The Network Service Descriptor (NSD) management (127) is a deployment template that include information used by the NFV NFVO (117) for the life cycle management of an NS. An NS is a composition of Network Functions (NF) arranged as a set of functions with unspecified connectivity between them or according to one or more forwarding graphs. Proper VNF Identification is further required across the VNF lifecycle from development to retirement / decommission.
[0055] The Application Programming Interface (API) specifications include NSD Management interface (as produced by the NFVO towards the OSS / BSS), NS Lifecycle Management interface (as produced by the NFVO towards the OSS / BSS), NS Performance Management interface (as produced by the NFVO towards the OSS / BSS), NS Fault Management interface (as produced by the NFVO towards the OSS / BSS), VNF Package Management interface (as produced by the NFVO towards the OSS / BSS), NFVI Capacity Information interface (as produced by the NFVO towards the OSS / BSS), VNF Snapshot Package Management interface (as produced by the NFVO towards the OSS / BSS) and NS LCM coordination interface (as produced by the OSS / BSS towards the NFVO).
[0056] Fig. 4 illustrates the existing systems of central management and orchestration. Central orchestration in cloud environments refers to the centralized automated coordination of disparate cloud resources, services, and workflows across public, private, or hybrid architectures through a unified management platform. This orchestration includes centralized automated coordination of disparate cloud resources, services, and workflows across public, private, or hybrid architectures via a unified management platform.
[0057] The SMO (129) is a specialized orchestration framework central to modern and virtualized Edge devices like vRANs. These devices are typically controlled and managed through centralized software-defined platforms that combine orchestration, monitoring, and automation. SMO (129) centrally provisions, configures, monitors, and optimizes vRANs. Edge devices (131) often have intermittent, unreliable network links, leading to issues like packet loss, jitter, and fluctuating bandwidth. Loss of connectivity can disrupt management operations, make updates harder to sync, and result in stale device states. The need to manage a large number and variety of devices across different locations creates operational complexity. Updating firmware, maintaining compatibility, ensuring regulatory compliance, and syncing configurations remotely is resource-intensive and error-prone, especially as networks scale up.
[0058] The technical problem addressed by this invention is the complexity and inefficiency of centralized orchestration systems in managing large-scale, geographically dispersed telecommunication environments. These systems struggle with latency, single points of failure, and the inability to handle localized decision-making and resource optimization effectively. Further, monitoring large-scale Containerized Network Functions (CNFs) is very resource-consuming in centralized orchestration, leading to significant performance bottlenecks and inefficiencies.
[0059] Scalability issues arise as the network grows, causing the central orchestrator to become a bottleneck, unable to handle the increasing load. Latency and performance bottlenecks occur due to the distance between the control plane and the network nodes, leading to delayed responses and suboptimal performance. The centralized nature of these systems makes them vulnerable to single points of failure; if the central orchestrator goes down, the entire network management process is disrupted. Resource consumption in monitoring large-scale CNFs centrally is resource-intensive, leading to inefficiencies and potential performance degradation.
[0060] Fig. 5 is a block diagram illustrating a distributed monitoring and correlation system according to the embodiments disclosed herein. The invention introduces a distributed orchestration solution for Container Infrastructure Service (CIS) clusters, integrating a NEO within each Managed CIS Cluster Object (MCCO). This approach decentralizes orchestration tasks, enabling localized decision-making, resource optimization, and lifecycle management of Network Functions (CNFs) across multiple nodes.
[0061] The implementation of distributed orchestration involves multiple steps and components for managing and monitoring CNFs. Within each MCCO, a NEO is deployed to handle network element lifecycle activities such as instantiation, scaling, updating, and termination. Each CNF cluster incorporates its own NEO embedded within the MCCO to provide localized orchestration and lifecycle management, ensuring that each cluster can autonomously manage its resources while remaining aligned with the overall orchestration framework.
[0062] NEOs within each CIS cluster retrieve policies and CNF packages from the management cluster, ensuring consistent deployment and configuration. Resolution techniques handle multiple overlapping policies. NEOs monitor CNFs and MCCOs, correlating fault and performance statistics to ensure optimal performance. They forward correlated statistics to the management cluster, distributing the monitoring workload evenly across CIS clusters.
[0063] The NEOs manage the entire lifecycle of vRAN NE CNFs, including instantiation, modification, and termination based on consumer infrastructure planning decisions and policies. The figure illustrates the proposed method for managing vRAN orchestration functionality in the cloud-native environment using a distributed orchestrator and dynamic clustering. The method includes creating multiple NEOs (402) at the central SMO entity (400) to monitor the plurality of vRAN clusters (300). The NEOs are scalable horizontally based on the demand of the plurality of vRAN clusters, creating at least one NEO (402) for each of the plurality of vRAN clusters. The NEOs (402) are responsible for the management and monitoring tasks specific to their respective vRAN clusters (300). The figure illustrates the implementation of the NEO (402) at central resources (403), edge devices (401), and at the vRAN clusters (300).
[0064] In an embodiment, the terms NEO and Telcom Orchestrator Monitoring (TelMon) are used interchangeably. When there is a change in network conditions, such as a significant increase in traffic load leading to higher latency and potential congestion in certain network segments, the Adaptive TelMon Placement and Federation Technique (ATPFA) updates its state representation based on the new traffic conditions and evaluates the current placement of TelMon instances. A deep reinforcement learning (DRL) model identifies that certain edge clusters are experiencing high load and that relocating some TelMon instances to less congested clusters would improve monitoring efficiency. The ATPFA migrates the affected TelMon instances to the selected clusters and adjusts federation settings to reduce the overhead of metrics aggregation. After the migration, the ATPFA monitors the impact of the action on network performance and updates its model based on the observed results.
[0065] Distributed orchestration offers several advantages over centralized orchestration, particularly in large-scale geographically dispersed environments. By distributing management tasks across multiple nodes, the system can scale efficiently without bottlenecks. Decentralization eliminates single points of failure, enhancing the system's overall reliability and resilience. Localized decision-making and resource optimization improve performance and resource utilization. Distributed monitoring reduces the resource consumption associated with centralized monitoring, enabling timely fault detection and performance optimization.
[0066] Distributed orchestration offers several advantages over CICD-based distributed deployment, particularly in monitoring and fault correlation perspectives. Localized decision-making and resource optimization improve performance and resource utilization, which cannot be fully achieved with a central CICD pipeline. Distributed monitoring reduces resource consumption associated with centralized monitoring, enabling timely fault detection and performance optimization, which is not possible with a CICD pipeline.
[0067] Fig. 6 is a block diagram illustrating a distributed monitoring and correlation system at different vRAN clusters (300) according to the disclosed embodiments. The NEO (402) is placed at the central SMO entity (400) within a vRAN cluster (300A) of a group of a plurality of vRANs or at individual vRANs, wherein each NEO (402) of the plurality of NEOs manages and monitors the plurality of vRANs. Block 701 illustrates the NEO located at the central SMO entity (400), responsible for managing and monitoring a group of edge vRANs. The NEO is scalable to accommodate new groups. Block 702 illustrates the NEO (402) located at one of the vRAN clusters, responsible for managing and monitoring that group’s vRANs. Further, block 703 illustrates the NEO located at all the vRAN clusters, responsible for managing and monitoring each vRAN. Block 704 depicts the creation of the NEO across all the cluster orchestrators (edge devices). In an embodiment, the terms SMO and central SMO entity are used interchangeably.
[0068] Fig. 7 illustrates the key components included in CIS and management services according to the disclosed embodiments. The figure illustrates the NFO MANO reference models of CNFs. The CNF extends the standard NFV MANO architecture to manage both virtual machines and containerized workloads. In this model, the MANO stack is composed of the NFVO, VNF Manager (VNFM), and Virtualized Infrastructure Manager (VIM), augmented with dedicated functions for container and cluster management, namely CISM (805) and Container Cluster Management (CCM) (807). The CISM (805) provides lifecycle management capabilities for container infrastructures. It interfaces with underlying cluster / container orchestrators such as Kubernetes, abstracting their specifics for higher layers of MANO. MCIO (809) represents a container infrastructure service managed by the CISM (805) function. It captures the desired and actual state of containerized workloads or subsets, including requested and allocated infrastructure resources, as well as applicable operational policies. The MCIOs are used to describe containerized workloads at an abstract level, facilitating OS container management and orchestration. The NEOs are integrated with the MCIOs.
[0069] The proposed invention describes a method for distributed orchestration within large-scale geographically dispersed telco environments, wherein management tasks are distributed across multiple nodes to enable localized decision-making and resource optimization. The NEO is deployed across selected clusters or each vRAN cluster or central orchestrators like the central SMO entity (400) and others. Deployment of the NEO within each MCCO (806) includes each CNF cluster housing its own NEO embedded within the MCCO for overseeing Network Element (NE) lifecycle management.
[0070] In each CIS cluster, the NEO interfaces with CISM to govern the MCIO, which serves as a CNF. The NEO (402) within each CIS cluster (300) retrieves policies and decides to retrieve CNF packages from a central Management cluster, ensuring consistent deployment and configuration with resolution techniques for multiple overlapping policies. The NEO (402) within each CIS cluster (300) monitors CNFs and MCIOs, correlating fault and performance statistics to ensure optimal performance and forwarding correlated statistics to the Management cluster for workload distribution.
[0071] Integration of the NEO with a centralized repository (including a Container Image Repository and version control system for realizing infrastructure as code configurations, policies, and security as code) and centralized monitoring (such as Prometheus) is proposed for storing and visualizing performance metrics. The centralized repository, often based on a Version Control System (VCS) and a container image registry, stores all the code and configuration files necessary for building, deploying, and managing cloud and containerized applications. Centralized monitoring systems aggregate and analyze metrics, logs, and other data from various sources across the cloud environment. Centralized repositories store information such as infrastructure definitions written in code, describing virtual machines, networks, storage, and other infrastructure components. Policy definitions (e.g., security, compliance, tagging) are also stored. Centralized monitoring primarily collects time-series metrics, which are numerical values recorded over time, including tracking user activity and system changes for security and compliance records, system-level events, and messages.
[0072] The NEO (402) on each vRAN cluster (300) polls the central repository (801) for policy configuration, retrieving necessary CNF package information for deployment. Distributed monitoring is implemented where each CIS Node in a CIS Cluster monitors the health of MCIOs, and CISM in the CIS cluster monitors the health of CIS Nodes and MCIOs, with NEO correlating metrics up to the CNF level and forwarding alarms and events to centralized monitoring.
[0073] The proposed method and system offer scalable and resilient orchestration approaches, distributing correlation and monitoring tasks across CIS clusters to alleviate the central orchestrator's burden, ensuring efficient lifecycle management and performance monitoring of CNFs. The NEO placement is based on Policy-Driven Grouping, AI / ML-Assisted NEO Placement, and dynamic NEO instantiation and termination. Adaptive NEO Placement Based on Policy-Driven Grouping includes a method where the distributed orchestration apparatus dynamically forms logical groups of CIS clusters based on pre-defined policies configured based on user requirements and SLA. Only a subset of CIS clusters within each group is designated to host a NEO, thereby distributing orchestration tasks across select clusters while offloading others. AI / ML-Assisted NEO placement is based on traffic predictions, dynamically determined using AI / ML techniques that predict traffic patterns, service demand, and network utilization trends, enabling selective NEO deployment only in clusters expected to experience higher loads while others operate without local NEO instances during periods of low activity.
[0074] Dynamic NEO Instantiation and Termination Based on Real-Time Monitoring includes a method of on-demand instantiation and termination of NEO instances within CIS clusters, triggered based on real-time monitoring data. NEOs are spun up when traffic thresholds or performance degradation indicators are detected and gracefully terminated when traffic falls below defined thresholds.
[0075] Distributed Group-Level Orchestration Using Selective NEOs includes a system where NEO instances deployed in selected CIS clusters (based on group policy or traffic prediction) act as group-level orchestrators, coordinating CNF lifecycle management, policy enforcement, and fault correlation not only for their own cluster but also for neighboring CIS clusters within the same logical group that do not host NEOs.
[0076] Resource Optimization by Offloading Monitoring and Orchestration Tasks is a method that frees up compute and storage resources on low-traffic CIS clusters by eliminating the need for persistent NEO presence, offloading monitoring, fault correlation, and CNF lifecycle tasks to the nearest available NEO in the same logical group.
[0077] Policy-Defined NEO Role Assignment is a mechanism where the role of each CIS cluster (whether to host a NEO or rely on neighboring clusters) is dynamically assigned by the central orchestrator using a policy engine that factors in traffic forecasts, service criticality, and available resources across clusters.
[0078] Hybrid Central-Distributed Orchestration with Selective NEO Placement is a hybrid orchestration model where the distributed orchestration apparatus (also referred to as central orchestrator or distributed orchestrator in the embodiments) directly manages CNFs in low-traffic clusters without local NEO presence, while high-traffic clusters retain embedded NEOs to handle localized decision-making, ensuring a balance between central control and distributed intelligence.
[0079] NEO Migration for Adaptive Load Balancing includes migrating active NEO instances between CIS clusters based on changing traffic patterns, cluster health, and resource availability, enabling continuous optimization of orchestration load across the infrastructure.
[0080] Fig. 8 is a block diagram illustrating the hardware features associated with the distributed orchestration apparatus. Examples of the distributed orchestration apparatus include, but are not limited to, SMO, Cloud Management Platforms (CMPs), container orchestration engines, operational support systems, NFVO, cloud and application orchestration platforms, and service life cycle management platforms, among others.
[0081] The distributed orchestration apparatus includes a memory (201), a processor (202), an I / O interface (204), and a distributed orchestration controller (203). The memory (201) stores instructions for the processor (202) and may include various non-volatile storage elements like hard disks, optical disks, flash memories, EPROM, or EEPROM. It is considered a non-transitory storage medium, meaning it is not in a carrier wave or propagated signal but can be moved. The memory (201) holds extensive information, including vRAN parameters, SLA details, traffic characteristics, failure rate predictions, CNF package information, and predefined policies.
[0082] The processor (202) can be a CPU, AP, GPU, VPU, or NPU and executes instructions stored in the memory (201), fetching various vRAN-related parameters, SLA details, traffic characteristics, failure rate predictions, CNF package information, and predefined policies.
[0083] The I / O interface (204) facilitates communication between the memory (201) and external devices, ensuring the processor's operating speed is synchronized with input and output devices. It connects peripheral devices like cluster orchestrators and the distributed orchestration controller, performing scenario-specific actions like deactivating or migrating NEO to enhance user experience.
[0084] In an embodiment, the distributed orchestration controller (203) of the distributed orchestration apparatus communicates with the processor (202), I / O interface (204), and memory (201) for managing and monitoring vRANs. Initializing a plurality of vRAN clusters, the distributed orchestration controller (203) determines the plurality of parameters associated with the plurality of vRANs and groups the plurality of vRANs into clusters based on these parameters. Creating a plurality of NEOs at the central SMO entity within a vRAN cluster or at individual vRAN clusters, each NEO manages and monitors the plurality of vRANs. Detecting data from the vRAN clusters to the NEO, the distributed orchestration controller (203) adjusts the number of NEOs in each vRAN cluster based on the parameters and data.
[0085] In an embodiment, the plurality of parameters comprises geographical location, service level agreement (SLA), traffic characteristics, and failure rate prediction for the plurality of vRANs. The NEO agents created by the distributed orchestration controller include creating multiple NEOs at the central SMO to monitor the vRAN clusters, scalable horizontally based on demand, and creating at least one NEO for each vRAN cluster. Each NEO is responsible for management and monitoring tasks specific to their respective vRAN clusters.
[0086] In an embodiment, each NEO created at the central SMO entity within a vRAN cluster acts as group-level orchestrators, coordinating lifecycle management, policy enforcement, and fault correlation for their own cluster or neighboring vRAN clusters within the same logical group that do not host NEOs. Monitoring changes in the plurality of parameters at the vRAN clusters, the distributed orchestration controller (203) determines low or high traffic conditions. During low traffic conditions, the apparatus deactivates at least one NEO from the vRAN clusters and activates or creates NEOs during high traffic conditions.
[0087] Upon deactivating NEOs during low traffic conditions, the distributed orchestration controller (203) manages and monitors the vRAN clusters without NEOs or assigns management to neighboring NEOs. Creating the plurality of NEOs includes adaptive placement based on predefined policy-driven grouping, AI / ML model-assisted placement, or dynamic placement based on real-time monitoring data.
[0088] Predefined policy-driven grouping includes grouping vRAN clusters based on user requirements, with policies flexible to change based on parameters. AI / ML model-assisted placement dynamically determines NEO placement using models that predict network conditions, traffic patterns, service demand, and network utilization trends, enabling selective NEO deployment in vRAN clusters expected to experience higher loads. Dynamic NEO placement based on real-time monitoring includes on-demand instantiation and termination of NEOs triggered by traffic thresholds or performance degradation indicators, continuously collecting, processing, and visualizing parameters based on real-time data and defined rules.
[0089] Fig. 9 is a block diagram illustrating the hardware features associated with the cluster orchestrator. Examples of the cluster orchestrator include, but are not limited to, SMO, CMPs, NFVO, cloud and application orchestration platforms, service life cycle management platforms, vRAN clusters, vRANs, central resources, central orchestration apparatus, edge devices, among others.
[0090] The vRAN clusters or CIS clusters, used interchangeably, are virtualized RAN components or a set of resources including CIS instances and CISM instances. These utilize compute resources as well as storage and network resources and are integrated at Distributed Unit (DU), Central Unit (CU), and Radio Unit (RU), deployed together on a shared pool of cloud or edge computing resources. Instead of being fixed hardware-based network elements, these RAN functions run as virtual network functions (VNFs) or cloud-native network functions (CNFs) on general-purpose servers.
[0091] The cluster orchestrator (300) comprises a memory (301), a processor (302), an I / O interface (304), and a distributed orchestration controller (303). The memory (301) stores instructions to be executed by the processor (302) and can include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard disks, optical disks, floppy disks, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In some examples, the memory (301) may be considered a non-transitory storage medium, indicating that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term non-transitory should not be interpreted to mean that the memory (301) is non-movable. In certain examples, the memory (301) stores larger amounts of information and may store data that can change over time (e.g., in Random Access Memory (RAM) or cache). The memory (301) stores a plurality of parameters associated with vRANs, a service level agreement (SLA) for the plurality of vRANs, traffic characteristics of the plurality of vRANs, failure rate prediction for the plurality of vRANs, CNF package information, predefined policies, and others.
[0092] The processor (302) may include one or a plurality of processors, such as a general-purpose processor like a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). Configured to execute the instructions stored in the memory (301), the processor (302) fetches the plurality of parameters associated with vRANs, a service level agreement (SLA) for the plurality of vRANs, traffic characteristics of the plurality of vRANs, failure rate prediction for the plurality of vRANs, CNF package information, predefined policies, and others. Further, the processor (302) retrieves instructions and executes them.
[0093] The I / O interface (304) transmits information between the memory (301) and external peripheral devices. Peripheral devices are the input-output devices associated with the cluster orchestrator (300). The I / O interface (304) receives several pieces of information from a plurality of UEs, network devices, servers, and the like. Ensuring that the operating speed of the processor is synchronized with respect to the input and output devices, the I / O interface (304) establishes a connection between different peripheral devices like cluster orchestrator memory, distributed orchestration controller (303), and others to perform distributed orchestration for any scenario-specific action like deactivating or migrating NEO or other functions to enhance the user experience.
[0094] In an embodiment, the distributed orchestration controller (303) of the distributed orchestration apparatus communicates with the processor (302), I / O interface (304), and memory (301) for managing and monitoring vRANs. The distributed orchestration controller (303) interfaces a plurality of NEO with a CISM. Integrated with the NEO, the cluster orchestrator (300) is at least one of a central SMO entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN cluster. Furthermore, the distributed orchestration controller (303) polls the central repository with policy configuration and retrieves the plurality of policies from the central repository for deployment. These policies include, but are not limited to, rules that define how infrastructure resources (compute, storage, network) are provisioned and managed, desired state configurations for applications and infrastructure ensuring consistency, versioning and compliance guidelines that enforce regulatory compliance requirements dynamically in cloud-native thresholds, and rules set for metrics, logs, and traces to trigger alerts and automated operational responses, and workflow rules that automate scaling, updates, and deployment processes within orchestration platforms. The plurality of policies are flexible and can be customized based on user requirements provided in the form of YAML intent and others.
[0095] Further, the distributed orchestration controller (303) monitors MCIO and CNFs, correlating fault and performance statistics to ensure optimal performance and forwarding correlated statistics to a distributed orchestration apparatus for workload distribution. Continuously monitoring and analyzing the operational state and metrics of containerized infrastructure and network functions, the distributed orchestration controller (303) ensures they are running efficiently and without errors.
[0096] Further, the distributed orchestration controller (303) performs fault correlation using multi-dimensional data analysis techniques and determines the root cause of network issues in faults correlated with events associated with the cluster orchestrator. Forwarding alarms correlating faults and performance statistics to the distributed orchestration apparatus, the fault correlation includes collecting and analyzing data such as errors, resource usage (CPU, memory), network throughput, latency, and other telemetry from MCIO and CNFs. By correlating the faults (e.g., failures, crashes) and performance metrics (e.g., throughput, latency), the distributed orchestration controller (303) gains insights into the health and efficiency of containerized workloads.
[0097] The distributed orchestration controller (303) determines the root cause of network issues in faults correlated with events associated with the cluster orchestrator (300) and forwards alarms correlating faults and performance statistics to the distributed orchestration apparatus (200).
[0098] Figs. 10 and 11 are sequence diagrams that illustrate the life cycle management through distributed orchestration according to the embodiments disclosed herein. Fig. 10 illustrates the life cycle management of managing and monitoring vRANs. At step S1101, the cluster orchestrator (300) downloads the policies from the central repository (500), validates the policies, and executes them. Further, the cluster orchestrator continuously polls for changes, as illustrated at step S1003. When there is any change, the new CNF package information is uploaded to the central repository (500), as illustrated at step S1004.
[0099] At step S1005, the cluster orchestrator (300) downloads the HELM chart followed by the CNF deployment. The cluster orchestrator (CO) (300) also downloads the CNF images from the central repository (500), as depicted at step S1007. The HELM charts include versioned application packages or directories includeing a specific structure of files and folders describing how to deploy an application, whereas the CNF image is a container image that implements network functionality (such as firewalls, VNFs, gateways, etc.) optimized for cloud-native environments. The CNF image includes the application code, supporting binaries, configuration files, and base libraries. Further, the cluster orchestrator (300) deploys the CNF or application based on descriptors (YAML, Helm charts, and others). This may include pulling images, configuring resources, networking, and storage. The cluster orchestrator continues to monitor for any change in the policy, as illustrated at step S1009. The cluster orchestrator (300) monitors the performance and health of the application, triggers automatic scaling (up / down, in / out), as well as healing (restart, redeploy, reschedule on other nodes) in response to failures or load changes. When there is a change in the policies, such as rolling updates, the CNF package information is modified. Again, the cluster orchestrator (300) downloads the HELM chart from the central repository (500) and informs the API server (600) to modify the CNF, as depicted in steps S1012 and S1013, respectively. The cluster orchestrator (300) further downloads the CNF images and forwards them to the Kubernetes API server (600), which modifies the instantiated CNF. The cluster orchestrator (300) further monitors for any change in the policies. The cluster orchestrator (300) continuously monitors metrics, logs, and traces, and correlates them with faults to proactively detect issues and ensure SLA compliance. When the service or application is no longer needed or is being replaced, a notification is received by the central repository (500) or central resources to remove the CNF or application and to clean up associated resources, as depicted at step S1017. The cluster orchestrator (300) deletes the HELM chart and sends a delete CNF message to the Kubernetes API server (600). At step S1120, the Kubernetes API server (600) deletes the CNF and related resources. The cluster orchestrator removes stale images, configurations, records, and releases associated licenses or quotas from the registry / catalog and continuously monitors for any change in the policy.
[0100] Fig. 11 illustrates the steps followed after the instantiation of the CNF at the edge clusters. The lifecycle steps after CNF instantiation revolve around the Kubernetes control plane (API server (600)) and the data plane (nodes) working in sync to deploy, monitor, scale, heal, and update the CNF workloads. The API server (600) acts as the authoritative source for the desired state, while the nodes execute the containers and report back actual states and health information, ensuring continuous reconciliation and optimal CNF operation within the cluster.
[0101] Upon CNF instantiation, the Kubernetes API server (600) receives the YAML manifests or Helm charts representing the CNF's desired state (deployments, services, configmaps, etc.) and stores them. The Kubernetes scheduler monitors the API server (600) for new pods related to the CNF and decides on which Kubernetes nodes the CNF pods will run, considering resource availability, affinity / anti-affinity rules, taints, and tolerations. The nodes continuously report the status of the containers to the Kubernetes API server (600).
[0102] The Kubernetes (k8s) API server continuously updates the cluster orchestrator with the current status of the pods and containers (Pending, Running, Succeeded, or Failed) upon receiving the request for POD status from the cluster orchestrator (300), as depicted at steps A1104 and 1105 respectively. The POD status includes health information, resource usage, and any lifecycle event updates.
[0103] At step S1106, the cluster orchestrator (300) correlates the POD status to CNF status. Correlating pod status to CNFs includes linking the operational health and lifecycle state of Kubernetes pods hosting CNFs to the overall status and performance of those CNFs. The steps include mapping PODs to the CNFs, health and readiness indication, fault detection and impact analysis, performance monitoring, event and log correlation, and others. The POD to CNF correlation helps in fault isolation, performance optimization, automated remediation, and others. The POD status is a fundamental source of truth for monitoring and managing CNFs within Kubernetes. Effective correlation between pod lifecycle data and CNF operational state enables comprehensive lifecycle management, fault detection, performance tracking, and automated orchestration in cloud-native environments.
[0104] At step S1107, the cluster orchestrator (300) reports the CNF status to the central monitoring based on the policy. The central repository receives requests from OSS or 3gpp systems for CNF status, to which the central repository responds with the CNF status.
[0105] Fig. 12 is a sequence diagram that illustrates the method of federated learning MOI creation and notification according to the embodiments disclosed herein. The life cycle management of the CNF broadly includes instantiation, modification, and termination of the CNFs. At step S1201, the consumer sends an instantiation request by updating the policy with inputs such as configuration parameters, target cluster / namespace, and resource requirements. These inputs are validated against policy compatibility and resource availability at the producer.
[0106] Continuous monitoring for any change in the policy is performed by the producer, as depicted at step S1202. Upon receiving an update of CNF package information through the CNFPakgInfo message from the consumer, the producer transmits CNF package info Response through CNFPkgInfo Response, as depicted at steps S1203 and S1204, respectively. At step S1205, the consumer transmits a Modify CNFPkgInfo request, which includes updated configuration, scaling parameters, or new CNF versions. The producer validates these parameters and assesses the impact on existing CNF instances to prevent service disruption, then transmits the CNFPkgInfo Response to the consumer. The CNF pods and services are updated incrementally at the producer, as depicted at step S1207.
[0107] Further, upon receiving the Delete CNFPkg Info message from the consumer (900), the producer is requested to stop and remove the CNF instance. The producer checks for any active sessions, dependencies, or policies before termination. The CNF is signaled to stop accepting new traffic, complete existing transactions, and shut down. The producer transmits the CNFPkg Info response to the consumer and deletes the CNF instance. Further, the pods, services, config maps, storage claims, and networking resources attached to the CNF are deleted.
[0108] Fig. 13 is a sequence diagram illustrating the distributed orchestration according to the disclosed embodiments. Federated learning includes the collaborative training of a machine learning model across multiple decentralized vRAN clusters while keeping the data local to each cluster for privacy. The federated learning MOI creation and notification subscription depicted in the figure enable producers to manage the lifecycle of vRAN NE CNFs (including instantiation, modification, and termination) and monitor their information. This information can be initiated or queried by consumers within the Management and Orchestration system. The producer (900) comprises a Centralized Repository and Monitoring.
[0109] Multiple CIS Clusters are in the ready state for the deployment of NE CNF workloads. These CIS Clusters are created manually or via the CIS Cluster Management (CCM) API. Consumers plan and dimension which CNF NE should be deployed on which CIS Cluster. Based on the consumer’s infrastructure planning decision, a PolicyInfo Managed Object Instance (MOI) is created. All available CIS Clusters download the PolicyInfo MOI from the Central Repository as depicted at step S1301.
[0110] At step S1302, the NEO at the CIS cluster (central CIS cluster) (300) validates, executes, and stores the policy. PolicyInfo MOI holds configuration on NEO polling settings with the Central Repository, forming critical information for NEO to manage the CNFs' lifecycle in a distributed manner. Upon execution of the PolicyInfo MOI at each CIS Cluster, the NEO on the CIS Cluster starts polling the Central Repository with agreed-upon policy config (e.g., with certain label / tag which can be implementation specific) as illustrated at step S1303.
[0111] The Central Repository has two key repositories: CIR for storing images needed for CNF to run and MCIOP repo for storing configuration and composition (e.g., Helm chart) of CNF to be run. When the consumer decides to deploy / instantiate a CNF, it uploads the CNFPkgInfo MOI to the Central Repository as per step S1304. As all the CIS Clusters (300) regularly poll the Central Repository, they notice the change made by the consumer. The NEO on each CIS Cluster (300), based on its policy, decides whether or not to use the new CNFPkgInfo. On a policy configuration match, the NEO downloads the respective CNFPkgInfo’s Helm chart from MCIOP repo and deploys it on its own CISM as illustrated at steps S1305 and S1306. CISM on the respective CIS clusters creates the necessary MCIO, which is to be composed as a CNF. During the MCIO creation, the respective container images are also pulled from CIR of the Central Repository (800). The NEO on each CIS Cluster ensures that all components of the CNF are in a running state and internally marks it as instantiated. After this, the polling continues with the Central Repository for any further changes (S1309).
[0112] If the running CNF needs to be modified for the sake of upgrading / downgrading the image version or for scaling it in or out by high-level components, then a CNFPkgInfo’s modification request is triggered from Consumer to Producer as depicted at step S1320. From the Consumer side, this modification request can originate in a planned or unplanned manner. The steps S1305 to S1307 are repeated to reflect the latest change in the configuration of the CNFs at steps S13011 to S1313. The same process is repeated for CNF Termination (S1318 to S1320), where the consumer deletes the CNFPkgInfo from the Central Repository (800). Similar to the lifecycle management of CNF, its monitoring is also done in a distributed manner. Each CIS Node in a CIS Cluster monitors the health of all MCIOs. The CISM in the CIS cluster monitors the health of all CIS Nodes and its managed MCIOs. The NEO frequently requests the CISM on all CIS Nodes and MCIOs status, then correlates the metric up to a CNF level. This is the heaviest role of an orchestrator in a centralized approach. In this approach, such correlation is distributed to each CIS Cluster. Only the CNF level alarms and events are sent to centralized monitoring. Consumers can then check the status of various CNFs and their status with the Producer (900).
[0113] The above scenario is only an example and is not specific to CNF monitoring alone; it can be applicable to multiple domains, including IoT monitoring and other Cloud-based Edge monitoring and others.
[0114] The present invention relates to a novel apparatus for improving the efficiency of solar panels. This apparatus includes a series of adjustable mirrors that can be oriented to reflect sunlight onto the solar panels, thereby increasing the amount of light that reaches the panels. The adjustable mirrors are controlled by a central processing unit (CPU) that calculates the optimal angles based on the position of the sun. Further, the apparatus features a cooling system to prevent overheating of the solar panels, which can degrade their performance over time. The cooling system utilizes a combination of air and liquid cooling to maintain the panels at an optimal temperature. This invention aims to enhance the overall efficiency and longevity of solar panels, making solar energy a more viable and sustainable option for power generation.
[0115] Attribute NameDescriptionPolicyNameA unique identifier for the policy, facilitating easy reference and management.PolicyTypeSpecifies the type or category of the policy, such as deployment, scaling, upgrade, or termination.TargetCISClusterIdentifies the target CIS cluster where the policy is intended to be applied. All CNF Packages with certain TargetCISCluster need to be applied onto the specific CIS cluster.TargetCISClusterLabelIdentifies the target CIS cluster and label where the policy is intended to be applied. All CNF Packages with certain label need to be applied onto specific CIS cluster.TargetCNFSpecifies the CNF targeted by the policy, allowing for granular control over specific functions or services.TriggerEventIndicates the event or condition that triggers the execution of the policy, such as a change in workload demand, resource availability, or system health status.ActionDefines the action to be taken by the system in response to the trigger event, including deployment, scaling, upgrade, downgrade, or termination of resources.ResourceThresholdSpecifies the threshold or criteria for resource utilization, performance metrics, or other conditions that trigger the policy action.ScheduleDefines the schedule or timing for policy execution, including frequency, recurrence, and time window for applying the policy.ApprovalRequiredIndicates whether manual approval is required before executing the policy, providing governance and control over automated actions.NotificationRecipientSpecifies the recipient or recipients who should be notified when the policy is executed or when certain conditions are met.LoggingEnabledIndicates whether logging or auditing of policy execution events is enabled, ensuring visibility and traceability of system actions.PolicyDescriptionProvides a detailed description or documentation of the policy, including its purpose, scope, and intended impact on the system.PolicyOwnerSpecifies the individual or team responsible for defining, managing, and overseeing the policy lifecycle.PolicyVersionIndicates the version number or revision history of the policy, enabling tracking of changes and ensuring compatibility with different system configurations.PolicyStatusProvides the current status or state of the policy, such as active, inactive, pending approval, or archived.
[0116] Attribute NameDescriptionPackageNameA unique identifier for the CNF package, facilitating easy reference and management.PackageVersionIndicates the version number or revision history of the CNF package, enabling tracking of changes and ensuring compatibility with different deployments.PackageDescriptionProvides a detailed description or documentation of the CNF package, including its purpose, features, dependencies, and configuration requirements.PackageTypeSpecifies the type or category of the CNF package, such as application, middleware, network function, or service component.PackageSourceIndicates the source or origin of the CNF package, such as a public repository, private registry, or custom build pipeline.PackageImagesSpecifies the container images of the CNF package, facilitating retrieval and distribution of package components. These images are stored in CIR of the Central repository.PackageDependenciesLists the dependencies or prerequisite components required for the CNF package to function properly, ensuring compatibility and seamless integration with the target environment.PackageConfigurationDefines the configuration settings, parameters, and options for customizing the behavior and performance of the CNF package during deployment and runtime. These can be represented as Helm charts or K8S spec files, these are stored in MCIOP of the Central repository.PackageSecuritySpecifies the security measures and protocols implemented within the CNF package to mitigate potential threats, vulnerabilities, and risks to the system.PackageLicensingIndicates the licensing terms and agreements associated with the CNF package, ensuring compliance with legal and regulatory requirements.PackageLifecycleDefines the lifecycle stages and states of the CNF package, including development, testing, staging, production, and decommissioning.PackageOwnerSpecifies the individual or organization responsible for creating, maintaining, and distributing the CNF package, providing accountability and support for users.PackageVisibilityIndicates the visibility or accessibility of the CNF package within the organization or across different deployment environments, ensuring appropriate access controls and permissions.PackageChecksumProvides a checksum or hash value calculated from the CNF package artifacts, enabling integrity verification and detection of tampering or corruption.PackageMetadataIncludes additional metadata or annotations associated with the CNF package, such as authorship, release notes, changelog, and usage guidelines.
[0117] These attribute names and descriptions given the Table 1 and Table 2, can be tailored and extended based on the specific requirements and use cases of the CNFPkgInfo IOC in the context of cloud-native CNF deployment and management.
[0118] Fig. 15 is a flow diagram that illustrates the flow of NEO agent placement method, as disclosed herein.At step S1501, the vRAN clusters are initiated by the distributed orchestration apparatus (200) with attributes like geography, SLA, and others. Further, at step S1502, the vRANs are clustered based on geography, SLA, and traffic similarity (e.g., K-means clustering). Score each placement option for each cluster using weighted metrics.
[0119] For each cluster, the distributed orchestration apparatus (200) calculates scores as illustrated at step S1503. The scores include central SMO placement score, cluster-level placement score, and per-vRAN placement score. The central SMO placement score offers several advantages, including low overhead and scalability, making it suitable for broad deployments. However, it suffers from high latency when servicing distant RANs. The score for this placement is calculated as SMO_Score = (1 - normalized_latency) * SLA_weight, emphasizing the importance of latency and service level agreements (SLAs). The cluster-level placement score provides a balanced approach by optimizing both latency and resource utilization. While it works well for medium to large clusters, it may introduce overhead in smaller cluster environments. Its score is determined by Cluster_Score = (traffic_density / failure_rate) * SLA_weight, reflecting the interplay between network traffic and reliability requirements. Whereas the per-vRAN placement score caters to scenarios demanding ultra-low latency and enhanced fault isolation. This method, though effective, comes with the drawback of high resource overhead due to its granular approach. The score here is computed as PerVRAN_Score = (SLA_criticality / cluster_size), highlighting the criticality of SLAs relative to the size of the cluster. Together, these scoring methods provide a comprehensive framework to evaluate and select the optimal placement strategy based on specific network needs and priorities.
[0120] At step S1504, the distributed orchestration apparatus (200) selects the optimal placement of the NEO on that vRAN or vRAN cluster having the highest score. If SMO_Score > threshold and latency < SLA, the NEO is placed at the central SMO entity. If the Cluster_Score > PerVRAN_Score, then the NEO is placed at the vRAN cluster. Else, the NEO is placed at the individual vRAN.
[0121] At step S1505, the central SMO score is validated, and if the SMO score is greater than the threshold, the NEO is deployed at the central SMO entity. If the SMO score is less than the threshold, then the cluster-level score is validated. If the score is greater than the threshold, the NEO is placed at the vRAN cluster. If the cluster score is less than the threshold, the NEO is placed at each individual vRAN as depicted at step S1507.
[0122] At step S1508, the distributed orchestration apparatus (200) monitors real-time traffic / failure rates. If the traffic change is greater than the threshold, the distributed orchestration apparatus (200) recomputes the scores and migrates NEOs if thresholds are breached. If the traffic change is less than the threshold, for example, less than 20% as illustrated in the figure, the distributed orchestration apparatus (200) takes no action and continues to receive the policy and monitor for any change in the policy.
[0123] The present disclosure relates to enhanced NEO agent placement with AI / ML, as illustrated in Fig. 16. The technique can be enhanced with new decision-making layers and optimized workflows for vRAN orchestration. Three advanced technique extensions address specific gaps: AI-driven predictive scaling for NEO, Federated Learning for Geo-Distributed NEOs, and Digital Twin for NEO Placement Simulation.
[0124] AI-driven predictive scaling for NEO includes the distributed orchestration apparatus (200) predicting future demand and adjusting resources in advance, rather than merely reacting to the current load. The distributed orchestration apparatus (200) collects past usage data (e.g., CPU requests, bandwidth) and trains the AI model to identify patterns and predict future demand for the next few minutes, hours, or days. Based on these predictions, the distributed orchestration apparatus (200) scales the NEO in advance to meet the anticipated load.
[0125] Federated learning is a distributed machine learning approach where multiple devices or organizations collaboratively train a shared model without sharing raw data. Instead of relying on a central monitoring and repository for NEO placement, the distributed orchestration apparatus (200) is trained locally on each vRAN, vRAN cluster, or central SMO entity. Only the model updates (parameters / gradients) are sent to the central server for aggregation. Once the optimal placement of the NEO is determined, the distributed orchestration apparatus (200) performs the NEO placement.
[0126] Fig. 17 illustrates the AI-driven predictive scaling of NEO agents. Static NEO placement may not adapt to sudden traffic spikes or failures; hence, time-series forecasting (e.g., LSTM / Prophet) is used to pre-scale NEO agents. At step S1801, the distributed orchestration apparatus (200) fetches historical traffic or failure data. This data is analyzed and predicted. For example, the distributed orchestration apparatus (200) predicts traffic and failure data for the next five minutes, as shown at step S1803. If the prediction exceeds the threshold, the NEOs are preemptively scaled. If the traffic prediction is below the threshold, the NEOs are deactivated. If the traffic rate is predicted to be stable, the current NEOs are maintained. Further, the distributed orchestration apparatus (200) continuously monitors active NEOs and traffic and failure history for optimal performance, as depicted at step S1807.
[0127] Fig. 18 is a flow diagram illustrating the federated learning process for geo-distributed NEO agents according to the embodiments disclosed herein. Centralized ML training violates data privacy for distributed vRANs; hence, federated learning is employed to train global models across edge clusters. Each vRAN cluster trains the global models across these edge clusters. The central SMO entity aggregates the models, as depicted at step S1802. Further, the global models, which are aggregates of the local models, are deployed at all NEOs at step S1803. At S1804, the model performance is monitored. If the accuracy drops, the local models are retrained; otherwise, the model updates are validated, and the performance is monitored continuously.
[0128] Fig. 19 illustrates the digital twin for NEO agent placement according to the disclosed embodiments. Real-world testing of NEO placement is costly; hence, GenAI-powered digital twins are employed to simulate outcomes. At step S1901, a synthetic vRAN is generated. This synthetic vRAN is a simulated environment built for testing, development, or training models without the need for physical RAN components. It is created using synthetic data or artificial traffic patterns to validate network performance, AI / ML model training, or LCM. NEO placement strategies are tested for central SMO placement, vRAN cluster placement, and individual vRAN placement, as illustrated at steps S1902, S1903, and S1904, respectively. SLA traffic and failure outcomes are predicted for each case. Based on the score, an optimal NEO placement is determined.
[0129] The proposed decentralized approach to life cycle management offers several benefits in the context of access telco domains. Decentralized management and monitoring of the vRANs enhance scalability by distributing workload and management overhead across multiple orchestrator instances, allowing for seamless expansion to accommodate growing network demands. Further, distributed orchestration improves fault tolerance and resilience by eliminating single points of failure and enabling self-healing capabilities at the edge clusters of the network. Performance and latency are enhanced by enabling localized processing and decision-making, minimizing the need for back-and-forth communication with a centralized controller. Adding certain types of monitoring parameters, such as packet loss, latency, and round-trip time (RTT), which are monitored locally rather than centrally, can significantly benefit performance optimization and fault detection, especially in distributed environments where real-time monitoring is critical for timely decision-making. The distributed orchestration aligns well with the inherent distributed nature of access telco networks, where resources are geographically dispersed and interconnected. By leveraging distributed orchestration, telco operators can achieve greater agility, efficiency, and responsiveness in managing access networks while ensuring consistent service quality and reliability across diverse geographical regions. This approach also enables operators to meet stringent latency requirements for emerging applications such as edge computing, IoT, and real-time multimedia services, among others.
[0130] Some of the use cases of the proposed distributed orchestration include, but are not limited to, 5G Network Slicing, where dynamic NEO placement ensures SLA compliance per slice (e.g., low-latency for URLLC, high-throughput for eMBB). Edge Computing distributed NEOs reduce latency for localized vRAN clusters (e.g., smart factories). Further, this can be implemented for disaster recovery, where failover NEOs in geo-distributed clusters maintain service during outages. The proposed distributed orchestration incorporates AI-driven scaling that handles traffic spikes, making the system adaptable. Federated learning reduces the data transfer overhead, and digital twin simulation pre-validates strategies, enhancing the efficiency and resilience of the system.
[0131] In an embodiment, the decentralized orchestration maintains a distributed ledger across multiple clusters, which periodically logs NEO decisions for auditability. Quorum-based updates are implemented for non-critical actions, maintaining eventual consistency. Policy synchronization between the central SMO entity and the NEO is standardized using the O-RAN A1 interface.
[0132] Timestamp-based prioritization is performed in another embodiment, wherein new policies override older ones. SLA-driven arbitration ensures that critical SLAs, like emergency services, take precedence. The central SMO entity acts as an arbitrator, resolving disputes between cluster orchestrators integrated with the NEOs via the A1 interface. Among multiple overlapping policies, the proposed invention implements policy tagging, where policies are labeled according to priority. Digital twin simulations are used before deployment to manage overlapping policies.
[0133] In an embodiment, the NEO monitors key statistics like Infra K8S RAW metrics, latency, PRB utilization, CPU load, failure rate, and CNF application-specific metrics, and creates alarms from raw metrics and root causes (e.g., high latency → DU overload).
[0134] While Fig. 20A illustrates the traditional central monitoring of vRAN clusters, Fig. 20B illustrates the distributed monitoring and correlation system TelMon according to the disclosed embodiments. In an embodiment, the terms NEO and Telcom Orchestrator Monitoring (TelMon) are used interchangeably.
[0135] Components of the ATPFA include the TelMon, Network Metrics, and TelMon States forming the input layer, which utilize input features such as latency, traffic load, fault rate, resource usage, TelMon configuration, and historical data. A Policy Network (Actor) processes these inputs through hidden layers to learn optimal action policies and outputs a probability distribution over possible actions, while a Value Network (Critic) processes inputs through hidden layers to generate state-value predictions and outputs a value estimate for the given state. An Experience Replay Buffer stores (state, action, reward, next state) tuples for training, and a DRL Training Module uses sampled experiences to update both the Policy Network and the Value Network. TelMon Instances are deployed across various edge clusters according to policy decisions and perform active monitoring to collect and correlate metrics. A Prometheus Federation Layer aggregates metrics from multiple clusters and ensures synchronization of coherent metric data across clusters. The overall environment represents a distributed telco network comprising multiple edge clusters with dynamic TelMon deployment and real-time network conditions reflected by continuously updated network metrics and events.
[0136] In an example scenario, a telecom network may experience a sudden spike in traffic due to a large-scale event, such as a live broadcast or natural disaster. When network conditions change, such as a significant increase in traffic load leading to higher latency and potential congestion in certain network segments, the Adaptive TelMon Placement and Federation Technique (ATPFA) updates its state representation based on the new traffic conditions and evaluates the current placement of TelMon instances. A deep reinforcement learning (DRL) model identifies that certain edge clusters are experiencing high load and that relocating some TelMon instances to less congested clusters would improve monitoring efficiency. Consequently, the ATPFA migrates the affected TelMon instances to the selected clusters and adjusts federation settings to reduce the overhead of metrics aggregation. Following the migration, the ATPFA monitors the impact of the action on network performance and updates its model based on the observed results.
[0137] The proposed system, TelMon, is designed as a decentralized monitoring and fault correlation solution specifically optimized for telecom workloads in cloud-native environments. Traditional monitoring approaches in telco environments rely heavily on centralized orchestrators, which introduce latency and bottlenecks due to the need for data transmission to and from a central point. In contrast, TelMon operates directly at the edge clusters where telco workloads, such as virtualized Central Units (vCUs) and Distributed Units (vDUs), are deployed.
[0138] By distributing monitoring tasks across edge clusters, TelMon ensures that data processing and fault detection occur close to the source of the data, minimizing latency of real-time observability. This architecture not only reduces the overhead associated with centralized data aggregation but also allows for localized decision-making, which is crucial for maintaining stringent SLAs in telecom networks.
[0139] TelMon utilizes a Kubernetes operator pattern to facilitate seamless integration with existing cloud-native infrastructure. The operator pattern is a method of packaging, deploying, and managing Kubernetes applications. It allows TelMon to extend Kubernetes capabilities by automating the deployment, scaling, and management of the monitoring solution across edge clusters.
[0140] The Kubernetes operator for TelMon is responsible for several key functions. Automated deployment includes automatically deploying TelMon instances across selected edge clusters based on predefined policies or dynamic AI-driven decisions. Scaling is dynamically managed, with TelMon instances being scaled up or down based on real-time network conditions and predicted load. Configuration management ensures that TelMon configurations are consistently applied across all instances, including metrics collection parameters, alerting rules, and data retention policies.
[0141] Prometheus, a widely used open-source monitoring solution, provides powerful capabilities for metrics collection and alerting. In the context of TelMon, Prometheus is leveraged for its federation feature, which enables the aggregation of metrics from multiple Prometheus instances deployed across different clusters.
[0142] The Prometheus Federation Configuration is enhanced by TelMon to support the unique requirements of telco environments. Key aspects include:
[0143] A Hierarchical Federation Model is implemented, where local Prometheus instances collect metrics at the pod and node levels within each cluster. TelMon then aggregates these metrics into a global view. This hierarchy ensures that only relevant data is federated, reducing network traffic and improving performance.
[0144] Custom Federation Rules are defined to select which metrics are federated based on their relevance to telco workloads. For example, metrics related to network latency, packet loss, and CPU usage are prioritized for vRAN deployments.
[0145] TelMon extends Prometheus with custom federation features to address the specific needs of telco workloads. Selective Metrics Aggregation is employed to aggregate only metrics that are critical for fault correlation and performance monitoring, thereby reducing data overhead. Multi-Cluster Synchronization ensures that metrics collected from different clusters are synchronized in real-time, allowing for a unified view of the network’s performance.
[0146] In an embodiment, a key innovation in TelMon is the use of an AI-based decision engine for dynamically selecting the optimal placement of TelMon instances across edge clusters. This engine leverages Deep Reinforcement Learning (DRL), a type of machine learning that learns optimal actions through interactions with an environment.
[0147] The DRL model is designed to optimize the placement of TelMon based on real-time network conditions and historical performance data. Key components of the model include:
[0148] State Representation: The state includes real-time metrics from the network, such as traffic patterns, fault rates, CPU and memory usage, and network latency.
[0149] Action Space: Actions correspond to the deployment or redeployment of TelMon instances across available edge clusters, adjusting federation settings, or altering monitoring parameters.
[0150] Reward Function: The reward function is defined to prioritize actions that minimize latency, maximize fault detection, and optimize resource usage. It incorporates penalties for actions that lead to increased latency or reduced fault correlation.
[0151] The DRL model is trained using a simulated network environment that mirrors the conditions of a live telecom network. Once trained, the model is deployed within the central orchestration system, where it continuously monitors network conditions and adjusts the placement and configuration of TelMon in real-time.
[0152] Fig. 21 is a block diagram that illustrates a ATPFA System Overview, according to the embodiments as disclosed herein. The ATPFA is designed to dynamically manage the placement of TelMon instances across edge clusters and to optimize the federation settings of metrics collection in a distributed cloud-native telco environment. The technique utilizes Deep Reinforcement Learning (DRL) (2103) to adaptively respond to changing network conditions, such as variations in traffic load, network faults, and resource availability. By learning from these conditions, the ATPFA ensures that monitoring is performed from the most effective locations within the network, thereby minimizing latency of collected metrics.
[0153] Components of ATPFA includes the following:
[0154] State Representation: This component defines the current state of the network and TelMon deployment. It includes real-time metrics such as network traffic patterns, latency, CPU and memory usage, fault rates, and the current configuration of TelMon instances.
[0155] Action Space: This represents the possible actions that can be taken by ATPFA. Actions include deploying a new TelMon instance in a different cluster, migrating an existing instance to a new location, adjusting the federation settings of Prometheus instances, and changing monitoring parameters (e.g., sampling rates, types of metrics collected).
[0156] Reward Function: This function defines the goals of ATPFA by assigning rewards or penalties to actions based on their impact on network performance. It is designed to encourage actions that reduce latency, enhance fault detection, and optimize resource utilization.
[0157] DRL Model: The core of ATPFA is a DRL model that learns from the environment (the telecom network) to make optimal placement and federation decisions. The model is trained using historical data and real-time feedback from the network.
[0158] Fig. 22 is a flow diagram that illustrates a method of distributed monitoring and correlation for telecom workloads in cloud-native environments according to the embodiments disclosed herein.
[0159] 1. Input Layer (S2301): Network metrics and TelMon states are fed into both the Policy Network (Actor) and the Value Network (Critic).
[0160] 2. Policy Network (2206): This network processes inputs to determine the optimal action, such as deploying or migrating TelMon instances and configuring federation settings.
[0161] 3. Value Network (2205): It estimates the value of the current state to assist in policy optimization and decision-making.
[0162] 4. Experience Replay Buffer (2202): The buffer stores the experiences (state, action, reward, next state) that are used for training the neural networks.
[0163] 5. DRL Training Module (2203): Experiences are sampled from the buffer, and the neural networks are updated using backpropagation and gradient descent.
[0164] 6. TelMon Instances (402): Based on the decisions from the Policy Network, TelMon instances are dynamically deployed or reconfigured across edge clusters.
[0165] 7. Prometheus Federation Layer (2208): This layer collects and synchronizes metrics from multiple clusters to provide a unified view for TelMon instances.
[0166] 8. Environment (2207): The distributed telco network is represented here, constantly providing new states and rewards to the DRL system.
[0167] ATPFA method includes the following Steps:
[0168] 1. Initialize DRL Model: Start by initializing the DRL model with historical data on network performance, traffic patterns, and previous TelMon placements.
[0169] 2. Monitor Network Conditions: Continuously monitor real-time metrics from the network, including traffic load, fault rates, latency, and resource usage.
[0170] 3. State Update: Update the state representation based on the collected metrics and the current configuration of TelMon instances.
[0171] 4. Decision Making: Using the DRL model, evaluate the current state and select the optimal action from the action space. This could include deploying a new TelMon instance, migrating an instance to a different cluster, or adjusting federation settings.
[0172] 5. Execute Action: Implement the selected action in the network. For example, if the DRL model suggests migrating a TelMon instance to a new cluster, perform the migration and adjust configurations as necessary.
[0173] 6. Feedback Loop: Collect feedback on the impact of the action on network performance. Update the DRL model with this feedback to improve future decision-making.
[0174] 7. Repeat: Continuously repeat steps 2-6, allowing the DRL model to learn and adapt to changing network conditions over time.
[0175] Initialization: The DRL model is initialized with pre-trained weights that are derived from historical network data. The technique runs in a continuous loop, collecting real-time network metrics such as traffic load, latency, fault rate, and resource usage. The current state of the network and TelMon deployment is represented by a set of metrics. This state is updated continuously as new metrics are collected. Further, the DRL model processes the current state and predicts the optimal action to take. Actions can include deploying a new TelMon instance, migrating an existing instance, adjusting federation settings, or modifying monitoring parameters. Based on the action chosen by the DRL model, the technique executes the corresponding function. For instance, if the action is to migrate a TelMon instance, the technique identifies the source and target clusters and performs the migration. After executing the action, the technique evaluates the network performance to assess the impact of the action. This feedback is used to calculate a reward, which is then fed back into the DRL model to update its weights, refining its decision-making process for future iterations. Actions and results are logged for future reference and training, and the technique waits for a predefined interval before starting the next monitoring cycle.
[0176] While pseudocode gives a high-level overview, here’s a basic mathematical representation of how the DRL model works within the technique:
[0177] 1. State Representation: St= [Lt, Ct, Ft, Rt, Mt] Where:
[0178] - L = Latency at time t
[0179] - C = Traffic load at time t
[0180] - F = Fault rate at time t
[0181] - R = Resource usage at time t
[0182] - M = Current TelMon configuration at time t
[0183] Action Space: A= [a1, a2, a3, a4] Where:
[0184] - a_1 = Deploy new TelMon instance
[0185] - a_2 = Migrate existing TelMon instance
[0186] - a_3 = Adjust federation settings
[0187] - a_4 = Adjust monitoring parameters
[0188] Reward Function: R(st,at)=∝ (latency_improvement) + β (fault_reduction) + γ(resource_efficiency), Where:
[0189] - (alpha, beta, gamma) are weight parameters that prioritize different aspects of network performance.
[0190] Policy Update: Using a policy gradient method (like Proximal Policy Optimization - PPO):θ_(t+1)=θ_t+ η∇_θj(θ), Where:
[0191] - ( θ theta ) are the model parameters,
[0192] - ( eta ) is the learning rate, and
[0193] - ( J(theta) ) is the expected reward under policy ( theta ).
[0194] This mathematical representation, along with the pseudocode, provides a comprehensive view of how ATPFA works to dynamically manage TelMon instances and optimize metrics collection in distributed telco environments.
[0195] Neural Model Architecture for ATPFA includes the following:
[0196] 1. Input Layer: The input layer will take in a feature vector representing the current state of the network and TelMon configuration. This feature vector includes various metrics and states that are critical for decision-making in a telco environment.
[0197] Input Features:
[0198] - Latency Metrics (L): Real-time latency measurements from different parts of the network.
[0199] - Traffic Load Metrics (C): Current traffic load data across the network and edge clusters.
[0200] - Fault Rate (F): The rate of detected faults or errors in the network.
[0201] - Resource Usage (R): CPU, memory, and bandwidth utilization data.
[0202] - TelMon Configuration (M): Current placement, active / inactive status, and monitoring parameters of TelMon instances.
[0203] - Historical Data: Past metrics to provide temporal context (e.g., moving averages, time series data).
[0204] Total Number of Input Nodes:Let’s assume the input vector size is (n) where n represents the total number of features extracted from the network metrics and TelMon states.
[0205] 2. Hidden Layers: To capture the complex relationships between network states and optimal actions, the model will use multiple hidden layers. These layers will be dense (fully connected) to ensure comprehensive learning of features.
[0206] First Hidden Layer:
[0207] - Number of Neurons: (h_1) (e.g., 128 neurons)
[0208] - Activation Function: Rectified Linear Unit (ReLU) - Helps in learning non-linear relationships.
[0209] Second Hidden Layer:
[0210] - Number of Neurons: (h_2) (e.g., 64 neurons)
[0211] - Activation Function: ReLU - Continues to learn more complex features.
[0212] Third Hidden Layer (Optional):
[0213] - Number of Neurons: (h_3) (e.g., 32 neurons)
[0214] - Activation Function: ReLU - Further refines learned features to focus on the most relevant data for decision-making.
[0215] Normalization Layer (Optional):
[0216] - Batch Normalization: Helps stabilize learning and improve performance by normalizing outputs from the previous layer.
[0217] 3. Output Layer: The output layer generates a set of possible actions that the TelMon system can take, based on the current state of the network. The action space in a DRL context is typically represented as a probability distribution over all possible actions.
[0218] - Number of Output Nodes: ( a ) (where ( a ) is the number of possible actions, e.g., deploy new instance, migrate instance, adjust federation, etc.)
[0219] - Activation Function: Softmax - Converts the output into a probability distribution over actions, allowing the DRL agent to select the action with the highest probability.
[0220] 4. Neural Network Representation for DRL: Given that DRL is being used, the neural network architecture will be integrated into a policy gradient or actor-critic framework. Here’s how these might be structured:
[0221] Policy Network (Actor): This network learns the policy function π(a | s;θ), which outputs the probability of taking action ( a ) given the state ( s ).
[0222] - Architecture: The architecture described above is an example of the policy network where the output layer provides a probability distribution over actions.
[0223] - Objective: Maximize expected rewards by learning optimal policies.
[0224] ∇θJθ) Espπaπθ∇θlogπθ(a│s)Qπθs a
[0225] - θ Parameters of the policy network.
[0226] - Qπθ(s,a)] Estimated advantage of taking action a in state s.
[0227] Value Network (Critic): This network estimates the value function V(s; Ø), which represents the expected cumulative reward from state s.
[0228] - Architecture: Similar to the policy network but typically with fewer output nodes. The output here is a single node representing the value estimate of the state.
[0229] - Objective: Provide a baseline for reducing variance in the policy gradient estimates, stabilizing training.
[0230] L(Ø)=Es~pπ[(Rt-VØ(s))]2]
[0231] - Ø : Parameters of the value network.
[0232] - Rt: Actual reward obtained from the environment.
[0233] - VØ(s): Predicted value of state s.
[0234] 5. Training the Neural Model: Training includes:
[0235] - Experience Replay: Storing past experiences as (state, action, reward, next state) tuples in a replay buffer and sampling from this buffer to break temporal correlations.
[0236] - Backpropagation and Gradient Descent: Updating network weights using gradients obtained from policy gradients or actor-critic methods (like Advantage Actor-Critic, A2C).
[0237] - Reward Function: The reward function is designed to maximize network performance, minimize latency, and optimize TelMon placement based on current and predicted network conditions.
[0238] Prometheus Configuration for Federation:
[0239] - Federated Scraping: Configuring Prometheus instances to scrape metrics from each other, enabling cross-cluster metric collection and synchronization.
[0240] - Alerting Rules: Custom alert rules are defined for TelMon events, ensuring standardized alarms (e.g., 3GPP, VES).
[0241] A custom federation implementation is provided. A custom metrics aggregation layer is implemented, which handles the synchronization of high-frequency data from edge clusters, thereby reducing data transfer overhead. Data collection intervals and retention policies are optimized based on network load and the criticality of monitored services.
[0242] TelMon’s fault correlation engine employs multi-dimensional data analysis techniques to identify and correlate faults across multiple layers of the network. This process includes analyzing metrics from various sources, including pod, node, network, and workload-specific metrics.
[0243] Multi-Layered Analysis Framework: Metrics are collected from different layers of the network, including application logs, network telemetry, and hardware performance counters. A correlation engine processes these metrics in real-time, identifying patterns and anomalies that indicate potential faults. Machine learning techniques are utilized by the engine to reduce false positives.
[0244] Once a fault is detected, the correlation engine performs root cause analysis by comparing metrics across different clusters and network layers, pinpointing the source of the issue.
[0245] In an embodiment, Intra and Inter model exchange / aggregation between CSPs with a Federated Learning framework (FLF) is a novel approach in machine learning that alleviates challenges in data collection. A blockchain-based approach for securely transmitting models between CSPs is introduced, representing a novel addition to the prior art. As an immediate step, relevant sections will be taken as a 3GPP SA5 study item.
[0246] In another embodiment, TelMon is designed for telecom workloads in cloud-native environments. It is characterized by its capability to perform real-time monitoring and fault correlation at edge clusters using a Kubernetes operator pattern. TelMon operates in a fully decentralized manner, distributing monitoring tasks across various edge clusters. This decentralized architecture allows for localized data processing and reduces latency in fault detection. Utilizing a Kubernetes operator pattern, TelMon is compatible with existing Kubernetes-managed environments, ensuring easy deployment, scalability, and management. TelMon is capable of real-time data processing and correlation, providing immediate insights into network performance and potential faults. This capability is critical for maintaining Service Level Agreements (SLAs) in telecom networks.
[0247] In an embodiment, the TelMon supports federation and can integrate with tools like Prometheus to collect, aggregate, and synchronize metrics across multiple clusters using the federation feature. Seamless federation integration allows for the aggregation of metrics from disparate clusters into a coherent dataset for analysis. TelMon provides customizable metrics collection, enabling telecom operators to select specific metrics relevant to their unique network configurations and performance criteria. The ability to synchronize metrics across clusters provides a unified and comprehensive view of overall network health and performance. Further, the Adaptive TelMon Placement and Federation Technique (ATPFA) is designed to dynamically manage the placement of TelMon instances across edge clusters and optimize the federation settings of metrics collection.
[0248] In an embodiment, a system includes a central orchestration component that employs an AI-based decision engine utilizing Deep Reinforcement Learning (DRL) to dynamically select optimal edge clusters for the placement of TelMon. TelMon performs federated metrics collection and correlates fault and performance events across these clusters. The AI-based decision engine leverages DRL techniques to continuously learn from network conditions and optimize the placement of TelMon across edge clusters. This approach allows for adaptive, context-aware decisions that consider factors like network load, latency, compute usage, and fault rates. Dynamic placement of TelMon is managed based on real-time analysis of network metrics and predicted load changes. This ensures that monitoring and correlation are always performed from the most effective locations within the network. TelMon engages in federated metrics collection, gathering data from various network clusters to provide a comprehensive view of network health and performance. The dynamic placement of TelMon optimizes its ability to correlate fault and performance events, enhancing the detection and diagnosis of network issues in a distributed, real-time fashion. By dynamically placing TelMon where it is most needed, the system ensures optimal utilization of network resources, minimizes unnecessary data movement, and reduces overall system load. Adaptability to changing network conditions, such as fluctuating traffic patterns or unexpected faults, is achieved through the use of DRL. The system relocates TelMon as needed to maintain effective monitoring and correlation. Dynamic placement of TelMon reduces latency in monitoring and correlation tasks by ensuring these processes occur close to the data source, thereby minimizing the overhead of data transmission across the network.
[0249] The invention relates to an AI-Based Dynamic TelMon Allocation system. Predictive analytics are employed to dynamically reconfigure TelMon based on real-time network conditions. TelMon autonomously adjusts its monitoring parameters, such as sampling rates and types of metrics collected, as well as federation settings, in response to changes in network traffic load or detected anomalies. A hierarchical clustering AI engine is utilized to identify the most suitable edge clusters for specific tasks, thereby improving the efficiency of workload management and fault correlation. The AI decision-making engine within TelMon can utilize not only Deep Reinforcement Learning (DRL) but also other forms of machine learning, including supervised learning, unsupervised learning, and reinforcement learning, to dynamically adjust its operations based on network conditions.
[0250] Fault and Performance Correlation Using Multi-Dimensional Data Analysis. The TelMon system performs fault correlation using multi-dimensional data analysis techniques. These techniques incorporate metrics from various layers of the telecom network, including pod node network and workload-specific metrics. A method for real-time fault isolation is employed by TelMon, wherein faults detected in one cluster are analyzed and correlated with events in neighboring clusters to rapidly determine the root cause of network issues.
[0251] Multi-Layered Analysis: TelMon’s fault correlation engine utilizes data from multiple network layers. This approach provides a holistic view of network performance, enabling more precise fault identification and resolution.
[0252] Seamless Integration with Existing Telecom Infrastructure: The TelMon is designed for seamless integration with existing telecom infrastructure, including legacy systems and modern cloud-native deployments, ensuring compatibility and ease of adoption. The TelMon generates alarms in standardized formats such as 3GPP, VES, or other telecom-specific event formats, ensuring compatibility with existing telecom network management systems.
[0253] In an embodiment, distributed monitoring offers several advantages over centralized monitoring, particularly in large-scale, geographically dispersed vRAN / oRAN environments. Real-time monitoring at the edge improves latency and fault detection, reducing the time required to detect and respond to faults. AI-driven dynamic placement optimizes resource utilization by ensuring monitoring resources are efficiently allocated, thereby reducing unnecessary data movement and processing. The Kubernetes operator pattern provides scalability and flexibility, allowing for easy deployment, scaling, and management in cloud-native environments. Enhanced SLA compliance is achieved through faster fault detection and performance monitoring, helping to maintain strict SLAs in telco networks.
[0254] The proposed invention addresses the complex challenge of providing real-time localized monitoring and fault correlation in distributed cloud-native telecom networks. By dynamically placing TelMon at the edge and synchronizing metrics across clusters, the system provides a comprehensive view of network health while supporting localized analysis. This approach enables telecom operators to meet strict SLA requirements, maintain consistent service quality, and respond effectively in dynamic, latency-sensitive environments.
[0255] Kubernetes Operators are utilized to manage the deployment and operation of TelMon instances across edge clusters, ensuring scalability and ease of management. Deep Reinforcement Learning, an AI technique, powers the dynamic placement of monitoring components, continuously optimizing their position based on real-time network data. Integrating with Prometheus, the proposed invention aggregates and synchronizes metrics across clusters to provide a holistic view of network health.
[0256] The proposed invention fundamentally changes the approach to network monitoring in telecom environments by shifting from a centralized to a decentralized model. This shift not only reduces latency and improves fault detection but also enhances the scalability and adaptability of the monitoring system, making it ideal for modern cloud-native deployments.
[0257] By decentralizing monitoring tasks through TelMon orchestration and applying AI-driven optimization, the invention delivers more than a fivefold improvement in fault detection speed and measurable gains in network performance. Localized processing at the edge minimizes transmission overhead, while adaptive AI dynamically positions TelMon instances in response to changing network states. These capabilities strengthen SLA compliance and ensure that operators can maintain reliable, high-quality services in distributed, latency-sensitive telecom environments.
[0258] Fig. 23 is a flow diagram that illustrates the method of managing and monitoring the vRANs according to the embodiments as disclosed herein. At step S2301, the method includes initializing by a distributed orchestration apparatus (200) a plurality of vRANs as depicted. At step S2302, the method includes determining by the distributed orchestration apparatus (200) a plurality of parameters associated with the plurality of vRANs. At S2303, the method includes grouping the plurality of vRANs into a plurality of clusters based on the plurality of parameters. After grouping the vRANs, at S2304, the method includes creating the NEOs at the central SMO entity, an individual vRAN, or within a vRAN cluster of a plurality of vRANs.
[0259] At step S2305, the method includes monitoring by the distributed orchestration apparatus (200) the parameters associated with the vRAN or vRAN cluster. At step S2306, the method includes adjusting the number of NEOs in each of the plurality of vRAN clusters based on the plurality of parameters associated with the plurality of vRAN clusters.
[0260] In an embodiment, the plurality of parameters comprises a geographical location of the plurality of vRANs, a service level agreement (SLA) of the plurality of vRANs, traffic characteristics of the plurality of vRANs, and failure rate prediction for the plurality of vRANs, among others.
[0261] In an embodiment, creating the NEOs includes creating the NEO or NEOs at the central SMO to monitor the plurality of vRAN clusters where the NEOs are scalable horizontally based on the demand of the plurality of vRAN clusters. The at least one NEO is created for each of the plurality of vRAN clusters, or the NEO is created at an individual vRAN. The NEO is responsible for the management and monitoring tasks specific to their respective vRAN clusters.
[0262] In an embodiment, the NEO acts as group-level orchestrators coordinating lifecycle management, policy enforcement, and fault correlation of at least one of their own vRAN cluster or for neighboring vRAN clusters within the same logical group that do not host NEO.
[0263] In an embodiment, the distributed orchestration apparatus (200) monitors for a change in the plurality of parameters at the vRAN clusters and determines at least one of a low traffic condition or a high traffic condition at the vRAN clusters. Further, the distributed orchestration apparatus (200) activates or creates at least one NEO at the vRAN cluster, central SMO entity, or at the individual vRAN during the high traffic conditions at the vRAN clusters. If the traffic conditions are low, the distributed orchestration controller (200) deactivates at least one NEO from the plurality of NEOs created at the vRAN clusters during low traffic conditions. Upon deactivating the NEO, the distributed orchestration apparatus (200) manages and monitors the vRAN clusters without NEOs during low traffic conditions or assigns the management and monitoring of the at least one vRAN cluster to a neighboring NEO during low traffic conditions after deactivating the NEO based on the plurality of parameters.
[0264] In an embodiment, creating the plurality of NEOs comprises at least one of an adaptive NEO placement based on predefined policy-driven grouping, Artificial Intelligence (AI) / Machine Learning (ML) model-assisted NEO placement, or dynamic NEO placement based on real-time monitoring data.
[0265] In an embodiment, the predefined policy-driven grouping comprises grouping of vRAN clusters based on predefined policies configured based on user requirements and SLA.
[0266] In an embodiment, the AI / ML model-assisted NEO placement comprises placement of a plurality of NEOs within selected vRAN clusters dynamically determined using AI / ML models that predict network conditions, traffic patterns, service demand, and network utilization trends, enabling selective NEO deployment only in vRAN clusters expected to experience higher loads while others operate without local NEOs during low traffic conditions.
[0267] In an embodiment, the dynamic NEO placement based on real-time comprises on-demand instantiation and termination of the plurality of NEOs triggered based on real-time monitoring data wherein the plurality of NEOs are placed when traffic thresholds or performance degradation indicators are detected and terminated when traffic falls below defined traffic thresholds.
[0268] Fig. 24 is a flow diagram that illustrates the method of managing and monitoring the vRANs by the cluster orchestrator integrated with the NEO according to the embodiments as disclosed herein. At step S2401, the method includes interfacing by the cluster orchestrator (300) the NEO with a CISM wherein the cluster orchestrator (300) is integrated with the NEO and the cluster orchestrator is at least one of a central SMO entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN cluster.
[0269] At step S2402, the method includes polling by the cluster orchestrator (300) a central repository with policy configuration and retrieves a plurality of policies from the central repository for deployment. Further, at step S2403, the method includes monitoring by the cluster orchestrator (300) MCIO and CNFs, correlating fault and performance statistics and forwarding correlated statistics to a distributed orchestration apparatus for workload distribution. At step S2404, the method includes performing by the cluster orchestrator (300) a fault correlation using multi-dimensional data analysis techniques. Further, at S2405, the method includes determining a root cause of network issues in faults correlated with events associated with the cluster orchestrator as depicted. Further, at S2405, the method includes forwarding by the cluster orchestrator (300) alarms correlating faults and performance statistics to the distributed orchestration apparatus (200).
[0270] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the embodiments as described herein.
Claims
1.A method for managing and monitoring virtual Radio Access Networks (vRANs) in a communication system, comprising:initialising, by a distributed orchestration apparatus (200), a plurality of vRANs;determining, by the distributed orchestration apparatus (200), a plurality of parameters associated with the plurality of vRAN;grouping, by the distributed orchestration apparatus (200), the plurality of vRANs into a plurality of clusters based on plurality of parameters;creating, by the distributed orchestration apparatus (200), a plurality of Network Element Orchestrator (NEO) at a central Service Management and Orchestration (SMO) entity, within a vRAN cluster of a group of a plurality of vRANs, or at individual vRAN, wherein each NEO of the plurality of NEOs manages and monitors the plurality of vRANs;monitoring, by the distributed orchestration apparatus (200), the plurality of parameters associated with the plurality of vRAN clusters; andadjusting, by the distributed orchestration apparatus (200), a number of NEOs in each of the plurality of vRAN clusters based on the plurality of parameters associated with the plurality of vRAN clusters.2.The method as claimed in claim 1, wherein the plurality of parameters comprises a geographical location for the plurality of vRANs, a service level agreement (SLA) for the plurality of vRANs, traffic characteristics of the plurality of vRANs, and failure rate prediction for the plurality of vRANs.3.The method as claimed in claim 1, wherein creating, by the distributed orchestration apparatus (200), the plurality of NEOs comprises at least one of:creating at least one NEOs at the central SMO entity to monitor the plurality of vRAN clusters, wherein the NEOs are scalable horizontally based on the demand of the plurality of vRAN clusters;creating at least one NEO for each of the plurality of vRAN clusters; andcreating a NEO at an individual vRAN, wherein the NEOs are responsible for the management and monitoring tasks specific to their respective vRAN clusters.4.The method as claimed in claim 1, wherein each NEOs created at the central SMO entity, within a vRAN cluster of a group of a plurality of vRANs, or at individual vRAN act as group-level orchestrators, coordinating lifecycle management, policy enforcement, and fault correlation of at least one of their own cluster or for neighboring vRAN clusters within the same logical group that do not host NEOs.5.The method as claimed in claim 1, comprising:monitoring, by the distributed orchestration apparatus (200), a change in the plurality of parameters at the vRAN clusters;determining, by the distributed orchestration apparatus (200), at least one of a low traffic condition or a high traffic condition at the vRAN clusters; andperforming, by the distributed orchestration apparatsus, at least one of:activating or creating at least one NEO at the vRAN cluster or central SMO entity or at the individual vRAN, during the high traffic conditions at the vRAN clusters, anddeactivating at least one NEOs from the plurality of NEOs created at the vRAN clusters during low traffic conditions.6.The method as claimed in claim 5, wherein upon deactivating at least one NEOs from the plurality of NEOs created at the vRAN clusters during low traffic conditions, comprising:managing and monitoring the at least one vRAN clusters without NEOs during low traffic conditions; orassigning the management and monitoring of the at least one vRAN clusters to a neighbouring NEO, during low traffic conditions after deactivating the NEO.7.The method as claimed in claim 1, creating, by the distributed orchestration apparatus (200), the plurality of NEOs comprises at least one of an adaptive NEOs placement based on predefined policy-driven grouping, Artificial Intelligence (AI) / Machine Learning (ML) model assisted NEO placement or dynamic NEOs placement based on real-time monitoring data.8.The method as claimed in claim 5, wherein the predefined policy-driven grouping comprises grouping of vRAN clusters based on predefined policies configured based on user requirement.9.The method as claimed in claim 5, wherein the AI / ML model assisted NEO placement comprises placement of plurality of NEOs within a selected vRAN clusters dynamically determined using AI / ML models that predict network conditions, traffic patterns, service demand, and network utilization trends, enabling selective NEO deployment only in vRAN clusters expected to experience higher loads, while others operate without local NEOs during low traffic conditions.10.The method as claimed in claim 5, wherein the dynamic NEOs placement based on real-time comprises on-demand instantiation and termination of the plurality of NEOs, triggered based on real-time monitoring data, wherein the plurality of NEOs are placed when traffic thresholds or performance degradation indicators are detected, and terminated when traffic falls below defined traffic thresholds.11.A method for managing and monitoring virtual Radio Access Networks (vRANs) in a communication system, comprising:interfacing, by a cluster orchestrator (300), at least one Network Element Orchestrator (NEO) with a CISM, wherein the cluster orchestrator (300) is integrated with the NEO and the cluster orchestrator (300) is at least one of a central Service Management and Orchestration (SMO) entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN cluster;polling, by the cluster orchestrator (300), a central repository with policy configuration and retrieves a plurality of policies from the central repository for deployment;monitoring, by the cluster orchestrator (300), Managed Container Infrastructure Object (MCIO) and Container Network Functions (CNFs) correlating fault and performance statistics to ensure optimal performance and forwarding correlated statistics to a distributed orchestration apparatus (200) for workload distribution;performing, by the cluster orchestrator (300), a fault correlation using multi-dimensional data analysis techniques;determining, by the cluster orchestrator (300), a root cause of network issues in faults correlated with events associated with the cluster orchestrator (300); andforwarding, by the cluster orchestrator (300), alarms, correlating faults and performance statistics to the distributed orchestration apparatus (200).12.A distributed orchestration apparatus (200) for managing and monitoring virtual Radio Access Networks (vRANs) in a communication system, comprising:a memory (201);a processor (202); anda distributed orchestration controller (203) coupled with the memory (201) and the processor (202), wherein the distributed orchestration controller (203):initialises a plurality of vRANs;determines a plurality of parameters associated with the plurality of vRAN;groups the plurality of vRANs into a plurality of clusters based on plurality of parameters;creates a plurality of Network Element Orchestrator (NEO) at a central Service Management and Orchestration (SMO) entity, within a vRAN cluster of a group of a plurality of vRANs, or at individual vRAN, wherein each NEO of the plurality of NEOs manages and monitors the plurality of vRANs;monitors the plurality of parameters associated with the plurality of vRAN cluster; andadjusts a number of NEOs in each of the plurality of vRAN clusters based on the plurality of parameters associated with the plurality of vRAN clusters.13.The distributed orchestration apparatus (200) as claimed in claim 12, wherein the distributed orchestration controller (203):monitors a change in the plurality of parameters at the vRAN clusters;determines at least one of a low traffic condition or a high traffic condition at the vRAN clusters; andperform at least one of:activates or creates at least one NEO at the vRAN cluster or central SMO entity, during the high traffic conditions at the vRAN clusters, anddeactivates at least one NEOs from the plurality of NEOs created at the vRAN clusters during low traffic conditions.14.A cluster orchestrator (300) for managing and monitoring virtual Radio Access Networks (vRANs) in a communication system, comprising:a memory (301);a processor (302); anda distributed orchestration controller (303), coupled with the memory (301) and the processor (302), wherein the distributed orchestration controller (303):interfaces at least one Network Element Orchestrator (NEO) with a Container Infrastructure Service Management (CISM), wherein the cluster orchestrator (300) is integrated with the NEO and the cluster orchestrator (300) is at least one of a central Service Management and Orchestration (SMO) entity, a vRAN cluster of a group of a plurality of vRANs, or an individual vRAN;polls a central repository with policy configuration and retrieves a plurality of policies from the central repository for deployment;monitors Managed Container Infrastructure Object (MCIO) and CNFs correlating fault and performance statistics to ensure optimal performance and forwards correlated statistics to a distributed orchestration apparatus (200) for workload distribution;performs a fault correlation using multi-dimensional data analysis techniques;determines a root cause of network issues in faults correlated with events associated with the cluster orchestrator (300); andforwards alarms, correlating faults and performance statistics to the distributed orchestration apparatus (200).