Synthesis of time series data representative of a communication network

The method addresses the limitations of existing synthetic data generation techniques by generating time series data that accurately represents the seasonality and anomalies of communication network data, enabling effective model training and anomaly detection.

WO2025115024A1PCT designated stage expired Publication Date: 2025-06-05TELEFONAKTIEBOLAGET LM ERICSSON (PUBL) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2023/051115
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing synthetic data generation techniques fail to accurately replicate the seasonality characteristics and anomalies present in actual time series data collected from communication networks, making them inadequate for training predictive models and detecting network anomalies.

Method used

The proposed method involves determining temporal buckets corresponding to equally spaced time instances, obtaining temporal statistics including distribution parameters for each bucket and outlier distributions, and generating synthetic data samples that accurately represent the seasonality and anomalies of the actual data.

Benefits of technology

This approach effectively generates synthetic data that mimics the multi-level seasonality and stochasticity of actual network traffic KPI patterns, enabling efficient training of predictive models and detection of anomalies, while being computationally efficient and scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2023051115_05062025_PF_FP_ABST
    Figure IN2023051115_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments include methods for generating synthetic data representative of characteristics of a communication network. Such methods include determining a plurality of temporal buckets corresponding to time instances equally spaced over a first duration at a frequency associated with a seasonality characteristic of data from the communication network. Such methods include obtaining the following temporal statistics for the synthetic data: for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and second parameters defining second distribution(s) of outliers from the first distributions associated with the temporal buckets. Such methods include, based on the temporal statistics, generating a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic. Such methods include adjusting a subset of the synthetic data samples to represent anomalies present in data from the communication network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYNTHESIS OF TIME SERIES DATA REPRESENTATIVE OF A COMMUNICATION NETWORK

[0002] TECHNICAL FIELD

[0003] The present disclosure relates generally to communication networks and more specifically to techniques for detecting operational anomalies (e.g., failures, etc.) that manifest themselves across multiple domains of a communication network.

[0004] BACKGROUND

[0005] The fifth generation (5G) of cellular systems, also referred to as New Radio (NR), was initially standardized 3GPP Rel-15 and continues to evolve in subsequent releases. NR is developed for maximum flexibility to support a variety of different use cases including enhanced mobile broadband (eMBB), machine type communications (MTC), ultra-reliable low latency communications (URLLC), side-link device-to-device (D2D), and several other use cases. 5G / NR technology shares many similarities with fourth-generation LTE.

[0006] At a high level, the 5G System (5GS) consists of an Access Network (AN) and a Core Network (CN). The AN provides UEs connectivity to the CN, e.g., via base stations such as gNBs or ng-eNBs. As described in more detail below, the CN includes a variety of Network Functions (NF) that provide a range of different functionalities such as session management, connection management, charging, authentication, etc.

[0007] The ever-increasing complexity of communication networks, including 5G networks, drives the evolution of analytics systems that support operation, optimization, and planning of these networks. This includes detecting and addressing sudden, undesired changes in network operation and / or performance (e.g., failures). These analytics systems, in turn, require collecting and processing of enormous amounts of data, particularly time series data.

[0008] In general, a time series is a sequence of data or information values, each of which has an associated time instance (e.g., when the data or information value was generated and / or collected). The data or information can be anything measurable that depends on time in some way, such as prices, humidity, or number of people. One important characteristic of a time series is frequency, which is how often the data values of the data set are recorded. Frequency is also inversely related to the period (or duration) between successive data values.

[0009] Time series analysis includes techniques that attempt to understand or contextualize time series data, such as to make forecasts or predictions of future data (or events) using a model built from past time series data. To best facilitate such analysis, it is preferrable that the time series consists of data values measured and / or recorded with a constant frequency or period. Time series datasets can be collected from geographic locations, such as from nodes of a communication network located in one or more geographic areas (e.g., countries, regions, provinces, cities, etc.). For example, values of performance measurement (PM) counters and key performance indicators (KPIs) can be collected from the various network nodes at certain time intervals. Time series data collected in this manner can be used to analyze, predict, and / or understand user behavior patterns as well as network performance trends.

[0010] SUMMARY

[0011] For example, time series data collected in this manner can be used to train a predictive model for network behavior. In some cases, however, actual time series data collected from a network may be insufficient for proper training of the model, or the cost and resources needed to collect and store network-generated time series data may be prohibitive. Moreover, in case the collected time series data contains user-specific data, storage of such data creates risks of security breaches and non-compliance with data privacy regulations.

[0012] Thus, there is a need for techniques that generate synthetic data that accurately represents the characteristics of actual time series data collected from a network. However, existing synthetic data generation techniques have various problems or issues that make them unable to replicate characteristics of actual time series data collected from a network, such as seasonality characteristics and anomalies typical present in collected data.

[0013] An object of embodiments of the present disclosure is to address these and other problems, issues, and / or difficulties by providing techniques for synthesis of time series data that accurately represents characteristics of actual time series data collected from a network, including but not limited to seasonality characteristics and anomalies.

[0014] Some embodiments include methods (e.g., procedures) for generating synthetic data representative of characteristics of a communication network.

[0015] These exemplary methods include determining a plurality of temporal buckets corresponding to time instances that are equally spaced over a first duration at a frequency associated with a seasonality characteristic of data collected from the communication network. These exemplary methods also include obtaining the following temporal statistics for the synthetic data:

[0016] • for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and

[0017] • second parameters defining one or more second distributions of outliers from the first distributions associated with the plurality of temporal buckets. These exemplary methods also include, based on the temporal statistics, generating a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic. These exemplary methods also include adjusting a subset of the plurality of synthetic data samples to represent anomalies present in data collected from the communication network.

[0018] In some embodiments, the synthetic data samples are representative of network performance management (PM) KPI samples generated by the communication network. In some embodiments, each time instance associated with a synthetic data sample corresponds to one of the temporal buckets and each synthetic data sample is generated based on the first parameters of the corresponding temporal bucket. In some of these embodiments, the first parameters defining the first distribution for each temporal bucket include the following:

[0019] • a distribution type of the first distribution;

[0020] • a minimum of the first distribution;

[0021] • a maximum of the first distribution;

[0022] • probability of a left outlier that is less than the minimum; and

[0023] • probability of a right outlier that is greater than the maximum.

[0024] In some of these embodiments, obtaining the temporal statistics includes the following operations:

[0025] • obtaining a source dataset produced by the communication network, wherein the source dataset includes a plurality of data samples associated with time instances that are equally spaced at the frequency over a reference duration;

[0026] • assigning each data sample to one of the temporal buckets based on a correspondence between the time instance associated with the data sample and the time instance associated with the assigned temporal bucket; and

[0027] • determining the first parameters for each temporal bucket based on the data samples assigned to the temporal bucket.

[0028] In other embodiments, obtaining the temporal statistics includes the following operations:

[0029] • obtaining a configuration that specifies the first parameters for a subset of the temporal buckets, wherein the subset corresponds to time instances distributed unequally over the first duration; and

[0030] • performing interpolation of the first parameters for the subset to obtain the first parameters for the plurality of temporal buckets.

[0031] In some of these embodiments, the subset of the temporal buckets correspond to seasonality inflection points at which a slope of the seasonality characteristic changes sign. In some embodiments, the second parameters defining the one or more second distributions of outliers include the following: a distribution of left outliers within any of the temporal buckets, a distribution of right outliers within any of the temporal buckets, and an outlier distribution type for the distributions of left and right outliers.

[0032] In some embodiments, adjusting the subset of the plurality of synthetic data samples to represent anomalies present in data produced by the communication network is based on an anomaly configuration that includes one or more of the following:

[0033] • a range of lengths for sequences of anomalous synthetic data samples;

[0034] • a first number of sequences of left anomalous synthetic data samples to be included;

[0035] • a second number of sequences of right anomalous synthetic data samples to be included;

[0036] • a minimum value for anomalous synthetic data samples, lmin;

[0037] • a maximum value for anomalous synthetic data samples, rmax;

[0038] • a first range of values usable to produce left anomalous synthetic data samples; and

[0039] • a second range of values usable to produce right anomalous synthetic data samples.

[0040] In some embodiments, these exemplary methods also include training, or causing to be trained, a machine learning (ML) model based on the synthetic data samples that are representative of the seasonality characteristic and the anomalies present in data collected from the communication network. In some of these embodiments, the ML model is trained to perform one or more of the following based on data collected from the communication network: prediction of network PM KPIs, and detection of anomalies.

[0041] Other embodiments include computing systems (e.g., CN nodes, SMO nodes, NM nodes, cloud systems, etc.) configured to perform operations corresponding to any of the exemplary methods described herein. Other embodiments include non-transitory, computer-readable media storing program instructions that, when executed by processing circuitry, configure such computing systems to perform operations corresponding to any of the exemplary methods described herein.

[0042] These and other embodiments described herein may provide various benefits and / or advantages. For example, embodiments may effectively represent actual network traffic KPI patterns that exhibit multi-level seasonality and stochasticity in a way that existing techniques cannot. Moreover, embodiments may be computationally efficient relative to existing techniques, thereby facilitating easy scaling to accommodate the amount of synthetic data required. In addition, synthetic data generated by various embodiments can be used to evaluate and / or train anomaly detection algorithms that utilize time series KPI inputs, such as when actual time series with labelled anomalies are unavailable. While some embodiments facilitate generate the synthetic data based on actual data that exhibits multi-level seasonality and stochasticity, other embodiments facilitate generating synthetic data based on user-configurable parameters, which may be beneficial when actual data with necessary characteristics is unavailable.

[0043] These and other objects, features, and advantages of embodiments of the present disclosure will become apparent upon reading the following Detailed Description in view of the Drawings briefly described below.

[0044] BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figures 1-2 illustrate various aspects of 5G / NR network architecture.

[0046] Figure 3 shows an exemplary multi-domain network comprising a RAN, a packet-based core network (CN), and an IP Multimedia Subsystem (IMS).

[0047] Figure 4 shows a high-level functional diagram of a synthetic data generation (SDG) system according to some embodiments of the present disclosure.

[0048] Figure 5 shows a functional diagram of a Pattern Extraction function according to some embodiments of the present disclosure.

[0049] Figure 6 shows a portion of an example source dataset that includes two months of network PM KPI samples with 15-minute result output period (ROP).

[0050] Figure 7 shows a subset of the portion shown in Figure 6, together with distribution parameters for temporal buckets determined according to some embodiments of the present disclosure.

[0051] Figures 8-9 illustrate distribution parameters for temporal buckets determined according to other embodiments of the present disclosure.

[0052] Figure 10 shows a functional diagram of a Pattern Reproduction function according to some embodiments of the present disclosure.

[0053] Figures 11-12 show synthetic mid-stage data generated for a range of timestamps based on temporal statistics obtained using two different embodiments of the present disclosure.

[0054] Figure 13 shows a functional diagram of a Pattern Mutation function according to some embodiments of the present disclosure.

[0055] Figure 14 shows an example of generated synthetic data including anomalies, according to some embodiments of the present disclosure.

[0056] Figure 15 shows an example where synthetic data generated according to embodiments of the present disclosure is used to train a model.

[0057] Figure 16 shows use of the model trained in Figure 15 during an operational phase.

[0058] Figure 17 (which includes Figures 17A-B) shows an exemplary method (e.g., procedure) for generating synthetic data representative of characteristics of a communication network, according to various embodiments of the present disclosure. Figure 18 shows a communication system according to various embodiments of the present disclosure.

[0059] Figure 19 shows a network node according to various embodiments of the present disclosure.

[0060] Figure 20 shows host computing system according to various embodiments of the present disclosure.

[0061] Figure 21 is a block diagram of a virtualization environment in which functions implemented by some embodiments of the present disclosure may be virtualized.

[0062] DETAILED DESCRIPTION

[0063] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject matter disclosed herein, the disclosed subject matter should not be construed as limited to only the embodiments set forth herein; rather, these embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.

[0064] In general, all terms used herein are to be interpreted according to their ordinary meaning to a person of ordinary skill in the relevant technical field, unless a different meaning is expressly defined and / or implied from the context of use. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise or clearly implied from the context of use. The operations of any methods and / or procedures disclosed herein do not have to be performed in the exact order disclosed, unless an operation is explicitly described as following or preceding another operation and / or where it is implicit that an operation must follow or precede another operation. Any feature of any embodiment disclosed herein can apply to any other disclosed embodiment, as appropriate. Likewise, any advantage of any embodiment described herein can apply to any other disclosed embodiment, as appropriate.

[0065] Note that the description given herein focuses on a 3GPP cellular communications system and, as such, 3GPP terminology or terminology similar to 3GPP terminology is generally used. However, the concepts disclosed herein are not limited to a 3GPP system, and can be applied in any system that can benefit from the concepts, principles, and / or embodiments described herein.

[0066] Figure 1 illustrates a high-level view of an exemplary 5G network architecture, which includes a Next Generation Radio Access Network (NG-RAN, 199) and a 5G Core (5GC, 198). The NG-RAN can include one or more gNodeB’s (gNBs, e.g., 100, 150) connected to the 5GC via one or more NG interfaces (e.g., 102, 152). More specifically, the gNBs can be connected to one or more Access and Mobility Management Functions (AMFs) in the 5GC via respective NG- C interfaces and to one or more User Plane Functions (UPFs) in the 5GC via respective NG-U interfaces. Various other network functions (NFs) can be included in the 5GC, as described in more detail below.

[0067] In addition, the gNBs can be connected to each other via one or more Xn interfaces (e.g., 140 between gNBs 100, 150). The radio technology for the NG-RAN is often referred to as “New Radio” (NR). With respect to the NR interface to UEs, each of the gNBs can support frequency division duplexing (FDD), time division duplexing (TDD), or a combination thereof. Each of the gNBs can serve a geographic coverage area including one or more cells and, in some cases, can also use various directional beams to provide coverage in the respective cells.

[0068] NG RAN logical nodes shown in Figure 1 include a Centralized Unit (CU or gNB-CU) and one or more Distributed Units (DU or gNB-DU). CUs (e.g., 110) are logical nodes that host higher-layer protocols and perform various gNB functions such controlling the operation of DUs. In contrast, DUs (e.g., 120, 130) are decentralized logical nodes that host lower layer protocols and can include various subsets of gNB functions, depending on the functional split. A CU connects to one or more DUs over respective Fl logical interfaces (e.g., 122, 132 in Figure 1).

[0069] One change in 5G networks (e.g., in 5GC) is that traditional peer-to-peer interfaces and protocols found in earlier- generation networks are modified and / or replaced by a Service Based Architecture (SB A) in which Network Functions (NFs) provide one or more services to one or more service consumers. This can be done, for example, by Hyper Text Transfer Protocol / Representational State Transfer (HTTP / REST) application programming interfaces (APIs). In general, the various services are self-contained functionalities that can be changed and modified in an isolated manner without affecting other services. Furthermore, the services are composed of various “service operations”, which are more granular divisions of the overall service functionality. The interactions between service consumers and producers can be of the type “request / response” or “subscribe / notify”.

[0070] Figure 2 shows an exemplary non-roaming architecture of a 5G network (200) with service-based interfaces and various 3GPP-defined NFs, including the following:

[0071] • Application Function (AF, with Naf interface) interacts with the 5GC to provision information to the network operator and to subscribe to certain events happening in operator's network. An AF offers applications for which service is delivered in a different layer (i.e., transport layer) than the one in which the service has been requested (i.e., signaling layer), the control of flow resources according to what has been negotiated with the network. An AF communicates dynamic session information to PCF (via N5 interface), including description of media to be delivered by transport layer. • Policy Control Function (PCF, with Npcf interface) supports unified policy framework to govern the network behavior, via providing PCC rules (e.g., on the treatment of each service data flow that is under PCC control) to the SMF via the N7 reference point. PCF provides policy control decisions and flow based charging control, including service data flow detection, gating, QoS, and flow-based charging (except credit management) towards the SMF. The PCF receives session and media related information from the AF and informs the AF of traffic (or user) plane events.

[0072] User Plane Function (UPF) supports handling of user plane traffic based on the rules received from SMF, including packet inspection and different enforcement actions (e.g., event detection and reporting). UPFs communicate with the RAN (e.g., NG-RNA) via the N3 reference point, with SMFs (discussed below) via the N4 reference point, and with an external packet data network (PDN) via the N6 reference point. The N9 reference point is for communication between two UPFs.

[0073] • Session Management Function (SMF, with Nsmf interface) interacts with the decoupled traffic (or user) plane, including creating, updating, and removing Protocol Data Unit (PDU) sessions and managing session context with the User Plane Function (UPF), e.g., for event reporting. For example, SMF performs data flow detection (based on filter definitions included in PCC rules), online and offline charging interactions, and policy enforcement.

[0074] • Charging Function (CHF, with Nchf interface) is responsible for converged online charging and offline charging functionalities. It provides quota management (for online charging), re-authorization triggers, rating conditions, etc. and is notified about usage reports from the SMF. Quota management involves granting a specific number of units (e.g., bytes, seconds) for a service. CHF also interacts with billing systems.

[0075] Access and Mobility Management Function (AMF, with Namf interface) terminates the RAN CP interface and handles all mobility and connection management of UEs (similar to MME in EPC). AMFs communicate with UEs via the N1 reference point, with SMFs via the Nil reference point, and with RAN (e.g., NG-RAN) via the N2 reference point.

[0076] • Network Exposure Function (NEF) with Nnef interface - acts as the entry point into operator's network, by securely exposing to AFs the network capabilities and events provided by 3GPP NFs and by providing ways for the AF to securely provide information to 3GPP network. For example, NEF provides a service that allows an AF to provision specific subscription data (e.g., expected UE behavior) for various UEs. In general, NEF provides services similar to services provided by SCEF in EPC. • Network Repository Function (NRF) with Nnrf interface - provides service registration and discovery, enabling NFs to identify appropriate services available from other NFs.

[0077] • Network Slice Selection Function (NSSF) with Nnssf interface - enables other NFs (e.g., AMF) to identify a network slice instance that is appropriate for a UE’s desired service. A network slice instance is a set of NF instances and the required network resources (e.g., compute, storage, communication) that provide the capabilities and characteristics of a network slice.

[0078] • Authentication Server Function (AUSF) with Nausf interface - based in a user’s home network (HPLMN), it performs user authentication and computes security key materials for various purposes.

[0079] • Network Data Analytics Function (NWDAF, 210) with Nnwdaf interface - interacts with other NFs to collect relevant data and provides network analytics information (e.g., statistical information of past events and / or predictive information) to other NFs.

[0080] • Location Management Function (LMF) with Nlmf interface - supports various functions related to determination of UE locations, including location determination for a UE and obtaining any of the following: DL location measurements or a location estimate from the UE; UL location measurements from the NG RAN; and non-UE associated assistance data from the NG RAN.

[0081] The Unified Data Management (UDM) function supports generation of 3GPP authentication credentials, user identification handling, access authorization based on subscription data, and other subscriber-related functions. To provide this functionality, the UDM uses subscription data (including authentication data) stored in the 5GC unified data repository (UDR). In addition to the UDM, the UDR supports storage and retrieval of policy data by the PCF, as well as storage and retrieval of application data by NEF. The terms “UDM” and “UDM function” are used interchangeably herein.

[0082] IP Multimedia Subsystem (IMS) is an architectural framework for delivering multimedia services to wireless devices based on these Internet-centric protocols. IMS was originally specified by 3GPP in Release 5 (Rel-5) as a technology for evolving mobile networks beyond GSM, e.g., for delivering Internet services over GPRS. IMS has evolved in subsequent releases to support other access networks and a wide range of services and applications.

[0083] At a high-level, the functionality of the IMS network can be sub-divided into two types: control and media, and application enablers. The control functionality comprises Call Session Control Function (CSCF) and Home Subscriber Server (HSS). The CSCF is used for session control for devices and applications that are using the IMS network. Session control includes the secure routing of the session initiation protocol (SIP) messages, subsequent monitoring of SIP sessions, and communicating with a policy architecture to support media authorization. CSCF functionality can also be divided into Proxy CSCF (P-CSCF), Serving CSCF (S-CSCF), and Interrogating CSCF (I-CSCF).

[0084] CSCF also interacts with the Home Subscriber Server (HSS), which is a master database containing user and subscriber information to support network entities handling calls and sessions. For example, HSS provides functions such as identification handling, access authorization, authentication, mobility management (e.g., which session control entity is serving the user), session establishment support, service provisioning support, and service authorization support.

[0085] A Media Resource Function (MRF) can provide media services in a user’s home network and can manage and process media streams such as voice, video, speech-to-text, and real-time transcoding of multimedia data. In general, a WebRTC Gateway allows native- and browser-based devices to access services in the network securely.

[0086] The increasing complexity of communication networks, including 5G networks, drives the evolution of analytics systems that support operation, optimization, and planning of these networks. This includes detecting and addressing sudden, undesired changes in network operation and / or performance (e.g., failures). These analytics systems, in turn, require collecting and processing enormous amounts of data.

[0087] Advanced analytics systems, such as Ericsson Expert Analytics (EEA), are based on collecting and correlating elementary network events from different network domains, such as core, radio, and transport networks. Such analytics systems calculate user- and session-level E2E service quality metrics (S-KPIs) as well as radio and network resource metrics (R-KPIs) that characterize the radio environment or network operation at user and session level. These types of solutions are suitable for session-based troubleshooting and analysis of network issues.

[0088] Event-based analytics require real-time collection and correlation of node and protocol events from different RAN and CN nodes, probing signaling interfaces, and sampling of userplane traffic. Additionally, event-based analytics require an advanced database, a rule engine, and a “big data” analytics platform.

[0089] Figure 3 shows an exemplary multi-domain network (300) comprising a RAN, a packetbased CN, and an IMS. As shown in Figure 3, the RAN includes eNBs that provide the LTE- Uu radio interface and gNBs that provide the NR-Uu interface to UEs. The CN includes SMF, AMF, and UPF in 5GC discussed above, as well as mobility management entity (MME), serving gateway (SGW), and packet gateway (PGW) that are part of the Evolved Packet Core (EPC) associated with LTE networks. The UPF connects to the IMS via the N6 interface, such that IMS in Figure 3 is an instance of the PDN shown in Figure 2. Figure 3 also shows various “tapping points” where data can be collected from the three domains of the network. For example, node events (e.g., PM counters) can be collected from eNBs, gNBs, AMF, SMF, UPF, MME, and PGW. Likewise, interface events can be collected from S5-U (user), S5-C (control), Sl-U, and S5-U interfaces in CN as well as from Mw interface between P-CSCF and IS-CSCSF in IMS. In addition to detecting events and / or conditions at the individual nodes and / or interfaces, some more advanced analytics systems combine information collected from the multiple domains to determine “user experience” analytics that represent performance experienced by an end user for a specific service.

[0090] Detecting and addressing sudden, undesired changes in network operation and / or performance (e.g., failures) is not the only reason for network monitoring by analytics systems. As briefly mentioned above, large OTT service providers (e.g., Netflix) often have SLAs with the network operators that specify required end-user QoE for OTT services delivered over the mobile networks. As part of these SLAs, network operations must monitor end-user QoE for these services and optimize the network as necessary to meet SLA requirements.

[0091] Time series datasets can be collected from various nodes and various interface in multiple domains of a communication network. For example, values of performance measurement (PM) counters and key performance indicators (KPIs) can be collected from the various network nodes at certain time intervals. Time series data collected in this manner can be used to analyze, predict, and / or understand user behavior patterns as well as network performance trends. For example, time series data collected in this manner can be used to train a predictive model for network behavior.

[0092] As an illustrative example, generative modelling involves automatically identifying and learning regularities or patterns in input data to produce a model, which can then be used to produce new examples that might have been reasonably derived from the original dataset. Generative adversarial networks (GANs) are an approach to generative modelling that involves deep learning techniques (e.g., convolutional neural networks).

[0093] In a GAN, two neural networks contest with each other in the form of a zero-sum game, where one agent's gain is another agent's loss. Given a training set, a GAN learns to generate new data with the same statistics as the training set. For example, a GAN trained on photographs can generate new photographs that look at least superficially authentic to human observers, with many realistic characteristics. Rather than being trained to minimize the distance to a specific image, the generator model is trained indirectly to “fool” a second model, called the “discriminator”. Put differently, the generator model is trained to create new instances and the discriminator model is trained to categorize examples as either real (from the training domain) or fake (from the generator). In general, when the two models are trained together in an adversarial zero-sum game the discriminator model is fooled about half the time. This indicates that the generator model is producing believable examples.

[0094] In some cases, however, actual time series data collected from a network may be insufficient for proper training of a model, or the cost and resources needed to collect and store network-generated time series data may be prohibitive. Moreover, in case the collected time series data contains user-specific or other sensitive data, storage of such data creates risks of security breaches and non-compliance with data privacy regulations. Thus, there is a need for techniques that generate synthetic data that accurately represents the characteristics of actual time series data collected from a network.

[0095] Use of synthetic data for various network and service operations can have various advantages. For example, as briefly mentioned above, synthetic data ensures compliance with data protection regulations by creating realistic but artificial datasets that safeguard sensitive customer information. Synthetic data also reduces the risk of data breaches by facilitating testing and development without exposing real customer data.

[0096] As another example, synthetic data can provide scalable and diverse datasets that facilitate thorough testing, in contrast to the limitations in coverage of actual data. As such, synthetic data can be used to train algorithms and models to recognize a wider range of conditions and / or scenarios compared with only using actual data for training. Similarly, models trained based on synthetic data can better facilitate network optimization and capacity planning.

[0097] As another example, synthetic data can be used to simulate user behavior to inform new product and / or service development, thereby reducing risks that can occur after product and / or service deployment. Similarly, synthetic data can facilitate comparing service performance with competitors, aiding strategic decisions and improvements. As another example, synthetic data generation can reduce the need for collection of actual data, which in some cases can be very time and / or resource intensive.

[0098] At a high level, synthetic data generation (SDG) of time series can be univariate or multivariate. Univariate SDG captures temporal characteristics of a data stream representing a single variable, along with statistical patterns exhibited by that variable. Multivariate SDG captures patterns in a system of multiple variables along with any relationships among the variables. Other than correlation statistics and joint-distribution goodness-of-fit tests, however, there are few techniques to validate whether a multivariate SDG generates synthetic time series data that follows the characteristics of a source dataset. Furthermore, multivariate SDGs have higher computational demands making them unsuitable to work with very large amounts of data (“big data”). Accordingly, SDG may be preferred for various applications. Several existing solutions are available for univariate SDG. Synthetic Data Vault (SDV) is a Python-based library of multiple SDGs including Gaussian Copula, Conditional Tabular GAN (CTGAN), Copula GAN, and Variational Autoencoder. Although these SDGs are able to work with sequential data, they are unable to provide data that has exactly one datapoint per interval and / or to capture seasonal trends (“seasonality”) involving temporal dependencies, as often found in actual data collected from communication networks.

[0099] The Mostly Al SDG learns patterns from original data to create realistic fictional datasets for customer profiles, transactions, etc. Although its techniques are not publicly disclosed, a related U.S. Pat. Pub. 2022 / 0237522 only discloses techniques for synthesizing mobility traces, which is a small subset of actual data that may be collected from a network and / or synthetic data needed for network-related purposes discussed above.

[0100] Various GANs have been adapted for SDG. For example, TSGAN uses a conditional GAN architecture while SigWGAN combines a Logsig-RNN with a Signature Wassersetein-1 metric. TimeGAN merges the adaptability of unsupervised GAN framework with the control provided by supervised training in autoregressive models.

[0101] Even so, these and other existing techniques for SDG have various problems, issues, and / or drawbacks that that make them unable to replicate characteristics of actual time series data collected from a network, such as seasonality characteristics and anomalies typical present in the collected data.

[0102] For example, existing techniques do not capture nuances of traffic variations across time of day and days of week, nor do they distinguish between seasonality and seasonally independent randomness. Furthermore, most existing techniques produce data with inconsistent frequency. Network KPI data has one data point (or record) for a given instant (e.g., timestamp) or period (e.g., result output period or ROP). In contrast, existing techniques may produce zero, one, or multiple records per ROP. In other words, there may be missing records in the synthetic data.

[0103] Furthermore, existing techniques provide no way to include specific anomalies (and corresponding descriptions) in the synthetic data, which may be useful for model training and validation. Additionally, most existing techniques are not “explainable”, in that it is not clearly understood how they generate the synthetic data given certain inputs.

[0104] Embodiments of the present disclosure address these and other problems, issues, and / or difficulties by novel, flexible, and efficient techniques for univariate SDG of time series data suitable for various uses associated with a communication network, such as training of models used for performance analysis, anomaly detection, etc. Some embodiments include an SDG procedure that may be implemented by a network node or function associated with a communication network, involving the following operations at a high level: • Pattern Extraction: extract seasonal statistics with stochastic variation from sample data or data configuration using temporal bucketing.

[0105] • Pattern Reproduction: synthesize data based on the extracted statistics for the user input synthesis period using bucket-specific distribution parameters.

[0106] • Pattern Mutation: inject anomalies into the synthetized data based on an anomaly configuration.

[0107] In more detail, the SDG assigns timestamps to temporal groups (“buckets”) and associates each bucket with a distribution that effectively capture multi-level seasonality of the bucket. In addition to the temporal buckets, the SDG also captures an expected outlier distribution that is independent of seasonality in the data; this outlier distribution is also referred to as seasonally independent stochasticity (or randomness). Both the bucket-specific distributions and the outlier distributions can be based on actual data collected from a network or a user-specified configuration.

[0108] Additionally, the SDG injects and labels anomalies into the synthesized time series. The anomalies can include isolated anomalies and / or sustained anomalies, according to a user- specified configuration. In some embodiments, the tail distributions (left, right, or both) of the anomalies can also be according to a user-specified configuration. Also, in some embodiments, the deviations of the anomalies can be dependent on the distribution of the temporal buckets of respective data points where the anomalies are injected.

[0109] Embodiments can provide various benefits and / or advantages. For example, embodiments of the SDG may effectively represent actual network traffic KPI patterns that exhibit multi-level seasonality and stochasticity in a way that existing techniques cannot. Moreover, embodiments of the SDG can be computationally efficient relative to existing techniques, thereby facilitating easy scaling to accommodate the necessary amount of synthetic data without requiring adaptation of complex neural networks such as GANs.

[0110] In addition, synthetic time series data generated by various embodiments can be used to evaluate and / or train anomaly detection algorithms that utilize time series KPI anomaly inputs, such as when actual data with labelled anomalies are unavailable. Put differently, injecting and labelling synthetic anomalies is significantly less complex than detecting and labeling anomalies in an existing dataset.

[0111] Furthermore, some embodiments provide the capability to generate the synthetic data based on actual data that exhibits multi-level seasonality and stochasticity. For example, embodiments enable a user to build an arbitrarily large dataset derived from a relatively small source dataset, which may be beneficial when actual data is scarce, difficult or expensive to collect, or subject to geographical and / or privacy constraints. Other embodiments provide the capability to generate the synthetic data based on user- configurable multi-level seasonality and stochasticity parameters. For example, these embodiments facilitate data synthesis when actual data is unavailable, which may be beneficial when privacy regulations restrict the sharing of user-identifiable data or its derivatives.

[0112] Figure 4 shows a high-level functional diagram of a synthetic data generation (SDG) system (400) according to some embodiments of the present disclosure. This exemplary system includes various functions that are described in more detail below. Note that these functions of an SDG may also be referred to by equivalent terms such as “module”, “sub-system”, “component”, etc. Each function (or equivalent term) may be implemented by any appropriate combination of hardware and / or software.

[0113] The Pattern Extraction function (410) extracts seasonality information and statistics of non-seasonal residuals from an input sample (or source) time series data. Alternately, the Pattern Extraction function generates seasonality information and statistics of non-seasonal residuals based on a user- specified configuration. The output of this function may be referred to as “temporal statistics”.

[0114] The Pattern Reproduction function (420) uses the temporal statistics to synthesize time series data for a given time period (also referred to as “synthesis period”), with a user- configurable sample granularity in time (e.g., one per hour, etc.). This synthesized time series data may be referred to as “mid-stage data”. The Pattern Mutation function (430) injects anomalies into the mid-stage data based on a user-specified configuration, with the output of this function being the synthetic time series data that can be used for various network-related purposes. For example, the synthetic time series data that can be network KPIs (with anomalies) over the synthesis period.

[0115] Figure 5 shows a functional diagram of a Pattern Extraction function (500), which is an example of Pattern Extraction function (410) shown in Figure 4.

[0116] In block 510, the Pattern Extraction function checks whether a source dataset (i.e., actual univariate time series data) is available and meets minimum size requirements. The required minimum size of the of the source dataset depends on time granularity or frequency (i.e., time interval between each data point) as well as a degree of seasonality expected to detected and / or generated. For example, network performance management (PM) KPIs are typically provided at 15 minute intervals (i.e., ROP). These KPIs generally exhibit both daily and weekly seasonality. In such case, it may be beneficial to have at least one month and preferably at least two months of source data available to be used as a basis for SDG. For KPIs or other data with a daily interval, it may be beneficial to have at least two years of to be used as a basis for the SDG, e.g., to be able to capture both weekly and yearly seasonality. If a sufficient amount of source data is available, the source dataset is obtained and provided to a Temporal Bucketing sub-function (520). This function assigns data values into groups (or “buckets”) based on the level of seasonality required for synthetic data. Table 1 below shows an example of temporal buckets or groupings according to some embodiments.

[0117] Table 1.

[0118] Each entry in the minimum frequency column gives the lowest time difference between adjacent samples in the time series needed to capture the degree of seasonality in the same row. Each entry in the right-most column suggests groupings consistent with the seasonality and the minimum frequency of the same row.

[0119] Some embodiments may support multi-level seasonality characterization. For example, if both weekly and daily seasonality should be considered, then the values are grouped into [day-of- week, hour-of-day] buckets. For network PM KPIs, the source data can further be grouped into minute-of-hour buckets level to capture hourly seasonality, i.e., [day-of-week, hour-of-day, minute-of-hour]. The degree of multi-level seasonality used can be pre-configured or part of a user-specified configuration.

[0120] The output of Temporal Bucketing is provided to the Seasonal Analysis sub-function (530), which determines a statistical distribution for all source data entries in each bucket. The type of distribution (e.g., Gaussian, Poisson, Uniform, etc.) can be pre-configured or estimated based on the actual source data using goodness-of-fit tests such as described in “Test of Hypothesis - Concise Formula Summary” by E. Isaac, published in 2015.

[0121] The Outlier Analysis sub-function (540) then determines a distribution of outliers across all buckets for the source dataset. Note that an “outlier” in a bucket is relative to the distribution determined for that bucket. If a bucket distribution has a mean / i and a standard deviation o. an “outlier” for this bucket is a value that is at least Uo away from the mean, where k can be preconfigured or based on a user-specified configuration. As an example, with k = 1.5, right-tailed outliers for each bucket with (p, o) are values greater than p + 1.5 a and left-tailed outliers are values less than p — ko.

[0122] More specifically, the Outlier Analysis sub-function outlier probability for each bucket, then determines a distribution type for all outliers in all buckets, and finally distribution parameters of right-tailed and left-tailed outliers, respectively, in all buckets. Accordingly, when the Pattern Extraction function utilizes a source dataset, the following temporal statistics output by the Pattern Extraction function:

[0123] • Distribution type for each bucket (e.g., Gaussian, Poisson, Uniform, etc);

[0124] • Distribution parameters for each bucket, including (p, o) o Distribution parameters, and o Outlier probabilities, P(0BR) and

[0125] • Outlier Statistics o Outlier distribution type, o Left-tailed outlier distribution parameters, and o Right-tailed outlier distribution parameters.

[0126] Note that the outlier statistics represent the stochasticity or randomness that is independent of the seasonality in the source dataset. The left- and right-tailed distribution parameters may include probability of each type of outlier (e.g., P(0L), P(0R) ) as well as parameters describing the distribution shape according to the outlier distribution type (e.g., (p, o) for Gaussian distribution).

[0127] The following example further illustrates that operation of the Pattern Extraction function based on a source dataset. A source dataset of network PM KPI samples with 15-minute ROP is used, which provides 96 ROPs per day and 672 ROPs per week. Thus, the following are a set of 672 unique temporal buckets during a week: {Monday, 00:00}, {Monday, 00:15}, ... {Monday, 23:45}, {Tuesday, 00:00}, {Tuesday, 00:15}, ... {Sunday, 23:30}, {Sunday 23:45}. Each temporal buckets can be assigned a numerical identifier (ID), e.g., 1 to 672 corresponding to temporal order.

[0128] Figure 6 shows a portion of a source dataset that includes two months of network PM KPI samples with 15-minute ROP. Given this dataset, the Temporal Bucketing sub-function groups its samples into [day-of-week, hour-of-day] buckets. In other words, Temporal Bucketing assigns to each sample a bucket ID corresponding to the time-of-day and day-of-week associated with that sample.

[0129] In this example, each temporal bucket is assumed to have a uniform distribution represented by parameters a and b, which correspond to the minimum and maximum values of the distribution. Seasonality Analysis then computes the distribution parameters, a and b, for each temporal bucket. Figure 7 shows a subset of the portion of the source dataset shown in Figure 6, spanning a period of approximately three days. For each sample during this period, Figure 7 also shows distribution parameters (a, b) determined for the temporal [day-of-week, hour-of-day] bucket to which the sample is assigned. The probabilities of outliers P(0BR) and P(0BL) for each bucket are also determined, followed by the distribution type and parameters for outliers in all buckets, as discussed above. Returning to Figure 5, if it is determined in block 510 that a sufficient amount of source data is not available, then the seasonality generation functionality in blocks 550-570 is utilized. Rather than using a source dataset, the Temporal Bucketing function (550) uses a built-in or user- specified configuration to generate temporal statistics for the buckets, including per-bucket distribution type / parameters and per-bucket outlier statistics. However, it is not necessary to explicitly specify all buckets in the configuration. In some embodiments, the configuration needs to only specify distribution parameters for a subset of buckets that correspond to seasonality inflection points, at which the slope of the seasonality changes sign (e.g., positive-to-negative). The parameters for other buckets can be generated based on this reduced configuration input.

[0130] For example, the Linear Imputation sub-function (560) can be used to interpolate parameters of the buckets that are not explicitly specified in the configuration, and the Curve Fitting sub-function (570) can be used to smooth variations of the linearly-interpolated parameters near inflection points. The Curve Fitting sub-function outputs the same or similar temporal statistics as extracted from a source dataset, discussed above.

[0131] The following example further illustrates that operation of the Pattern Extraction function based on a configuration, e.g., when an adequate source dataset is not available. In this example, the configuration specifies various parameters corresponding to the seasonality inflection points of the time series data to be generated.

[0132] Although it is possible for the configuration to specify a per-bucket distribution for each day, this example assumes that each weekday follows a first distribution and each day of the weekend follows a second distribution, i.e., different than the first distribution. Accordingly, only two sets of buckets - weekday and weekend - need to be specified in the configuration. If each bucket for a given day can be encoded as (time-of-day, a, b), Table 2 below shows an exemplary configuration with parameters for each bucket corresponding to a seasonality inflection point.

[0133] Table 2.

[0134] Outlier probabilities for each of these buckets corresponding to a seasonality inflection point can also be specified in a similar way (e.g., as additional parameters in the tuples above), with the outlier statistics across all buckets being specified separately in the configuration. Given this input, the corresponding parameters for the other buckets will be produced by the Linear Imputation sub-function. Figure 8 shows a plot of distribution parameters (a, b) for the various buckets of a weekday distribution, after applying the Linear Imputation. A similar plot is easily obtained for a weekend distribution.

[0135] In this example, the Curve Fitting sub-function applies a Savitzky-Golay filter to the output of the Linear Imputation sub-function, thereby producing more realistic (e.g., smoother) transitions at the various inflection points in the weekday and weekend buckets. Figure 9 shows a plot of distribution parameters (a, b) for the various buckets of a weekday distribution, after applying the Curve Fitting.

[0136] Figure 10 shows a functional diagram of a Pattern Reproduction function (1000), which is an example of the Pattern Reproduction function (420) shown in Figure 4.

[0137] The Bucket Expansion sub-function (1010) expands the possible buckets for the configured synthesis period, as needed, based on the configured level(s) of seasonality. This produces a data frame indexed by timestamps along with the bucket ID assigned to each timestamp. For example, the output of Bucket Expansion can be a set of {time, bucket ID} tuples.

[0138] In block 1020, the temporal statistics for each bucket are appended to each timestamp during the data frame. This includes the distribution parameters (e.g., a, b) and the right / left outlier probabilities P(0BR) and P(0BL) for each bucket. This sub-function is also referred to as Attribute Appendage and may facilitate parallelism in subsequent operations described below. For example, in the case of a uniform distribution with parameters (a, b), the output of this operation can be a set of {time, bucket ID, a, b, P(0BR), P(OBL) } tuples.

[0139] The Bounded Random Number Generator (RNG) sub-function (1030) builds the distribution for each bucket using the distribution type from the temporal statistics and the appended parameters for each timestamp. This RNG is “bounded” in the sense that it can be configured to not produce outliers, which are added later by the Data Coloration sub-function (1040). More specifically, for each of the input tuples, the Bounded RNG sub-function generates a synthetic data value according to the distribution parameters in the tuple, e.g., a value between (a, b) with a uniform distribution. This generated data value can be appended to the tuple to form a new tuple, e.g., {time, bucket ID, a, b, P(0BR), P(0BL), value}. The Data Coloration sub-function uses the outlier statistics to adjust or perturb (“color”) the values generated by the Bounded RNG sub-function. For example, if P(0BR) is the right-tailed outlier probability of bucket B and P (0R') is the right-tailed outlier probability of the entire dataset, the right-tailed outliers are added by the Data Coloration sub-function to bucket B with the probability P(0BR) ■ P(0R). Similarly, if P(0BL) is the left-tailed outlier probability of bucket B and P(0L) is the left-tailed outlier probability of the entire dataset, the left-tailed outliers are added by the Data Coloration sub-function to bucket B with probability P(0BL) ■ P(0L). When outliers are added, the values are generated according to the outlier distribution parameters that are part of the temporal statistics. The resulting data frame may be referred to as “mid-stage data”.

[0140] Figure 11 shows synthetic mid-stage data generated for a range of timestamps based on temporal statistics extracted from the exemplary two-month source dataset discussed above, of which Figure 6 shows a portion. Figure 12 shows synthetic mid-stage data generated for a range of timestamps based on temporal statistics generated based on a configuration, as illustrated in Figures 8-9.

[0141] Figure 13 shows a functional diagram of a Pattern Mutation function (1300), which is an example of the Pattern Mutation function (430) shown in Figure 4.

[0142] The Candidate Grouping sub-function (1310) groups sequences of timestamps that are candidates for addition of anomalies according to an anomaly configuration, which can include the following configuration parameters:

[0143] • Sustained Range, (s1;s2) for each anomalous occurrence, where s2> s2;

[0144] • Minimum value for left-tail anomalies, lmin> 0;

[0145] • Maximum value for right-tail anomalies, > 0;

[0146] • Number of left-tail anomalous occurrences to be added;

[0147] • Number of right-tail anomalous occurrences to be added;

[0148] • Left-tail anomaly multipliers, and l2; and

[0149] • Right-tail anomaly multipliers, and r2.

[0150] An anomalous occurrence is a sequence of anomalous values with < sequence length < s2.

[0151] Selection of candidate timestamps for addition of anomalies can be based on one or more of the following criteria, in accordance with the configuration parameters:.

[0152] • Each candidate must be a series of contiguous intervals within the Sustained Range;

[0153] • Candidates for left-tailed anomalies must have values greater than 11O’ +

[0154] • Candidates for right-tailed anomalies must have values lesser than rxo — rmax; and

[0155] • Number of left- and right-tailed anomalies added must be consistent with the corresponding configuration parameters. Note that sigma (o) in the above selection criteria represents the standard deviation of the distribution for a given bucket.

[0156] Once the candidates are selected, the Candidate Grouping sub-function adds anomaly label attributes to the mid-stage data entries, e.g., {time, bucket ID, a, b, P(0BR), P(0BL), value, anomaly_label}. The default value for the anomaly_label is “0”, which indicates “not anomalous”. Right-tailed anomalies are indicated by “1” and left-tailed anomalies are indicated by “-1”, which will be added in a subsequent operation below.

[0157] Subsequently, the left-tailed anomalies are added by the Left-Tail Injection sub-function (1320) and the right-tailed anomalies are added by the Right-Tail Injection sub-function (1330). If M is a time series and M(t) is the value of M at time t, then the Left-Tail Injection performs the following operations:

[0158] • Randomly select value d G [Z1;Z2];

[0159] • M(t) «- M(t) — do-,

[0160] • If M(t) < lmin, then

[0161] • Specify that a left-tailed anomaly is injected by changing anomaly_label attribute for time t to “-1”.

[0162] Similarly, the Right-Tail Injection performs the following operations:

[0163] • Randomly select value d G [r1;r2];

[0164] • M(t) «- M(t) + do-,

[0165] • If M(t) > rmax, then M(t) «- rmax; and

[0166] • Specify that a right-tailed anomaly is injected by changing anomaly_label attribute for time t to “1”.

[0167] After anomaly injection, the Attribute Removal sub-function (1340) eliminates from the tuples the attributes that are no longer needed (e.g., distribution parameters and outlier probability) leaving only the {time, value, anomaly_label} tuples.

[0168] The following example further illustrates operation of the Pattern Mutation function, based on the following example anomaly configuration:

[0169] • Sustained Range (4, 7)

[0170] • min 0

[0171] •rmax=None (or infinity)

[0172] • Number of left-tail anomalous occurrences = 2

[0173] • Number of right-tail anomalous occurrences = 2

[0174] • Left-tail multipliers (2.4, 4)

[0175] • Right-tail multipliers (2.4, 4) In this example, Candidate Grouping selects two sequences of timestamps for left-tailed anomalies and two sequences of timestamps for right-tailed anomalies. Each anomalous sequence is four to seven consecutive ROPs. Left-Tail and Right-Tail injection injects anomalous values at the ROPs of the selected sequences, based on the relevant configuration parameters. Figure 14 shows a portion of the final synthetic data with anomalies added, as indicated by the circles.

[0176] To summarize, embodiments of the present disclosure can generate synthetic data that is representative of actual data collected from a communication network, e.g., network PM KPIs. This synthetic data can be used for various purposes associated with operation of the communication network and / or development of products and / or services supported by the communication network. Figure 15 shows an example where synthetic data generated by an SDG (1510) according to embodiments of the present disclosure is used to train a model (W), which may be an artificial intelligence / machine learning (AI / ML) model. As shown in Figure 15, the data synthesis can be based on either actual collected data (e.g., KPIs) or a configuration for synthetic data, in accordance with the above description of these two options. The synthetic data used to train the model may include labels identifying samples in which anomalies have been injected, as discussed above. By taking into account these labels during training, the model can be tuned to recognize anomalies in unlabeled data (e.g., KPIs) collected later from the communication network.

[0177] Figure 16 shows use of the trained model (W) during an operational phase. Unlabeled data (e.g., KPIs) is collected from the network and input to the trained model. The model outputs predictions based on these inputs. For example, the model may predict future KPIs (or trends). Alternately or in addition, the model may detect unlabeled anomalies in the collected data. The output of the model can be used, e.g., by an operations / administration / management (0AM) function in the communication network to adjust configurations of network elements, shut down faulty network elements, move data traffic between network elements, etc.

[0178] Although Figures 15-16 show training and operation of the model as separate phases, skilled persons will readily comprehend that training can also be performed during the operational phase, based on actual collected data and / or synthetic data generated by an SDG (1610) according to embodiments of the present disclosure.

[0179] Various features of the embodiments described above correspond to various operations illustrated in Figure 17 (including parts A and B), which depicts an exemplary method e.g., procedure) for generating synthetic data representative of characteristics of a communication network, according to various embodiments of the present disclosure. In other words, various features of the operations described below correspond to various embodiments described above. Although Figure 17 shows specific blocks in a particular order, the operations of the exemplary method can be performed in a different order than shown and can be combined and / or divided into blocks having different functionality than shown. Optional blocks or operations are indicated by dashed lines.

[0180] The following description is based on the exemplary method being performed by a computing system, which may be associated with and / or part of the communication. For example, the computing system can be a service management and orchestration (SMO) system for a RAN, an analytics-related CN node such as NWDAF, a network management node in an OAM system, or a host computing system external to the network (e.g., public or private cloud environment).

[0181] The exemplary method can include the operations of block 1710, where the computing system determines a plurality of temporal buckets corresponding to time instances that are equally spaced over a first duration at a frequency associated with a seasonality characteristic of data collected from the communication network. The exemplary method also includes the operations of block 1720, where the computing system obtains the following temporal statistics for the synthetic data:

[0182] • for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and

[0183] • second parameters defining one or more second distributions of outliers from the first distributions associated with the plurality of temporal buckets.

[0184] The exemplary method also includes the operations of block 1730, where based on the temporal statistics, the computing system generates a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic. The exemplary method includes the operations of block 1740, where the computing system adjusts a subset of the plurality of synthetic data samples to represent anomalies present in data collected from the communication network.

[0185] In some embodiments, the synthetic data samples are representative of network PM key KPI samples generated by the communication network. In some embodiments, each time instance associated with a synthetic data sample corresponds to one of the temporal buckets and each synthetic data sample is generated based on the first parameters of the corresponding temporal bucket. In some of these embodiments, the first parameters defining the first distribution for each temporal bucket include the following:

[0186] • a distribution type of the first distribution;

[0187] • a minimum of the first distribution;

[0188] • a maximum of the first distribution;

[0189] • probability of a left outlier that is less than the minimum; and

[0190] • probability of a right outlier that is greater than the maximum. In some variants of these embodiments, for each bucket, the distribution type of the first distribution is one of the following: uniform, Gaussian, or Poisson. Even so, skilled persons will recognize that these distribution types are merely examples and any distribution type appropriate for the first distributions can be used.

[0191] In some variants of these embodiments, generating the plurality of synthetic data samples based on the temporal statistics in block 170 includes the following operations (labelled with corresponding sub-block numbers) for each synthetic data sample:

[0192] • (1731) generating, as the synthetic data sample, a random value according to the first distribution for the temporal bucket corresponding to the time instance associated with the synthetic data sample;

[0193] • (1732) selectively modifying the generated synthetic data sample to be a left outlier less than the minimum of the first distribution for the corresponding temporal bucket, based on the probability of a left outlier for the corresponding temporal bucket; and

[0194] • (1733) selectively modifying the generated synthetic data sample to be a right outlier greater than the maximum of the first distribution for the corresponding temporal bucket, based on the probability of a right outlier for the corresponding temporal bucket.

[0195] In some further variants, selectively modifying the generated synthetic data sample to be a left outlier in sub-block 1732 is further based on a probably of a left outlier occurring in any of the temporal buckets. Likewise, selectively modifying the generated synthetic data sample to be a right outlier in sub-block 1733 is further based on a probably of a right outlier occurring in any of the temporal buckets.

[0196] In some of these embodiments, adjusting the subset of the plurality of synthetic data samples to represent anomalies present in data produced by the communication network in block 1740 is based on an anomaly configuration that includes one or more of the following:

[0197] • a range of lengths for sequences of anomalous synthetic data samples;

[0198] • a first number of sequences of left anomalous synthetic data samples to be included;

[0199] • a second number of sequences of right anomalous synthetic data samples to be included;

[0200] • a minimum value for anomalous synthetic data samples, lmin;

[0201] • a maximum value for anomalous synthetic data samples, rmax;

[0202] • a first range of values usable to produce left anomalous synthetic data samples; and

[0203] • a second range of values usable to produce right anomalous synthetic data samples.

[0204] In some variants of these embodiments, adjusting the subset of the plurality of synthetic data samples in block 1740 includes the following operations, labelled with corresponding subblock numbers: • (1741) selecting the first and second number of sequences of the synthetic data samples, wherein each synthetic data sample of the selected first number of sequences is at least a first amount greater than lmtn, and each synthetic data sample of the selected second number of sequences is at least a second amount less than rmax;

[0205] • (1742) decreasing respective synthetic data samples of the selected first number of sequences, by respective first amounts determined based on the first range of values; and

[0206] • (1743) increasing respective synthetic data samples of the selected second number of sequences, by respective second amounts determined based on the second range of values. In some further variants, each of the first and second ranges is a range of scaling factors, and one or more of the following applies:

[0207] • each first amount is determined based on the following: a scaling factor randomly selected from the first range, a standard deviation for the temporal bucket to which the synthetic data sample is assigned,

[0208] • each second amount is determined based on the following: a scaling factor randomly selected from the second range, a standard deviation for the temporal bucket to which the synthetic data sample is assigned, and rmax.

[0209] In some of these embodiments, obtaining the temporal statistics in block 1720 includes the following operations, labelled with corresponding sub-block numbers:

[0210] • (1721) obtaining a source dataset produced by the communication network, wherein the source dataset includes a plurality of data samples associated with time instances that are equally spaced at the frequency over a reference duration;

[0211] • (1722) assigning each data sample to one of the temporal buckets based on a correspondence between the time instance associated with the data sample and the time instance associated with the assigned temporal bucket; and

[0212] • (1733) determining the first parameters for each temporal bucket based on the data samples assigned to the temporal bucket.

[0213] For example, the source dataset can include a plurality of network PM KPI samples generated by the communication network.

[0214] In some variants of these embodiments, determining the first parameters for each temporal bucket in sub-block 1733 includes the following operations:

[0215] • determining the distribution type, the maximum, and the minimum of the first distribution based on the data samples assigned to the temporal bucket;

[0216] • determining the probability of a left outlier based on how many of the assigned data samples are less than the determined minimum; and • determining the probability of a right outlier based on how many of the assigned data samples are greater than the determined minimum.

[0217] In some further variants, the second parameters defining the one or more second distributions of outliers include the following: a distribution of left outliers within any of the temporal buckets, a distribution of right outliers within any of the temporal buckets, and an outlier distribution type for the distributions of left and right outliers. In some further variants, obtaining the temporal statistics in block 1720 also includes the operations of sub-block 1724, where the computing system determine the second parameters defining the one or more second distributions, based on the following:

[0218] • for each of the temporal buckets, respective first differences between the minimum of the first distribution and the assigned data samples that are less than the minimum; and

[0219] • for each of the temporal buckets, respective second differences between the maximum of the first distribution and the assigned data samples that are greater than the maximum. For example, determining the second parameters defining the one or more second distributions in block 1724 can include the following operations:

[0220] • determining the distribution of left outliers based on the first differences and the outlier distribution type; and

[0221] • determining the distribution of right outliers based on the second differences and the outlier distribution type.

[0222] In other embodiments, obtaining the temporal statistics in block 1720 includes the following operations, labelled with corresponding sub-block numbers:

[0223] • (1725) obtaining a configuration that specifies the first parameters for a subset of the temporal buckets, wherein the subset corresponds to time instances distributed unequally over the first duration; and

[0224] • (1726) performing interpolation of the first parameters for the subset to obtain the first parameters for the plurality of temporal buckets.

[0225] In some of these embodiments, the subset of the temporal buckets corresponds to seasonality inflection points at which a slope of the seasonality characteristic changes sign. In some of these embodiments, performing interpolation in sub-block 1726 includes the following operations:

[0226] • performing linear imputation of the specified first parameters for the subset to obtain first parameters for the temporal buckets not included in the subset; and

[0227] • performing curve fitting on the specified first parameters and the linearly imputed first parameters to obtain the first parameters for the plurality of temporal buckets. In some of these embodiments, the subset of temporal buckets comprises a first subset that correspond to weekday time instances and a second subset that correspond to weekend time instances. In such case, performing interpolation in sub-block 1726 includes the following operations:

[0228] • performing interpolation of the first parameters for the first subset to obtain the first parameters for all temporal buckets that correspond to weekday time instances; and

[0229] • performing interpolation of the first parameters for the second subset to obtain the first parameters for all temporal buckets that correspond to weekend time instances.

[0230] In some embodiments, the seasonality characteristic is an integer multiple of the frequency. In some embodiments, the seasonality characteristic is one or more of the following: hourly, daily, weekly, monthly, quarterly, and yearly. Table 1 above shows various examples of these embodiment. In some embodiments, each temporal bucket corresponds to a plurality of time instances of the second duration (i.e., of the synthetic data).

[0231] In some embodiments, the exemplary method also includes the operations of block 1750, where the computing system trains, or causes to be trained, an ML model based on the synthetic data samples that are representative of the seasonality characteristic and the anomalies present in data collected from the communication network. In some of these embodiments, the ML model is trained to perform one or more of the following based on data collected from the communication network: prediction of network PM KPIs, and detection of anomalies.

[0232] Although various embodiments are described herein above in terms of methods, apparatus, devices, computer-readable medium and receivers, the person of ordinary skill will readily comprehend that such methods can be embodied by various combinations of hardware and software in various systems, communication devices, computing devices, control devices, apparatuses, non-transitory computer-readable media, etc.

[0233] Figure 18 shows an example of a communication system 1800 in accordance with some embodiments. In this example, communication system 1800 includes a telecommunication network 1802 that includes an access network 1804 (e.g., RAN) and a core network 1806, which includes one or more core network nodes 1808. Access network 1804 includes one or more access network nodes, such as network nodes 1810a-b (one or more of which may be generally referred to as network nodes 1810), or any other similar 3GPP access nodes or non-3GPP access points. Moreover, as will be appreciated by those of skill in the art, a network node is not necessarily limited to an implementation in which a radio portion and a baseband portion are supplied and integrated by a single vendor. Thus, it will be understood that network nodes include disaggregated implementations or portions thereof. For example, in some embodiments, telecommunication network 1802 includes one or more Open-RAN (ORAN) network nodes. An ORAN network node is a node in telecommunication network 1802 that supports an ORAN specification (e.g., a specification published by the O-RAN Alliance, or any similar organization) and may operate alone or together with other nodes to implement one or more functionalities of any node in telecommunication network 1802, including one or more network nodes 1810 and / or core network nodes 1808.

[0234] Examples of an ORAN network node include an open radio unit (O-RU), an open distributed unit (O-DU), an open central unit (O-CU), including an O-CU control plane (O-CU- CP) or an O-CU user plane (O-CU-UP), a RAN intelligent controller (near-real time or non-real time) hosting software or software plug-ins, such as a near-real time control application (e.g., xApp) or a non-real time control application (e.g., rApp), or any combination thereof (the adjective “open” designating support of an ORAN specification). The network node may support a specification by, for example, supporting an interface defined by the ORAN specification, such as an Al, Fl, Wl, El, E2, X2, Xn interface, an open fronthaul user plane interface, or an open fronthaul management plane interface. Moreover, an ORAN access node may be a logical node in a physical node. Furthermore, an ORAN network node may be implemented in a virtualization environment (described further below) in which one or more network functions are virtualized. For example, the virtualization environment may include an O-Cloud computing platform orchestrated by a Service Management and Orchestration Framework via an 0-2 interface defined by the O-RAN Alliance or comparable technologies. Network nodes 1810 facilitate direct or indirect connection of UEs, such as by connecting UEs 1812a-d (one or more of which may be generally referred to as UEs 1812) to core network 1806 over one or more wireless connections.

[0235] In some embodiments, telecommunication network 1802 can also include one or more Network Management (NM) nodes 1818, which can be part of an operation support system (OSS), a business support system (BSS), and / or an operation / administration / maintenance (0AM) system. The NM nodes can monitor and / or control operations of other nodes in access network 1804 and core network 1806. Although not shown in Figure 18, NM node 1818 is configured to communicate with other nodes in access network 1804 and core network 1806 for these purposes.

[0236] Example wireless communications over a wireless connection include transmitting and / or receiving wireless signals using electromagnetic waves, radio waves, infrared waves, and / or other types of signals suitable for conveying information without the use of wires, cables, or other material conductors. Moreover, in different embodiments, communication system 1800 may include any number of wired or wireless networks, network nodes, UEs, and / or any other components or systems that may facilitate or participate in the communication of data and / or signals whether via wired or wireless connections. Communication system 1800 may include and / or interface with any type of communication, telecommunication, data, cellular, radio network, and / or other similar type of system.

[0237] UEs 1812 may be any of a wide variety of communication devices, including wireless devices arranged, configured, and / or operable to communicate wirelessly with network nodes 1810 and other communication devices. Similarly, network nodes 1810 are arranged, capable, configured, and / or operable to communicate directly or indirectly with UEs 1812 and / or with other network nodes or equipment in telecommunication network 1802 to enable and / or provide network access, such as wireless network access, and / or to perform other functions, such as administration in telecommunication network 1802.

[0238] In the depicted example, core network 1806 connects network nodes 1810 to one or more hosts, such as host 1816. These connections may be direct or indirect via one or more intermediary networks or devices. In other examples, network nodes may be directly coupled to hosts. Core network 1806 includes one or more core network nodes (e.g., 1808) that are structured with hardware and software components. Features of these components may be substantially similar to those described with respect to the UEs, network nodes, and / or hosts, such that the descriptions thereof are generally applicable to the corresponding components of core network node 1808. Example core network nodes include functions of one or more of a Mobile Switching Center (MSC), Mobility Management Entity (MME), Home Subscriber Server (HSS), Access and Mobility Management Function (AMF), Session Management Function (SMF), Authentication Server Function (AUSF), Subscription Identifier De-concealing function (SIDE), Unified Data Management (UDM), Security Edge Protection Proxy (SEPP), Network Exposure Function (NEF), and / or a User Plane Function (UPF).

[0239] Host 1816 may be under the ownership or control of a service provider other than an operator or provider of access network 1804 and / or telecommunication network 1802, and may be operated by the service provider or on behalf of the service provider. Host 1816 may host a variety of applications to provide one or more service. Examples of such applications include live and pre-recorded audio / video content, data collection services such as retrieving and compiling data on various ambient conditions detected by a plurality of UEs, analytics functionality, social media, functions for controlling or otherwise interacting with remote devices, functions for an alarm and surveillance center, or any other such function performed by a server.

[0240] In some embodiments, access network 1804 can include a service management and orchestration (SMO) system or node 1820, which can monitor and / or control operations of the access network nodes 1810. This arrangement can be used, for example, when access network 1804 utilizes an O-RAN architecture. SMO system 1820 can be configured to communicate with core network 1806 and / or host 1816, as shown in Figure 18. In some embodiments, one or more of core network node 1808, host 1816, NM node 1818, and SMO system 1820 can be configured to perform various operations of exemplary methods (e.g., procedures) for generating synthetic data representative of characteristics of a communication network, such as described above in relation to other figures.

[0241] As a whole, communication system 1800 of Figure 18 enables connectivity between the UEs, network nodes, and hosts. In that sense, the communication system may be configured to operate according to predefined rules or procedures, such as specific standards that include, but are not limited to: Global System for Mobile Communications (GSM); Universal Mobile Telecommunications System (UMTS); Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, 5G standards, or any applicable future generation standard (e.g., 6G); wireless local area network (WLAN) standards, such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards (WiFi); and / or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave, Near Field Communication (NFC) ZigBee, LiFi, and / or any low-power wide-area network (LPWAN) standards such as LoRa and Sigfox.

[0242] In some examples, telecommunication network 1802 is a cellular network that implements 3GPP standardized features. Accordingly, telecommunication network 1802 may support network slicing to provide different logical networks to different devices that are connected to telecommunication network 1802. For example, telecommunication network 1802 may provide Ultra Reliable Low Latency Communication (URLLC) services to some UEs, while providing Enhanced Mobile Broadband (eMBB) services to other UEs, and / or Massive Machine Type Communication (mMTC) / Massive loT services to yet further UEs.

[0243] In some examples, UEs 1812 are configured to transmit and / or receive information without direct human interaction. For instance, a UE may be designed to transmit information to access network 1804 on a predetermined schedule, when triggered by an internal or external event, or in response to requests from access network 1804. Additionally, a UE may be configured for operating in single- or multi-RAT or multi-standard mode. For example, a UE may operate with any one or combination of Wi-Fi, NR (New Radio) and LTE, i.e. being configured for multi-radio dual connectivity (MR-DC), such as E-UTRAN (Evolved-UMTS Terrestrial Radio Access Network) New Radio - Dual Connectivity (EN-DC).

[0244] In the example, hub 1814 communicates with access network 1804 to facilitate indirect communication between one or more UEs (e.g., 1812c and / or 1812d) and network nodes (e.g., network node 1810b). In some examples, hub 1814 may be a controller, router, content source and analytics, or any of the other communication devices described herein regarding UEs. For example, hub 1814 may be a broadband router enabling access to core network 1806 for the UEs. As another example, hub 1814 may be a controller that sends commands or instructions to one or more actuators in the UEs. Commands or instructions may be received from the UEs, network nodes 1810, or by executable code, script, process, or other instructions in hub 1814. As another example, hub 1814 may be a data collector that acts as temporary storage for UE data and, in some embodiments, may perform analysis or other processing of the data. As another example, hub 1814 may be a content source. For example, for a UE that is a VR headset, display, loudspeaker or other media delivery device, hub 1814 may retrieve VR assets, video, audio, or other media or data related to sensory information via a network node, which hub 1814 then provides to the UE either directly, after performing local processing, and / or after adding additional local content. In still another example, hub 1814 acts as a proxy server or orchestrator for the UEs, in particular if one or more of the UEs are low energy loT devices.

[0245] Figure 19 shows a network node 1900 in accordance with some embodiments. Examples of network nodes include, but are not limited to, access points (APs) (e.g., radio access points), base stations (e.g., radio base stations, Node Bs, eNBs, gNBs), and O-RAN nodes or components of an O-RAN node (e.g., O-RU, O-DU, O-CU).

[0246] Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and so, depending on the provided amount of coverage, may be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be a relay node or a relay donor node controlling a relay. A network node may also include one or more (or all) parts of a distributed radio base station such as centralized digital units, distributed units (e.g., in an O-RAN access node) and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS).

[0247] Other examples of network nodes include multiple transmission point (multi-TRP) 5G access nodes, multi-standard radio (MSR) equipment such as MSR BSs, network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi-cell / multicast coordination entities (MCEs), Operation and Maintenance (O&M) nodes, Operations Support System (OSS) nodes, Self-Organizing Network (SON) nodes, positioning nodes (e.g., Evolved Serving Mobile Location Centers (E-SMLCs)), and / or Minimization of Drive Tests (MDTs).

[0248] In some embodiments, network node 1900 can be configured to perform various operations of exemplary methods e.g., procedures) for generating synthetic data representative of characteristics of a communication network, such as described above in relation to other figures. Network node 1900 includes processing circuitry 1902, memory 1904, communication interface 1906, and power source 1908. Network node 1900 may be composed of multiple physically separate components (e.g., a NodeB component and a RNC component, or a BTS component and a BSC component, etc.), which may each have their own respective components. In certain scenarios in which network node 1900 comprises multiple separate components (e.g., BTS and BSC components), one or more of the separate components may be shared among several network nodes. For example, a single RNC may control multiple NodeBs. In such a scenario, each unique NodeB and RNC pair, may in some instances be considered a single separate network node. In some embodiments, network node 1900 may be configured to support multiple radio access technologies (RATs). In such embodiments, some components may be duplicated (e.g., separate memory 1904 for different RATs) and some components may be reused (e.g., a same antenna 1910 may be shared by different RATs). Network node 1900 may also include multiple sets of the various illustrated components for different wireless technologies integrated into network node 1900, for example GSM, WCDMA, LTE, NR, WiFi, Zigbee, Z-wave, LoRaWAN, Radio Frequency Identification (RFID) or Bluetooth wireless technologies. These wireless technologies may be integrated into the same or different chip or set of chips and other components within network node 1900.

[0249] Processing circuitry 1902 may comprise a combination of one or more of a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application-specific integrated circuit, field programmable gate array, or any other suitable computing device, resource, or combination of hardware, software and / or encoded logic operable to provide, either alone or in conjunction with other network node 1900 components, such as memory 1904, to provide network node 1900 functionality.

[0250] In some embodiments, processing circuitry 1902 includes a system on a chip (SOC). In some embodiments, processing circuitry 1902 includes radio frequency (RF) transceiver circuitry 1912 and / or baseband processing circuitry 1914. In some embodiments, RF transceiver circuitry 1912 and baseband processing circuitry 1914 may be on separate chips (or sets of chips), boards, or units, such as radio units and digital units. In alternative embodiments, part or all of RF transceiver circuitry 1912 and / or baseband processing circuitry 1914 may be on the same chip or set of chips, boards, or units.

[0251] Memory 1904 may comprise any form of volatile or non-volatile computer-readable memory including, without limitation, persistent storage, solid-state memory, remotely mounted memory, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), mass storage media (for example, a hard disk), removable storage media (for example, a flash drive, a Compact Disk (CD) or a Digital Video Disk (DVD)), and / or any other volatile or non-volatile, non-transitory device-readable and / or computer-executable memory devices that store information, data, and / or instructions that may be used by processing circuitry 1902. Memory 1904 may store any suitable instructions, data, or information, including a computer program, software, an application including one or more of logic, rules, code, tables, and / or other instructions (collected denoted computer program 1904a, which may be in the form of a computer program product) capable of being executed by processing circuitry 1902 and utilized by network node 1900. Memory 1904 may be used to store any calculations made by processing circuitry 1902 and / or any data received via communication interface 1906. In some embodiments, processing circuitry 1902 and memory 1904 is integrated.

[0252] Communication interface 1906 is used in wired or wireless communication of signaling and / or data between a network node, access network, and / or UE. As illustrated, communication interface 1906 comprises port(s) / terminal(s) 1916 to send and receive data, for example to and from a network over a wired connection. Communication interface 1906 also includes radio frontend circuitry 1918 that may be coupled to, or in certain embodiments a part of, antenna 1910. Radio front-end circuitry 1918 comprises filters 1920 and amplifiers 1922. Radio front-end circuitry 1918 may be connected to an antenna 1910 and processing circuitry 1902. The radio front-end circuitry may be configured to condition signals communicated between antenna 1910 and processing circuitry 1902. Radio front-end circuitry 1918 may receive digital data that is to be sent out to other network nodes or UEs via a wireless connection. Radio front-end circuitry 1918 may convert the digital data into a radio signal having the appropriate channel and bandwidth parameters using a combination of filters 1920 and / or amplifiers 1922. The radio signal may then be transmitted via antenna 1910. Similarly, when receiving data, antenna 1910 may collect radio signals which are then converted into digital data by radio front-end circuitry 1918. The digital data may be passed to processing circuitry 1902. In other embodiments, the communication interface may comprise different components and / or different combinations of components.

[0253] In certain alternative embodiments, network node 1900 does not include separate radio front-end circuitry 1918, instead, processing circuitry 1902 includes radio front-end circuitry and is connected to antenna 1910. Similarly, in some embodiments, all or some of RF transceiver circuitry 1912 is part of communication interface 1906. In still other embodiments, communication interface 1906 includes one or more ports or terminals 1916, radio front-end circuitry 1918, and RF transceiver circuitry 1912, as part of a radio unit (not shown), and communication interface 1906 communicates with baseband processing circuitry 1914, which is part of a digital unit (not shown).

[0254] Antenna 1910 may include one or more antennas, or antenna arrays, configured to send and / or receive wireless signals. Antenna 1910 may be coupled to radio front-end circuitry 1918 and may be any type of antenna capable of transmitting and receiving data and / or signals wirelessly. In certain embodiments, antenna 1910 is separate from network node 1900 and connectable to network node 1900 through an interface or port.

[0255] Antenna 1910, communication interface 1906, and / or processing circuitry 1902 may be configured to perform any receiving operations and / or certain obtaining operations described herein as being performed by the network node. Any information, data and / or signals may be received from a UE, another network node and / or any other network equipment. Similarly, antenna 1910, communication interface 1906, and / or processing circuitry 1902 may be configured to perform any transmitting operations described herein as being performed by the network node. Any information, data and / or signals may be transmitted to a UE, another network node and / or any other network equipment.

[0256] Power source 1908 provides power to the various components of network node 1900 in a form suitable for the respective components (e.g., at a voltage and current level needed for each respective component). Power source 1908 may further comprise, or be coupled to, power management circuitry to supply the components of network node 1900 with power for performing the functionality described herein. For example, network node 1900 may be connectable to an external power source (e.g., the power grid, an electricity outlet) via an input circuitry or interface such as an electrical cable, whereby the external power source supplies power to power circuitry of power source 1908. As a further example, power source 1908 may comprise a source of power in the form of a battery or battery pack which is connected to, or integrated in, power circuitry. The battery may provide backup power should the external power source fail.

[0257] Embodiments of network node 1900 may include additional components beyond those shown in Figure 19 for providing certain aspects of the network node’s functionality, including any of the functionality described herein and / or any functionality necessary to support the subject matter described herein. For example, network node 1900 may include user interface equipment to allow input of information into network node 1900 and to allow output of information from network node 1900. This may allow a user to perform diagnostic, maintenance, repair, and other administrative functions for network node 1900.

[0258] Figure 20 is a block diagram of a host 2000, which may be an embodiment of host 1816 of Figure 18, in accordance with various aspects described herein. Hoes 2000 may be or comprise various combinations hardware and / or software, including a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, container, or processing resources in a server farm.

[0259] Host 2000 includes processing circuitry 2002 that is operatively coupled via a bus 2004 to an input / output interface 2006, a network interface 2008, a power source 2010, and a memory 2012. Other components may be included in other embodiments. Features of these components may be substantially similar to those described with respect to the devices of previous figures, such as Figure 18, such that the descriptions thereof are generally applicable to the corresponding components of host 2000.

[0260] Memory 2012 may include one or more computer programs including one or more host application programs 2014 and data 2016, which may include user data, e.g., data generated by a UE for host 2000 or data generated by host 2000 for a UE. Embodiments of host 2000 may utilize only a subset or all of the components shown. Host application programs 2014 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG, VP9) and audio codecs (e.g., FLAC, Advanced Audio Coding (AAC), MPEG, G.711), including transcoding for multiple different classes, types, or implementations of UEs (e.g., handsets, desktop computers, wearable display systems, heads-up display systems). Host application programs 2014 may also provide for user authentication and licensing checks and may periodically report health, routes, and content availability to a central node, such as a device in or on the edge of a core network. Accordingly, host 2000 may select and / or indicate a different host for over-the-top services for a UE. Host application programs 2014 may support various protocols, such as the HTTP Live Streaming (HLS) protocol, Real-Time Messaging Protocol (RTMP), Real- Time Streaming Protocol (RTSP), Dynamic Adaptive Streaming over HTTP (MPEG- DASH), etc.

[0261] In some embodiments, host 2000 can be configured to perform various operations of exemplary methods e.g., procedures) for generating synthetic data representative of characteristics of a communication network, such as described above in relation to other figures.

[0262] Figure 21 is a block diagram illustrating a virtualization environment 2100 in which functions implemented by some embodiments may be virtualized. In the present context, virtualizing means creating virtual versions of apparatuses or devices which may include virtualizing hardware platforms, storage devices and networking resources. As used herein, virtualization can be applied to any device described herein, or components thereof, and relates to an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 2100 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node may be entirely virtualized. In some embodiments, the virtualization environment 2100 includes components defined by the O-RAN Alliance, such as an O-Cloud environment orchestrated by a Service Management and Orchestration Framework via an 0-2 interface.

[0263] Applications 2102 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in the virtualization environment 2100 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein. In some embodiments, one or more applications 2102 can be configured to perform various operations of exemplary methods (e.g., procedures) for generating synthetic data representative of characteristics of a communication network, such as described above in relation to other figures.

[0264] Hardware 2104 includes processing circuitry, memory that stores software and / or instructions (collected denoted computer program 2104a, which may be in the form of a computer program product) executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 2106 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 2108a-b (one or more of which may be generally referred to as VMs 2108), and / or perform any of the functions, features and / or benefits described in relation with some embodiments described herein. Virtualization layer 2106 may present a virtual operating platform that appears like networking hardware to VMs 2108.

[0265] VMs 2108 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer 2106. Different embodiments of the instance of a virtual appliance 2102 may be implemented on one or more of VMs 2108, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment.

[0266] In the context of NFV, each VM 2108 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, non-virtualized machine. Each VM 2108, and that part of hardware 2104 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs 2108 on top of the hardware 2104 and corresponds to application 2102. Hardware 2104 may be implemented in a standalone network node with generic or specific components. Hardware 2104 may implement some functions via virtualization. Alternatively, hardware 2104 may be part of a larger cluster of hardware (e.g. such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration function 2110, which, among others, oversees lifecycle management of applications 2102. In some embodiments, hardware 2104 is coupled to one or more radio units that each include one or more transmitters and one or more receivers that may be coupled to one or more antennas. Radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with the virtual components to provide a virtual node with radio capabilities, such as a radio access node or a base station. In some embodiments, some signaling can be provided with the use of a control system 2112 which may alternatively be used for communication between hardware nodes and radio units.

[0267] The foregoing merely illustrates the principles of the disclosure. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous systems, arrangements, and procedures that, although not explicitly shown or described herein, embody the principles of the disclosure and can be thus within the spirit and scope of the disclosure. Various embodiments can be used together with one another, as well as interchangeably therewith, as should be understood by those having ordinary skill in the art.

[0268] The term unit, as used herein, can have conventional meaning in the field of electronics, electrical devices and / or electronic devices and can include, for example, electrical and / or electronic circuitry, devices, modules, processors, memories, logic solid state and / or discrete devices, computer programs or instructions for carrying out respective tasks, procedures, computations, outputs, and / or displaying functions, and so on, as such as those that are described herein.

[0269] Any appropriate steps, methods, features, functions, or benefits disclosed herein may be performed through one or more functional units or modules of one or more virtual apparatuses. Each virtual apparatus may comprise a number of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessor or microcontrollers, as well as other digital hardware, which may include Digital Signal Processor (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as Read Only Memory (ROM), Random Access Memory (RAM), cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory includes program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the respective functional unit to perform corresponding functions according one or more embodiments of the present disclosure.

[0270] As described herein, device and / or apparatus can be represented by a semiconductor chip, a chipset, or a (hardware) module comprising such chip or chipset; this, however, does not exclude the possibility that a functionality of a device or apparatus, instead of being hardware implemented, be implemented as a software module such as a computer program or a computer program product comprising executable software code portions for execution or being run on a processor. Furthermore, functionality of a device or apparatus can be implemented by any combination of hardware and software. A device or apparatus can also be regarded as an assembly of multiple devices and / or apparatuses, whether functionally in cooperation with or independently of each other. Moreover, devices and apparatuses can be implemented in a distributed fashion throughout a system, so long as the functionality of the device or apparatus is preserved. Such and similar principles are considered as known to a skilled person.

[0271] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning that is consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0272] In addition, certain terms used in the present disclosure, including the specification and drawings, can be used synonymously in certain instances (e.g., “data” and “information”). It should be understood, that although these terms (and / or other terms that can be synonymous to one another) can be used synonymously herein, there can be instances when such words can be intended to not be used synonymously.

Claims

CLAIMS1. A computer-implemented method for generating synthetic data representative of characteristics of a communication network, the method comprising: determining (1710) a plurality of temporal buckets corresponding to time instances that are equally spaced over a first duration at a frequency associated with a seasonality characteristic of data collected from the communication network; obtaining (1720) the following temporal statistics for the synthetic data: for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and second parameters defining one or more second distributions of outliers from the first distributions associated with the plurality of temporal buckets; based on the temporal statistics, generating (1730) a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic; and adjusting (1740) a subset of the plurality of synthetic data samples to represent anomalies present in data collected from the communication network.

2. The method of claim 1, wherein: each time instance associated with a synthetic data sample corresponds to one of the temporal buckets, and each synthetic data sample is generated based on the first parameters of the corresponding temporal bucket.

3. The method of claim 2, wherein the first parameters defining the first distribution for each temporal bucket include the following: a distribution type of the first distribution; a minimum of the first distribution; a maximum of the first distribution; probability of a left outlier that is less than the minimum; and probability of a right outlier that is greater than the maximum.

4. The method of claim 3, wherein generating (1730) the plurality of synthetic data samples based on the temporal statistics comprises, for each synthetic data sample:generating (1731), as the synthetic data sample, a random value according to the first distribution for the temporal bucket corresponding to the time instance associated with the synthetic data sample; selectively modifying (1732) the generated synthetic data sample to be a left outlier less than the minimum of the first distribution for the corresponding temporal bucket, based on the probability of a left outlier for the corresponding temporal bucket; and selectively modifying (1733) the generated synthetic data sample to be a right outlier greater than the maximum of the first distribution for the corresponding temporal bucket, based on the probability of a right outlier for the corresponding temporal bucket.

5. The method of claim 4, wherein: selectively modifying (1732) the generated synthetic data sample to be a left outlier is further based on a probably of a left outlier occurring in any of the temporal buckets; and selectively modifying (1733) the generated synthetic data sample to be a right outlier is further based on a probably of a right outlier occurring in any of the temporal buckets.

6. The method of any of claims 2-5, wherein adjusting (1740) the subset of the plurality of synthetic data samples to represent anomalies present in data produced by the communication network is based on an anomaly configuration that includes one or more of the following: a range of lengths for sequences of anomalous synthetic data samples; a first number of sequences of left anomalous synthetic data samples to be included; a second number of sequences of right anomalous synthetic data samples to be included; a minimum value for anomalous synthetic data samples, lmin; a maximum value for anomalous synthetic data samples, rmax; a first range of values usable to produce left anomalous synthetic data samples; and a second range of values usable to produce right anomalous synthetic data samples.

7. The method of claim 6, wherein adjusting (1740) the subset of the plurality of synthetic data samples comprises:selecting (1741) the first and second number of sequences of the synthetic data samples, wherein: each synthetic data sample of the selected first number of sequences is at least a first amount greater than lmin, and each synthetic data sample of the selected second number of sequences is at least a second amount less thandecreasing (1742) respective synthetic data samples of the selected first number of sequences, by respective first amounts determined based on the first range of values; and increasing (1743) respective synthetic data samples of the selected second number of sequences, by respective second amounts determined based on the second range of values.

8. The method of claim 7, wherein each of the first and second ranges is a range of scaling factors, and one or more of the following applies: each first amount is determined based on the following: a scaling factor randomly selected from the first range, a standard deviation for the temporal bucket to which the synthetic data sample is assigned, and; and each second amount is determined based on the following: a scaling factor randomly selected from the second range, a standard deviation for the temporal bucket to which the synthetic data sample is assigned, and rmax.

9. The method of any of claims 3-8, wherein obtaining (1720) the temporal statistics comprises: obtaining (1721) a source dataset produced by the communication network, wherein the source dataset includes a plurality of data samples associated with time instances that are equally spaced at the frequency over a reference duration; assigning (1722) each data sample to one of the temporal buckets based on a correspondence between the time instance associated with the data sample and the time instance associated with the assigned temporal bucket; and determining (1723) the first parameters for each temporal bucket based on the data samples assigned to the temporal bucket.

10. The method of claim 9, wherein determining (1723) the first parameters for each temporal bucket comprises:determining the distribution type, the maximum, and the minimum of the first distribution based on the data samples assigned to the temporal bucket; determining the probability of a left outlier based on how many of the assigned data samples are less than the determined minimum; and determining the probability of a right outlier based on how many of the assigned data samples are greater than the determined minimum.

11. The method of claim 10, wherein the second parameters defining the one or more second distributions of outliers include the following: a distribution of left outliers within any of the temporal buckets; a distribution of right outliers within any of the temporal buckets; and an outlier distribution type for the distributions of left and right outliers.

12. The method of claim 11, wherein obtaining (1720) the temporal statistics further comprises determining (1724) the second parameters defining the one or more second distributions, based on the following: for each of the temporal buckets, respective first differences between the minimum of the first distribution and the assigned data samples that are less than the minimum; and for each of the temporal buckets, respective second differences between the maximum of the first distribution and the assigned data samples that are greater than the maximum.

13. The method of claim 12, wherein determining (1724) the second parameters defining the one or more second distributions comprises: determining the distribution of left outliers based on the first differences and the outlier distribution type; and determining the distribution of right outliers based on the second differences and the outlier distribution type.

14. The method of any of claims 1-8, wherein obtaining (1720) the temporal statistics comprises: obtaining (1725) a configuration that specifies the first parameters for a subset of the temporal buckets, wherein the subset corresponds to time instances distributed unequally over the first duration; andperforming (1726) interpolation of the first parameters for the subset to obtain the first parameters for the plurality of temporal buckets.

15. The method of claim 14, wherein performing (1726) interpolation comprises: performing linear imputation of the specified first parameters for the subset to obtain first parameters for the temporal buckets not included in the subset; and performing curve fitting on the specified first parameters and the linearly imputed first parameters to obtain the first parameters for the plurality of temporal buckets.

16. A computing system (400, 1510, 1610, 1808, 1816, 1818, 1820, 1900, 2000, 2100) configured to generate synthetic data representative of characteristics of a communication network (198, 199, 200, 300, 1802), the computing system comprising: communication interface circuitry (1906, 2008, 2104) configured to communicate with one or more network nodes or functions of the communication network; and processing circuitry (1902, 2002, 2104) that is operably coupled to the communication interface circuitry, whereby the processing circuitry and the communication interface circuitry are configured to: determine a plurality of temporal buckets corresponding to time instances that are equally spaced over a first duration at a frequency associated with a seasonality characteristic of data collected from the communication network; obtain the following temporal statistics for the synthetic data: for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and second parameters defining one or more second distributions of outliers from the first distributions associated with the plurality of temporal buckets; based on the temporal statistics, generate a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic; and adjust a subset of the plurality of synthetic data samples to represent anomalies present in data collected from the communication network.

17. The computing system of claim 16, wherein the processing circuitry and the communication interface circuitry are further configured to perform operations corresponding to any of claims 2-15.

18. A computing system (400, 1510, 1610, 1808, 1816, 1818, 1820, 1900, 2000, 2100) configured to generate synthetic data representative of characteristics of a communication network (198, 199, 200, 300, 1802), the computing system being further configured to: determine a plurality of temporal buckets corresponding to time instances that are equally spaced over a first duration at a frequency associated with a seasonality characteristic of data collected from the communication network; obtain the following temporal statistics for the synthetic data: for each temporal bucket, first parameters defining a first distribution of synthetic data associated with the temporal bucket; and second parameters defining one or more second distributions of outliers from the first distributions associated with the plurality of temporal buckets; based on the temporal statistics, generate a plurality of synthetic data samples associated with time instances that are equally spaced over a second duration at the frequency associated with the seasonality characteristic; and adjust a subset of the plurality of synthetic data samples to represent anomalies present in data collected from the communication network.

19. The computing system of claim 18, being further configured to perform operations corresponding to any of the methods of claims 2-15.

20. A non-transitory, computer-readable medium (1904, 2002, 2104) storing computerexecutable instructions that, when executed by processing circuitry (1902, 2002, 2104) of a computing system (400, 1510, 1610, 1808, 1816, 1818, 1820, 1900, 2000, 2100) configured to generate synthetic data representative of characteristics of a communication network (198, 199, 200, 300, 1802), configure the computing system to perform operations corresponding to any of the methods of claims 1-15.

21. A computer program product (1904a, 2014, 2104a) comprising computer-executable instructions that, when executed by processing circuitry (1902, 2002, 2104) of a computing system (400, 1510, 1610, 1808, 1816, 1818, 1820, 1900, 2000, 2100) configured to generate synthetic data representative of characteristics of a communication network (198, 199, 200, 300,1802), configure the computing system to perform operations corresponding to any of the methods of claims 1-15.

Citation Information

Patent Citations

  • Automatic time series exploration for business intelligence analytics

    US20160342910A1

  • Systems and methods for detecting anomalous behaviors based on temporal profile

    US20210185068A1