Abnormal behavior detection in mobile networks using GPT language model

By employing a GPT Language Model to transform network data into tokens for pattern detection, the system addresses the limitations of current network analytics in detecting irregularities, enabling proactive measures and improving anomaly detection efficacy.

WO2025125876A1PCT designated stage expired Publication Date: 2025-06-19TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Patent Information

Application Number
PCT/IB2023/062717
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current network analytics techniques struggle to detect irregularities in mobile networks in advance, often generating alarms only after issues have occurred, and are challenged by the complexity of network data and the large number of possible parameter combinations.

Method used

The use of a Generative Pre-trained Transformer (GPT) Language Model to transform network events and performance metrics into tokens, enabling pattern detection and allowing for the identification of irregularities in network behavior before they become major issues.

Benefits of technology

This approach enables real-time detection of irregular patterns, allowing network operators to take preventive actions, and provides a more effective and explainable method for anomaly detection compared to traditional machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2023062717_19062025_PF_FP_ABST
    Figure IB2023062717_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A system and method are described for detecting irregular network behavior, service quality or traffic patterns, subscriber behavior or deviation from typical operation of a 3GPP mobile data network. Convolutional techniques are applied to transform time series data of network events, network performance metrics, and traffic metrics into token sequences to enable pattern detection using a generative pre-trained transformer (GPT) Language Model (LM).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ABNORMAL BEHAVIOR DETECTION IN MOBILE NETWORKS USING GPT LANGUAGE MODEL

[0002] TECHNICAL FIELD

[0003] The present disclosure relates generally to management of communication networks and, more particularly, related to the detection of irregularities in communication networks using natural language processing techniques.

[0004] BACKGROUND

[0005] Wireless telecommunication networks are configured to run with best performance by using correct configuration management parameters, which may be set and / or tuned to address varied cell sizes, deployment topographies, traffic load patterns, etc. Different configurations induce different behavior in the network, ideally to optimize network throughput, minimize interference and dropped / interrupted calls or data sessions, and otherwise optimize user experience. Normally, the parameters are configured based on guidelines provided by the respective network vendors for running the networks effectively.

[0006] Modern telecommunication networks - especially mobile networks - are extremely complex and this complexity is continuously increasing. Performance perceived by end users can be impacted by configuration, load and interactions of thousands of different networks elements. To keep operations cost low and end-user perceived performance high, it is critical to automate network operations and network optimization as much as possible.

[0007] Network Analytics systems, which are part of the Network Management domain in communications systems such as the 4G and 5G wireless communication systems specified by members of the 3rd-Generation Partnership Project (3GPP), monitor and analyze service and network quality at session level in mobile networks. Network Analytics systems are increasingly used for automatic network operation as well, improving the network, or eliminating service or network issues.

[0008] Basic network Key Performance Indicators (KPIs) are continuously monitored by Network Analytics systems. The KPIs are based on node and network events and counters. KPIs are aggregated in time and often for node or other dimensions (e.g., device type, service provider, or network dimensions, like cell, network node, etc.). KPIs can indicate node or network failures but usually are not detailed enough for troubleshooting purposes, and they are typically not suitable for identifying end-to- end (e2e), user-perceived service quality issues. In addition to calculating and monitoring KPIs, Network Analytics systems also generate “incidents.” Incidents are either based on explicit network failure indications (e.g., failure cause codes in signaling events) or based on KPI thresholds. Anomaly detection algorithm and methods are used to identify unexpected changes of KPI values, or outlying values for different node, network function, subscriber, terminal groups. Incidents are used for generating alarms in Fault Management (FM) systems.

[0009] Advanced, event-based analytics systems collect and correlate elementary network events as well as e2e service quality metrics and computing user level e2e service quality KPIs and incidents. Per-subscriber and per-session and network-level KPIs and incidents are aggregated for different network and time dimensions. These types of solutions are suitable for session-based troubleshooting and analysis of network issues. The monitored network domains include both the radio and core networks.

[0010] Problems remain with current network analytics techniques, however. For example, explicit network failure indications, node alarms, usually are generated when the issue has already occurred, which may mean that issues result in temporary or permanent service quality degradation. There are often no indications in advance before the issue occurs.

[0011] Threshold-based anomaly detection systems can simultaneously monitor multiple low-level indicators in the network and indicate defects when one of these metrics is outside of its valid range. KPI threshold-based fault indications or anomaly detection methods are typically based on time aggregated KPIs values. If the issue is characterized by a relatively short time period within the aggregation period, the issue may escape detection. For the same reason, if the KPI is aggregated for a parameter, such as node, network function, subscriber group, terminal types, etc., and the issue is related to few parameters, the issue may not be detected. Monitoring the KPIs separately for different parameter values is challenging, or not possible, due to the large number of parameters and possible parameter values or value ranges. If an issue is related to two or more specific parameter combinations, the issue may become very difficult to detect with legacy methods due to the large number of possible combinations.

[0012] In the state of the art, machine learning (ML) methods are often used for anomaly detection. In this traditional ML approach, the input of the ML model is the data itself with some transformation, like feature selection etc. ML methods typically learn the normal behavior by using data with normal behavior at the input in the learning phase, and then at the operational phase they fire an alarm if the new incoming data is different from what the method learnt in the learning phase.

[0013] In the case of network data, supervised methods are difficult to apply for this task, because a target function is needed for them to recognize what is the anomaly. A lot of manual work is required to identify and label anomalies. Enumeration of all possible combinations is nearly impossible because of the large feature space.

[0014] Unsupervised methods are more powerful to find outliers using N-dimensional Euclidean distance, clustering, nearest neighbors, or statistical techniques. Data session records have a large number of feature spaces with sparse filled content so many records will have unused fields that are empty. Further, there are many variations in the fields used in the same data records, and there may also be many overlapping parts among the observable structures. This characteristic makes the data session record comparison by Euclidean distance or clustering to reveal outliers difficult.

[0015] Natural language model applications can be used to handle the large variability of patterns in the data of an observed network. The co-pending PCT Application No. PCT / IB2022 / 056639 titled "Early detection of irregular patterns in mobile networks", (hereinafter “Early Detection Of Irregular Patterns”), discloses the use of a language model (LM) to detect anomalies in time series data collected from the network. The LM is based on N-grams. In this approach, certain token collocations occurring at a high rate in good quality training data are assumed to represent the expected language of the evaluated data and any unusual token transitions are possible anomalies. Because of the dimensionality of time series network data, random sampling is used to select sequences in both the training and evaluation phases. The evaluation can be slow, because thousands of random sequences are needed to enumerate from one data row to provide enough number of sequential token pairs to find unlikely transitions. For example, if one data row is the daily KPI values of one subscriber, for the evaluation the data of millions of subscribers there need to evaluate billions of token sequences.

[0016] Another drawback is that the N-gram approach focuses on the relation between the following tokens. Far time distance relations cannot be modelled with N- gram because the coexistence probability of non-following tokens is faded in the computed N-gram probability chain of the token sequence. It is the same for more complex relations between more KPIs at far distances or casual effects. The tokenization does not differentiate the magnitude of KPI values. If one KPI value is higher than the other, they are encoded into two different tokens with the same importance. However, some tokens may be more important than others. For example, if the KPI is related to packet loss, the higher value is most likely indicating issues than the lower one. Token importance cannot be handled in direct way with the N-gram approach.

[0017] Handling changing correlation effects with N-gram approach is also challenging. The correlation between some KPI pairs can be different at morning, midday and evening depending on the habit of network users. Encoding the hour would solve the correlation problem, but it may cause other problems, such as a much larger vocabulary and independent tokens (with missing relation) for the same KPI at different hours.

[0018] Traditional ML models based on the KPIs as input data also have several drawbacks. Most importantly, they are typically non-explainable, since the black-box model that is learned can classify the new input as abnormal correctly, but there is no hint on why it was abnormal. In addition, the computational footprint of these methods is rather large so that re-training and maintenance of the models is difficult and time-consuming. Finally, the capabilities of these models are limited by the algorithms used, and these typically cannot compete with the language-based methods.

[0019] SUMMARY

[0020] A system and method are described for detecting irregular network behavior, service quality or traffic patterns, subscriber behavior or deviation from typical operation of a 3GPP mobile data network. Convolutional techniques are applied to transform network events, network performance metrics, and traffic metrics into tokens to enable pattern detection using a generative pre-trained transformer (GPT) Language Model (LM).

[0021] The detection of irregular patterns is performed to enable the network operators to take preventive actions before major network issues occur affecting large number of subscribers. The irregular pattern detection is done in two phases. First, in the training phase, the normal behavior according to Quality of Service (QoS) is learned offline. Second, the irregularities are continuously detected and reported by the system in real time during operation. In the evaluation phase, the LM processes collected network data and estimates the probability that the collected data was observed during normal operational period. The network operator can differentiate good and bad periods using the LM’s response.

[0022] A first aspect of the disclosure comprises a method of detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module. The method 400 comprises transforming a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The method further comprises transforming the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds to a set of KPIs for a time slot in the time interval. The method further comprises detecting irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

[0023] A second aspect of the disclosure comprises an analytics system for detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module. The analytics system is configured to transform a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The analytics system is further configured to transform token sequences, where each token sequence corresponds to a set of KPIs for a time slot in the time interval. The analytics system is configured to detect irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

[0024] A third aspect of the disclosure comprises an analytics system for detecting irregularities in the performance of a wireless communication network using a GPT- based LM as may be performed by the GPT anomaly detection module. The analytics system comprises processing circuitry and memory cooperatively coupled to the processing circuitry. The memory stores program instructions that when executed by the processing circuitry cause the analytics system to transform a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The program instructions, when executed by the processing circuitry, further cause the analytics system to transform token sequences, where each token sequence corresponds to a set of KPIs for a time slot in the time interval. The program instructions, when executed by the processing circuitry, further cause the analytics system to detect irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

[0025] A fourth aspect of the disclosure comprises a computer program for an analytics system in a wireless communication network. The computer program comprises executable instructions that, when executed by processing circuitry in the analytics system, causes the analytics system to perform the method according to the first aspect.

[0026] A fifth aspect of the disclosure comprises a carrier containing a computer program according to the fourth aspect. The carrier is one of an electronic signal, optical signal, radio signal, or a non-transitory computer readable storage medium.

[0027] A sixth aspect of the disclosure comprises a method of detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module. The method 400 comprises transforming a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The method further comprises transforming the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds to a set of KPIs for a time slot in the time interval. The method further comprises training the token language model using the KPI token sequences generated from the intensity maps..

[0028] A seventh aspect of the disclosure comprises an analytics system for detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module. The analytics system is configured to transform a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The analytics system is further configured to transform token sequences, where each token sequence corresponds to a set of KPIs for a time slot in the time interval. The analytics system is configured to train the token language model using the KPI token sequences generated from the intensity maps.

[0029] An eighth aspect of the disclosure comprises an analytics system for detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module. The analytics system comprises processing circuitry and memory cooperatively coupled to the processing circuitry. The memory stores program instructions that when executed by the processing circuitry cause the analytics system to transform a set of KPIs for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. The program instructions, when executed by the processing circuitry, further cause the analytics system to transform token sequences, where each token sequence corresponds to a set of KPIs for a time slot in the time interval. The program instructions, when executed by the processing circuitry, further cause the analytics system to train the token language model using the KPI token sequences generated from the intensity maps.

[0030] A ninth aspect of the disclosure comprises a computer program for an analytics system in a wireless communication network. The computer program comprises executable instructions that, when executed by processing circuitry in the analytics system, causes the analytics system to perform the method according to the sixth aspect.

[0031] A tenth aspect of the disclosure comprises a carrier containing a computer program according to the ninth aspect. The carrier is one of an electronic signal, optical signal, radio signal, or a non-transitory computer readable storage medium.

[0032] BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 illustrates a wireless communication implementing anomaly detection using a GPT-based LM.

[0034] Figure 2 illustrates a GPT anomaly detection module.

[0035] Figure 3 illustrates a method of selecting training data to train the GPT-based LM.

[0036] Figure 4 illustrates a method of training the GPT-based LM.

[0037] Figure 5 illustrates an exemplary GPT architecture.

[0038] Figure 6 illustrates a method of detecting anomalies during network operation. Figure 7 illustrates a KPI intensity map during normal network operation. Figure 8 illustrates a KPI intensity map during irregular network operation. Figure 9 illustrates time correlation heatmaps for 5, 10, 15, and 20 time slots. Figure 10 illustrates an exemplary KPI correlations for unordered KPI matrix. Figure 11 illustrates an exemplary KPI correlations for an ordered KPI matrix.

[0039] Figure 12 illustrates an abnormality heat map for a full time interval.

[0040] Figure 13 illustrates an abnormality heat map for a partial time interval.

[0041] Figure 14 illustrates a method of detecting anomalies in network performance using a token language model.

[0042] Figure 15 illustrates a method of training a token language model for detecting network anomalies.

[0043] Figure 16 is a block diagram showing an example network node for carrying out one or more of the presently disclosed techniques. Figure 17 illustrates a virtualization environment, in which parts of or all of any of the techniques disclosed herein may be implemented.

[0044] DETAILED DESCRIPTION

[0045] A system and method are described for detecting irregular network behavior, service quality or traffic patterns, subscriber behavior or deviation from typical operation of a 3GPP mobile data network. Convolutional techniques are applied to transform network events, network performance metrics, and traffic metrics into tokens to enable pattern detection using a generative pre-trained transformer (GPT) Language Model (LM).

[0046] The detection of irregular patterns is performed to enable the network operators to take preventive actions before major network issues occur affecting large number of subscribers. The irregular pattern detection is done in two phases. First, in the training phase, the normal behavior according to Quality of Service (QoS) is learned offline. Second, the irregularities are continuously detected and reported by the system in real time during operation. The input KPI values are crossnormalized so that values close to 1.0 indicate high activity. Only those session records are selected where the estimated QoS level was above a predefined threshold. In the evaluation phase, the LM processes collected network data and estimates the probability that the collected data was observed during normal operational period. The network operator can differentiate good and bad periods using the LM’s response.

[0047] The proposed GPT approach builds a LM on a normal operational period with attention to both the short-range and long-range relations between KPIs and time slots. The measurement of uncertainty is extracted from the perplexity maps that can be generated by the LM for any session record. The aggregated uncertainty yields an overall metric in one value and makes different session records comparable, orderable, and selectable. For example, the overall metric enables selection of the top N session records most likely to reveal issues with network performance. The perplexity map also provides detailed information about potential network issues and indicates the network elements most affected, the KPIs impacted by the network issues, and the timing and duration of the network issue.

[0048] LM theory is an intensively researched field for natural language processing. There are thousands of publications about the variants, solutions, and their successful applications in Natural Language Processing (NLP). In earlier LM implementations, LMs are based on N-gram probabilities, where the N-grams are the N length subsequences of words in the given sentence. State of the arts, LM applications are based on Generative Pretrained Transformer (GPT) architecture and Multi-head Attention (MHA) technology.

[0049] Generative pretraining was a long-established concept in ML applications, but the transformer architecture used in state-of-the art LMs was not introduced until 2017. The development of the transformer architecture led to the emergence of large language models which are pre-trained transformers but not designed to be generative (they were "encoder-only"). OpenAI later introduced the first generative pre-trained transformer (GPT) system as "GPT-1", and the ChatGPT was built on the GPT-3.5 and GPT-4 releases.

[0050] Autoencoders are neural networks for learning the compressed representation of input data. Such model contains an encoder and decoder part as part of the transformer architecture. The encoder transforms the input feature vector to a much more compact latent vector keeping the important information in a dimensionally reduced space. The decoder tries to reconstruct the original input from the latent vector. GPT is also a transformer variant but includes a self-attention mechanism. Due to the self-attention mechanism, the GPT outperforms earlier LMs for NLP tasks.

[0051] NLP analogies can be found for the “language” of non-natural sequences as well. For example, GPT-based LMs have been applied to the task of image recognition, where the token sequences are generated with a convolutional transformer (CT). The GPT-based LM applies a MHA module to reveal long-distance relations between image parts. According to the reported results on a reference image recognition dataset, the CT outperforms more traditional methods, such as Convolutional Neural Networks (CNN).

[0052] Anomaly detection in a communication network is an unsupervised task for analyzing network data. It differs from issue classification, which is a supervised task and demands human assistance in labeling issues, which is a very hard task on network data. For anomaly identification, there need a ML training strategy, such as training on good quality network data and evaluate this model on full data and check what are the significant differences in comparison of good period.

[0053] Many similarities can be found comparing human conversation (NLP) with machine-to-machine communication. In natural languages, the word order (syntax) is determined by grammar rules, while in network traffic most messages are statistically coming in the same order because there is a higher protocol that handles the same cases in the same way. It is possible to model the network traffic with NLP technologies. For example a natural language model (NLM) can be applied to identify unusual sequences in the network flow that may indicate possible malfunctions. For example, call flows can be modelled with KPI token N-grams to detect early anomalies by real-time monitoring of LTE and Voice over LTE (VoLTE) call flows that will later cause the dialog to end with an error message. Rejecting such dialog immediately will save resources and time in the IMS network.

[0054] Mobile networks produce huge amount of time series data when network entities (e. g., subscribers, cells, or nodes) are measured by KPIs sampled at regular time intervals (e. g. at each 5 minutes or hours). According to one aspect of the disclosure, the enumeration of this data yields a sequence of values that can be modelled by ML methods. In “Early detection of irregular patterns” random sampling was used to extract KPI token sequences for building N-gram based LMs to identify anomalies in the time series data. Based on the law of large numbers, the LM assigns higher probability to bigrams, three-grams, etc. occurring most often in the random samples. In evaluation phase, the statistically unusual subsequences can be identified as anomalies. Subscriber classification-based network service level indicators (SLI) were used to select good quality network data for training model.

[0055] The present disclosure improves the existing solutions for LM-based anomaly detection on network time series data using QoS based training data selection, a GPT- based LM, and convolutional tokenization of network data activation maps.

[0056] Figure 1 illustrates an example system architecture of a wireless communication according to standards published by the Third Generation Partnership project (3GPP) network 10 and including an analytics system. 100. The wireless communication network 10 comprises a radio access network (RAN) 20, core network 30, and an Operations Support System (OSS). The analytics system 100 forms part of the OSS 40. The wireless communication network may further include an internet Protocol (IP) Multimedia Subsystem (IMS) 50.

[0057] The RAN 20 comprises one or more base stations 25 providing access network access to user equipment (UEs) in their respective service areas. The base stations, also referred to as access nodes, are referred to in Long Term Evolution networks as Evolved NodesBs (eNBs) and in New Radio (NR) networks as gNodeBs (gNBs).

[0058] The core network 30 comprises a collection of network functions (NFs) performing different network tasks. Figure 1 illustrates various NFs relevant to this disclosure including the UPF 35, Access and Mobility Management Function (AMF) 40, and Session Management Function (SMF) 45.

[0059] The UPF 35 supports handling of user plane traffic, including packet inspection, packet routing and forwarding, traffic usage reporting, and QoS handling. The UPF 35 connects with external IP networks and serves as an IP anchor point for user equipment (UEs) served by the UPF 35 so that the UEs are reachable even when moving around in the network 10. The UPF 35 processes data being forwarded. Such processing may include packet inspection, classification, and QoS marking of forwarded packets. The UPF 35 generates traffic usage reports, which the SMF 45 includes in charging reports, and is involved in policy enforcement.

[0060] The AMF 40 is a network function that manages access to the 5G network and handles mobility-related functions for the UEs. Its role is similar to that of the Mobility Management Entity (MME) in Fourth Generation (4G) networks. When a UE is not in idle mode, the gNB handovers are exposed to the AMF 40. Additionally, the AMF 40 can use other signaling elements to determine which gNB a UE 15 is attached to. The AMF 40 can expose these events by means of a standardized interface as well.

[0061] The SMF 45 manages Packet Data Unit (PDU) sessions for the UEs, which includes the establishment, modification, and release of PDU sessions. The SMF 45 selects the UPF 35 to handle a PDU session and controls the UPF 35. The SMF 45 receives Policy and Charging Control (PCC) Rules from the PCF 50 and configures the UPF 35 for various data flow tasks, such as shaping, policing to provide bandwidth, and charging functions.

[0062] The OSS 40 comprises management systems for monitoring, analyzing, and managing the other components in the wireless communication networklO. The OSS 40 includes the analytics system 100 and a fault manager 45. The analytics system 100, part of the OSS 40, collects real-time events from various network nodes and network functions from data sources in different network domains (i.e. , RAN 20, core network 30, and IMS 50). The most relevant data sources are 4G and 5G base stations 25 in the RAN 20, the UPF 35, AMF 40 and SMF 45 in the core network 30, and Call State Control Functions (CSCFs) 55 from the IMS 50. The analytics system 100 analyzes the collected data and sends alarms to the sends alarms to FM 45, or other alarm consumers.

[0063] The analytics system 100 includes an event correlator 110, aggregator 115, KPI calculation module 120, rule engine 125, QoS module 130, alarm generator 135, and GPT detection system 140. The event correlator correlates network events per session into a per session correlated records, which is streamed to the KPI calculation module and GPT detection module. The rule engine contains expert rules, explicit logics to identify network or service quality issues. The aggregator aggregates KPIs, incidents for different time periods and parameters. Based on a single incident, or aggregated incidents, alarm generator sends alarms to the Fault Manager (FM), or other alarm consumers. The GPT detection system extends the rule-based alarms in conventional analytic systems 100 with early warnings as hereinafter described. The GPT detection system uses a trained LM to evaluate the data session records in real time and detect anomalies in network operation. The QoS module aids in the selection of data session records used for training the LM. Generally, the QoS module ensures that the data session records selected for training the LM represent normal or good operating periods.

[0064] Figure 2 illustrates the main functional components of the GPT detection system 140. Figure 2 illustrates components of an example GPT detection system 140. GPT detection system 140 includes a training subsystem 150, as well as an evaluation subsystem 180. The training subsystem 150 includes a training data selection module 160, and a GPT model training module 170. The functions and operations of these modules will be described below. The evaluation subsystem 180 includes an early detection module 190 as will be detailed below to evaluate the KPI token sequences in real time using a language model trained by training data selection module 160. The evaluation of token sequences in real time facilitates early detection of anomalies, which in turn allows corrective actions to be taken sooner than would otherwise be possible.

[0065] The input data to GPT detection system 140 may be collected by an underlying existing network management system, such as Ericsson Expert Analytics. The input, called “data session records”, contains records of each individual data session of each subscriber. One data session describes one service usage transaction with all the details. As used herein, the term “service” refers to a particular service offered to subscriber stations (e.g., UEs) by the wireless communication network 10, such as video streaming, web browsing, etc.

[0066] The anomaly detection process implemented by the analytics system has three main phase referred to herein as the learning phase, operation phase, and problem identification phase. During the learning phase, the normal network operation is learned and modelled. The baseline for normal network operation is identified by existing methods. KPIs are collected during the normal operating phases for training the LM. The KPIs are organized into KPI matrices where each column represents one time slot in a time interval and each row corresponds to a network parameter that is being monitored, which is also referred to herein as a performance attribute. The KPI matrices are converted into images referred to as KPI intensity maps where each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval. Convolutional techniques are applied to transform the KPI intensity maps into token sequences that will be input to a GPT to train the LM. By applying the token sequences collected during normal operation, the GPT is able to generate a trained LM that represents normal network operation.

[0067] During the operational phase, the trained GPT is applied on new token sequences (generated in the same way as above from raw KPI values) to detect any abnormally behaving network entities. A network entity can be a subscriber, a cell, a network function, or a network node The abnormal behavior detection is done by computing a generic loss function based on perplexity.

[0068] In the problem identification phase, abnormally behaving network entities are identified and ranked based on the loss. The exact problem per network entity is identified by generating a special KPI abnormality heatmap per network entity that shows the exact time period(s) and KPI(s) that are behaving abnormally.

[0069] The system as herein described can be applied on sliding window data to detect the possible misbehavior in early phases.

[0070] The following provides further details regarding the learning and operation phases.

[0071] For both the learning and operation phases, the input record has a fixed (flat) structure and all possible fields that can be relevant in any of the possible service usage transactions. Some fields can be common for multiple service types, e.g., a service provider field is a valid field both for web browsing and video streaming. Some fields may be specific to a certain service type, e.g., a video resolution field is only relevant to video service. In the latter case, technically, the non-relevant fields of the input record may be “null”.

[0072] The input of the system is per session, per cell, per NF etc. In case the network entity for which the KPIs are collected is a subscriber, the KPIs contain records of each individual data session of each subscriber. One data session describes one service usage transaction with all the details.

[0073] Subscribers may use a variety of services provided by their devices (e.g., UEs) and the mobile network operator. These services may include, but are not limited to the following:

[0074] • Video streaming

[0075] • Video conferencing

[0076] • Video chat I video call

[0077] • Voice call

[0078] • Messaging

[0079] • E-mail • Web browsing

[0080] • File download

[0081] • File upload

[0082] • Location services

[0083] • Presence

[0084] • Software update

[0085] • Social networking

[0086] The quality of services provided by the wireless communication network 10 can be measured by several Key Performance Indicators (KPIs) and / or Quality of Experience (QoE) metrics. These performance indicators are collectively or separately referred to herein as simply KPIs. Service-level KPIs indicate the subscriber perceived service quality. Some of these KPIs are common to several service types, e.g., downlink throughput KPI, can be relevant for Web browsing, video streaming, file transfer, and several other service types. In many cases, though, the KPI is unique for the given service. For example, the KPI “video stall ratio” is relevant to and can be computed for video streaming service only.

[0087] Geneal service quality KPIs indicating service quality include, but are not limited to, the following:

[0088] • Downlink I uplink throughput of data traffic [0-100 Gbps]

[0089] • Video quality [1-5 MOS type of score]

[0090] • Video stall time ratio [0-100%]

[0091] • Video initial buffering time [0-100 s]

[0092] • Web page access time [0-100 s]

[0093] • Web page download time [0-100 s]

[0094] • Web page download success ratio [0-100%]

[0095] • Downloaded-uploaded bytes [0-1000 MB]

[0096] • Voice quality [1-5 MOS type of score]

[0097] • Call setup time [0-100 s]

[0098] • Call setup success ratio [0-100%]

[0099] • Call drop ratio [0-100%]

[0100] • Radio KPIs indicative of the quality of the radio interface include, but are not limited to, the following:

[0101] • Reference Singal Received Power (RSRP

[0102] • Reference Singal Received Quality RSRQ

[0103] • Signal to Interference plus noise ratio (SI NR)

[0104] • Handover success or failure ratio [0-100%] • Handover execution time [0-100 ms]

[0105] • Session setup success or failure ratio [0-100%]

[0106] • Session setup time [0-100 s]

[0107] • Registration success or failure ratio [0-100%]

[0108] • Session failure ratio (drop) [0-100%]

[0109] • Transport KPIs indicative of the quality of the transport network include, but are not limited to, the following:

[0110] • Bitrate [0-100 Mbps]

[0111] • Throughput [0-1000 Mbps]

[0112] • Packet loss ratio [0-100%]

[0113] • Packet delay or round-trip time [0-1000 ms]

[0114] • Jitter [0-1000 ms]

[0115] • Burst length [0-1000 ms]

[0116] • Burst size [0-100 kB]

[0117] • Burst rate [0-1000 1 / s]

[0118] In both the training and evaluation phases, the KPIs are organized into a KPI matrix per network entity. The network entity is typically a subscriber, but the techniques herein described can be applied to another network entities such as network nodes, network functions, cells, etc. The KPI matrix serves as input to the LM model during both the training and evaluation phases. Below is an example of a 4-timeslot long KPI matrix:

[0119] Header: subscriber identifier: IMSI=1234567890, start time: 1695263500, resolution: 60 sec

[0120] Time stamps: 1695263500 1695263560 1695263620 1695263680 KP matrix:

[0121] The KPIs can be aggregated for different network nodes (e.g. NFs, cells, etc.), which can be added to the per session, per node, per cell, etc. Per session KPIs can be extended with the aggregated KPIs (e.g., bitrate of the session is 500 kbps, but the average cell value is 200 kbps) and used as input to the GPT model.

[0122] In addition to aggregated KPIs node, cells can have node or cell specific KPIs. Examples of node-specific or NF-specific KPIs include, but are not limited to:

[0123] • Processing load [0-100%]

[0124] • Memory usage [0-100%]

[0125] • Slice load [0-100%]

[0126] Examples of cell-specific KPIs include, but are not limited to:

[0127] • PRB usage [0-100%]

[0128] • Radio resource sharing [0-100%]

[0129] • Slice load [0-100%]

[0130] These aggregated KPIs enable the LM to identify bad / abnormal node, NF or cell instances.

[0131] In addition to KPIs, the analytics system 100 can collect additional parameters and / or information for the data sessions. Some examples of this additional information include user equipment information, subscription groups, radio access technology (RAT), etc. KPIs can be calculated and used in separate input matrix for these instances, e.g., separate matrix for 4G radio, 5G radio, or Samsung or Apple devices, or separate LM for tablet users and for smartphone users. The goal is to have homogenous groups for which the normal behavior is learned.

[0132] Some parameters are related to specific KPIs only. For example, throughput, bitrate or service quality can be obtained for different service types, or processor load can be calculated for different NFs. Therefore, it is possible to apply separate throughput KPIs for web browsing and for video streaming, or separate video Mean Opinion Score (MOS) KPI for YouTube and Netflix video service provider, and so on.

[0133] LMs can be used as a generator of trained language. If there is an initial sequence, the most likely continuation tokens can be predicted by the LM. Using this information, a random selector can generate artificial sequences (for example Shakespeare style conversations). This feature makes LMs applicable for anomaly detection. Ttraining data collected sequences from “good” operational periods represents the expected language of the LM. In the evaluation phase, the LM assigns high probability to sequences that are similar to the observed “good” sequences, and a low probability to the unexpected “bad” sequences that contain unseen token combinations.

[0134] The previously described input data contains several KPIs that alone cannot provide overall picture about the quality of the provided service. These KPIs provide insight into the performance of subservices, such as throughput, delay, setup time, call drops, voice quality video quality, etc. An aggregated measurement is needed that automatically assigns quality information for all session records. This metric is given in co-pending PCT Application No. PCT / IB2022 / 057855 titled "Unified Service Quality Model For Mobile Networks" (hereinafter “Unified Service Quality Model”) where the QoS is calculated by a loss model that aggregates various network operational features like accessibility, retainability and integrity. The QoE calculation focuses on network degradation that occurs in the setup, service, and termination phases, which are handled separately. If no issue is detected, it is assumed that the QoE is not degraded and the maximum value is assigned. The QoE is expressed in a scale of 1-5, which is typically used for MOS KPI.

[0135] Service degradation can occur in the setup, service, and termination phases, which are handled separately. Note that the importance of these degradations strongly depends on the service types.

[0136] During training, the QoS calculation module 130 receives the input session records data and extends it with the estimated loss values. The input session record is selected into the training set when: where the limit value is a predefined system parameter. This parameter can be set experientially. If the limit is too low, the modelled network language will accept sequences from poor-quality periods as normal behavior and does not highlight any deviation. If the limit is too high, the model will be overly sensitive to deviations from excellent-quality sequences.

[0137] The QoS Calculation module 130 is only needed for training data selection. Generally, using the QoS for training data collection can be good approach for lots of ML tasks in network environment.

[0138] Figure 3 illustrates a method 200 of training data selection carried out by the analytics system 100. The KPIs computed by the KPI calculation module 120 and aggregator 115 are input to the QoS calculation module 130 and training data selection module 160 (block 210). QoS calculation module 130 computes a loss metric according to the left side of Eq. 1 above and outputs the loss metric to the training data selection module 160 (block 220). The training data selection module 160 selects data for training is the loss metric is greater than the threshold, which is assumed to represent a normal operating period (block 230). The selected data is output the GPT training module 170 as training data (block 240). Feature selection is also important because unnecessary fields may overload the input data and model generation, which can slow down the evaluation process and decrease the quality of generated models. Enrichment fields like identifiers (IDs), locations, IP addresses, and timestamps should not be selected directly for ML model building. However, these fields often provide contextual information for one or more KPI values. For example, cell frequency may affect some KPIs performance. Therefore, enrichment fields can be used to define entity groups, such as network usage subscriber profiles, cell types, etc. The KPI relations for different groups can be slightly different and occasionally significantly different. Therefore, a more accurate ML model can be generated if these groups are handled separately in the training and evaluation phases. For example preserving the enrichment context from the original session records may help the further analysis of the model evaluation results, such as aggregating the identified issues by one or more enrichment fields, or giving extra information for Root Cause Analysis (RCA).

[0139] KPIs are generated by the network analytics system 100 from the observed network entity, network events and counters during a specified time interval. The time interval for the sample matrix given above is one minute. The raw KPI values can be very different scale and need to be normalized. After the normalization the KPIs become comparable to each other. KPI values near to 0 should indicate low traffic or low error level, and KPI values near to 1 should indicate high traffic or potential issues. Some KPIs may need to be reversed so that the normalized values are comparable. For example, a KPI based on the count of packet loss should be reversed. The value update can be executed with this formula:

[0140] „ ( KPnrmnormal case

[0141] KPI xvnnoorrmm = 1 . > KP nI jnormreverse case Eq. (2)

[0142] During both the training and evaluation phases, the KPI matrices are transformed into KPI intensity maps. The KPI intensity map is a two-dimensional (2D) matrix representation of the data rows of the same entity (subscriber, node, cell) in the observed time period. The columns correspond to sample time slots (e.g., 5 mins, 1 hour) and the rows correspond to the KPI values for a given attribute as these are changing in time (e.g., during a day). This format is used to visualize the dynamics of KPI changes by replacing the normalized values with gray scale values as in Figures 7 and 8 and to build ML models for analyzing or classifying the intensity maps. The KPI values are normalized float numbers and the greyscale conversion of KPI matrix is achieved by mapping the normalized KPI values [0.0 ... 1.0] to gray scale values [0 ... 255] mapping where the R, G, B color numbers are the same. Figure 7 shows a KPI intensity map for a normal operating period and Figure 8 show a KPI intensity map for an abnormal operating period.

[0143] Time is a special dimension in this 2D representation of the KPIs. The same KPI sequence repeats in each time slot with the actual values at the given time slots. Based on the KPI intensity map, ML models can learn hidden relations between KPIs. For example, if one KPI value is increasing or decreasing what is the effect of this for other KPIs, what are the most correlating other KPIs.

[0144] The KPI intensity maps enable network data to be represented as images. The KPI intensity maps are applicable for visualization, but also can be analyzed using image processing methods. In real image processing tasks (e.g., animal recognition) convolutional methods, such as CNNs, provide efficient method to achieve high accuracy. The real (photo) images have a strong 2D local structure where spatially neighboring pixels are usually highly correlated. Use of convolutional techniques forces the capture of this local structure while also achieving some degree of shift, scale, and distortion invariance.

[0145] The KPI intensity maps can be highly correlated with the context at each pixel. Figure 9 illustrates time correlation heatmaps and show the time correlation that naturally occurs between same KPI values sampled in neighboring timeslots due to the fact that the same things are most likely to occur in the network at the next timeslot compared to the previous time slot. With rolling timeslots (e.g., 0-24 hour) the beginning and ending timeslots are also correlating because these are close to each other in time (end of day, beginning of next day), despite the fact that these time slots are far apart on the KPI intensity map.

[0146] There can be many correlating KPI pairs among the features of selected network data, but it is not obvious that these KPI rows are close to each other on the KPI intensity map. The order of KPIs may have been compiled according to other aspects or simply it was not essential in another data processing tasks. However, in convolutional approaches, the vertical correlation would be also important to provide appropriate input for image processing methods. Figures 10 and 11 are heat maps showing vertical correlation when the KPIs are not ordered (Figure 10) and when the KPI order is optimized for context processing by focusing on one KPI (Figure 11).

[0147] Reordering the KPI order in the KPI intensity map to support the context correlation of one KPI, can impact the order of some other KPIs. Therefore, a global optimum should be found for ordering KPIs that keeps the context correlation for all KPIs on an appropriate level. The following expression is used in the reordering algorithm as maximum target:

[0148] ' t ' j corr(KPlLj) / dist (KPILj)' -> max Eq. (3) where corr KPIif) and dist (KPIt ) notations are the correlation and row distance between the i-th and j-th KPIs in the KPI activation map.

[0149] During training and evaluation, the KPI intensity maps, after reordering, are used as input to a GPT.

[0150] Figure 4 illustrates a method 250 of model training as may be performed by the GPT model training module 170. The GPT model training module 170 implements a GPT to learn the “language” embodied in the training data output by the training data selection module 160. The GPT model training module 170 receives the training data selected by the training data selection module 160 as input (block 260). The GPT model training module 170 generates KPI intensity maps that are input to a GPT having a convolutional tokenizer (block 270). As will be described in more detail below, the convolutional tokenizer transforms the KPI intensity map for a time interval into discrete KPI token sequences for each time slot within the time interval. The KPI token sequences are then used to train the GPT-based LM (block 280). Following training, the GPT model training module 170 outputs the trained GPT-based LM (block 290).

[0151] More formally, given an unsupervised data S that contains {s1;...,sn} sequences of tokens from the T = {t1, vocabulary. The language modeling task is to maximize the following likelihood: where k is the size of the context window. The conditional probability P is modeled using a neural network with parameters 0, what are trained using stochastic gradient descent method. The P can be estimated with the output of last transformer block as follows: hQ= CWe+ WpEq. (5) ht= trans former _block(hi_1) Vi e [1, layers] Eq. (6)

[0152] P(s) = softmax hiWg) Eq. (7) where C = {t_k, t_x} is the context vector of tokens, Weis the token embedding matrix, and Wpis the position embedding matrix.

[0153] The attention mechanism makes use of three main components, the queries Q, the keys K of dimension dk, and the values V of dimension dv. It takes the query vector attributed to some specific token in the sequence and scores it against each key in the database. In this way it captures how the token under consideration relates to the others in the sequence. Then it scales the values according to the attention weights (computed from the scores) to retain focus on those tokens relevant to the query. The attention function is computed as the dot products of the query with all keys, divide each by yfd^, and a softmax function is applied to obtain the weights on the values:

[0154] Instead of performing a single attention function the model applies a Multiheaded Attention (MHA) operation over the input context tokens followed by position- encoded feedforward layers. The queries, keys and values are learned linear projections in n times. These projections are concatenated and once again projected, resulting in the final values: multi_head Q, K, 7) = concat(head1, ... , / ieadn) Eq. (9) are projection parameter matrices.

[0155] After training the model with the maximum likelihood target the parameters are adapted to the supervised target task. We assume a labeled dataset Y, where each instance consists of a sequence of input tokens along with a label y. The inputs are passed through our pre-trained model to obtain the final transformer block’s activation hn, which is then fed into an added linear output layer with parameters Wyto predict y.

[0156] This gives the following objective to maximize:

[0157] Figure 5 illustrates in simplified form the GPT architecture 400 used by the by the GPT detection module 140 for both training the GPT LM and using the GPT- based LM for anomaly detection. The GPT architecture has two parts: the GPT part 420 and the ML part 450. The GPT part 420 includes a convolutional layer 430 and MHA layer 440. The GPT part 420 receives the KPI intensity map form of session records as input and the convolutional layer 430 enumerates the overlapping chunks of 2D data column-wise to generates token sequences according to the following formula: where zlklvis the token input for Q / K / V multi-headed attention matrices from the z(session record that is transformed to intensity map with the Reashape2DQ function. The Conv2D0 is a depth-wise separable convolution, and s refers to the convolution kernel size. During the training process the convolution layer 430 learns the most accurete filters. The flatten function enumerates the extracted tokens columnwise. In the training phase all possible sequences are extracted from the beginning until the last token of the selected timeslot.

[0158] The ML part 450 shown in Figure 5 performs the supervised finetuning procedure as described above. The ML part 450 implements a perplexity mapper 460. In this proposed solution the main focus on the abnormal behaviour detection what is calculated from the prediction confidence of GPT model for the input session records. In the finetuning phase, there is no need to perform additional supervised ML training. The main task of the ML part 450 is to optimize the parameters and thresholds that determine the severity and amount of identified anomalies.

[0159] Once the GPT-based LM is trained, it can be used to detect irregularities or anomalies in network operation. Figure 6 illustrates a method 300 of detecting anomalies or irregularities in a network using a GPT-based LM as may be performed by the early detection module 190 in the evaluation subsystem 180 as shown in Figure 2. This method 300 uses the same GPT architecture 400 shown in Figure 5. Data collected in real time by the analytics system 100 is input to the GPT anomaly detection unit 140 (block 310). The early detection module 190 of the GPT anomaly detection unit 140 generates KPI intensity maps that are input to a GPT 420 along with the trained GPT-based LM (blocks 320, 330). The convolutional tokenizer 430 transforms the KPI intensity map for a time interval into discrete KPI token sequences for each time slot within the time interval. The KPI token sequences are then evaluated to detect anomalies as described in more detail below (block 340). The early detection module 190 may optionally generate abnormality heat maps that can be used for Route Cause Analysis (RCA) and troubleshooting problems in the network (block 350).

[0160] The objective of evaluation part is finding anomalies in the network data. This contains two subtasks. First, the session records that may contain anomalies should be collected and ordered by severity, second the exact place of unusual KPI activation part (which KPI(s), when the issue happened) should be identified for later processing with Route Cause Analysis (RCA). The RCA itself is out of scope from this patent.

[0161] The perplexity (PPL) is the best metric for finding anomalies with a trained LM. It measures how well a model predicts the likelihood of a given sequence of tokens. It gives back the average number of choices at each token position when it is predicted by the LM using the sequence of previous tokens. This means the lower the perplexity score, the better the model is at predicting the likelihood of a sequence. In other words, the model has a better understanding of the language of input sequence. For a given tokenized sequence s = ..., tn, the PPL is given by: where according to this formula the lower PPL indicate better fit to the trained language.

[0162] In the case of a set of multiple token sequences the formula above can be used by creating the mean of the PPL of individual sequences. The PPL can be also used to predict the next possible tokens after a given subsequence. This is the basis of individual identification of a uncertain part in a token sequence by incrementally evaluating each next token, where it is changing (the PPL is much more higher) this may indicate potential anomalies.

[0163] With GPT model the PPL can be computed with the exponential of crossentropy loss function. The language modeling task is to maximize the likelihood with the estimated PQ probability distribution. The loss is coming from the difference comparing with the T() target probability distribution, less loss means better modelling performance. Using these parameters, the PPT for a given s token sequence can be calculated according to:

[0164] PPLGPT(S) = exp cross_entropy P s K(s)) Eq. (14)

[0165] High PPL indicates uncertainty in evaluating the subsequences of the given KPI intensity map as it is valid part of the trained network language. The PPL can be calculated at every KPI-time map position. Putting together these PPL values KPI abnormality heatmaps are generated in the same dimension as the original input KPI intensity map. Figures 12 and 13 show exemplary abnormality heat maps. Figure 12 shows a full heat map and Figure 13 shows a partial heatmap. The abnormality is defined as a relatively high PPL at certain positions.

[0166] In the convolutional layer 130 the token extraction is executed in columnwise direction, that means all KPI tokens are enumerated in the same KPI order from the processed time slot before the processing of the next time slot. Furthermore, the GPT model is trained on all possible subsequences from the beginning, cutting them at the last tokens of timeslots. This enumeration method makes possible the partial model evaluation of GPT model until the current time in real-time monitoring systems. An example of partial KPI abnormality heatmap is shown on Figure 13. By analyzing the partial results, information can be extracted for early detection of network issues.

[0167] The analytic system 100 as herein described detects problems that are hidden (e.g., cannot be recognized by humans or are not possible to capture by monitoring single or a few parameter values, deterministic numerical thresholds). The analytic system 100 detects problems that cannot be described by rule parameters or crossing numerical thresholds. The analytic system 100 discovers complex patterns that are occurring at far distance in time and provides a set of new, unique type of signals of irregular network behavior.

[0168] The analytic system 100 detects irregularities well before possible real incidents or anomalies occur. Thus, the analytic system serves as a filtering and / or early warning apparatus. It reduces the number of possible real problems that occur. The analytic system can aggregate the detected irregularities periodically, which can serve as input to an early warning system. Like anomaly detection based on alarms or incidents, the early warning system analyze (postprocessed) irregularities by the available dimensions and helps to avoid real problems. The analytic system has a very low footprint due to the compression achieved by the applied convolution compared to existing methods. Results are explainable: KPI abnormality heatmaps show why exactly an entity was marked as not normal.

[0169] Figure 14 illustrates an exemplary method 500 of detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the analytic system 100. The analytic system 100 transforms a set of key performance indicators (KPIs) for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval (block 510). Analytic system 100 further transforms the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval (block 520). Analytic system 100 further detects irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data (530).

[0170] In some embodiments of method 500, transforming a set of KPIs for a network entity collected over a time interval into an intensity map comprise ordering time series of KPI values for one or more performance attributes into a two- dimensional KPI matrix, where each row of the KPI matrix corresponds to one performance attribute and each column corresponds to a time slot in the time interval; and transforming the KPI matrix into a KPI intensity map.

[0171] In some embodiments of method 500, ordering time series of KPI samples for one or more performance attributes into a two-dimensional KPI matrix comprises ordering the rows of the KPI matrix to optimize context correlation between KPIs according to an optimization criteria. In some embodiments of method 500, transforming the KPI matrix into a KPI intensity map comprises converting the KPI values in the KPI matrix into grayscale values to generate the KPI intensity map.

[0172] Some embodiments of method 500 further comprise normalizing the KPI values prior to transforming the KPI values into a KPI intensity map.

[0173] In some embodiments of method 500, transforming the KPI intensity map into one or more KPI token sequences comprises applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences.

[0174] In some embodiments of method 500, applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences comprises applying the convolutional transformer to extract tokens from the KPI intensity maps; and enumerating the extracted tokens column-wise to generate the KPI token sequences.

[0175] In some embodiments of method 500, detecting irregularities in network performance comprises evaluating the KPI token sequences using a generative pretrained transformer.

[0176] In some embodiments of method 500, evaluating the KPI token sequences using a generative pre-trained transformer comprises, for one or more time slots in the time interval, calculating a perplexity for each KPI in the time slot based on the token language model; and detecting irregularities based on the calculated perplexities.

[0177] In some embodiments of method 500, detecting irregularities based on the calculated perplexities comprises detecting irregularities based on a set of token sequences representing less than the entire time interval.

[0178] Some embodiments of method 500 further comprise generating a KPI heat map for at least a portion of the time interval based on the calculated perplexities.

[0179] In some embodiments of method 500, the set of KPIs for the network entity comprises data session records generated at predetermined time intervals.

[0180] In some embodiments of method 500, the network entity comprises one of: a subscriber; a cell; or a network function.

[0181] Figure 15 illustrates an exemplary method 550 of detecting irregularities in the performance of a wireless communication network using a GPT-based LM as may be performed by the GPT anomaly detection module 140. Analytic system 100 transforms a set of key performance indicators (KPIs) for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval (block 560). Analytic system 100 further transforms the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval (block 570). Analytic system 100 further trains the token language model using the KPI token sequences generated from the intensity maps (block 580). GPT anomaly detection module 140 may optionally detect irregularities in network performance by applying the token language model to token sequences generated from network data collected during network operation (block 590).

[0182] In some embodiments of method 550, transforming a set of KPIs for a network entity collected over a time interval into an intensity map comprises ordering time series of KPI values for one or more performance attributes into a two- dimensional KPI matrix, where each row of the KPI matrix corresponds to one performance attribute and each column corresponds to a time slot in the time interval; and transforming the KPI matrix into a KPI intensity map.

[0183] In some embodiments of method 550, ordering time series of KPI samples for one or more performance attributes into a two-dimensional KPI matrix comprises ordering the rows of the KPI matrix to optimize context correlation between KPIs according to an optimization criteria.

[0184] In some embodiments of method 550, transforming the KPI matrix into a KPI intensity map comprises converting the KPI values in the KPI matrix into grayscale values to generate the KPI intensity map.

[0185] Some embodiments of method 550 further comprise normalizing the KPI values prior to transforming the KPI values into a KPI intensity map.

[0186] In some embodiments of method 550, transforming the KPI intensity map into one or more KPI token sequences comprises applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences.

[0187] In some embodiments of method 550, applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences comprises applying a convolutional transformer to extract from the KPI intensity map; and enumerating the extracted tokens column-wise to generate the KPI token sequences.

[0188] In some embodiments of method 550, training the token language model using the KPI token sequences generated from the intensity maps comprises performing unsupervised training with a first maximum likelihood target using the KPI token sequences; and performing supervised training with a second maximum likelihood target using labeled KPI token sequences.

[0189] Some embodiments of method 550 further comprise detecting irregularities in network performance by applying the trained token language model to KPI token sequences generated from network data collected during network operation. In some embodiments of method 550, detecting irregularities in network performance by applying the trained token language model to KPI token sequences generated from network data collected during network operation comprises evaluating the KPI token sequences using a generative pre-trained transformer.

[0190] In some embodiments of method 550, evaluating the KPI token sequences using a generative pre-trained transformer comprises, for each of one or more time slots in the time interval calculating a perplexity for each KPI in the time slot based on the token language model; and detecting irregularities based on the calculated perplexities.

[0191] In some embodiments of method 550, detecting irregularities based on the calculated perplexities comprises detecting irregularities based on a set of token sequences representing less than the entire time interval.

[0192] Some embodiments of method 550 further comprise generating a KPI heat map for at least a portion of the time interval based on the calculated perplexities.

[0193] In some embodiments of method 550, the set of KPIs for the network entity comprises data session records generated at predetermined time intervals.

[0194] In some embodiments of method 550, the network entity comprises one of: a subscriber; a cell; or a network function.

[0195] An apparatus can perform any of the methods described herein by implementing any functional means, modules, units, or circuitry. In one embodiment, for example, the apparatuses comprise respective circuits or circuitry configured to perform the steps dedicated to performing certain functional processing and / or one or more microprocessors in conjunction with memory. For instance, the circuitry may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory may include program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein, in several embodiments. In embodiments that employ memory, the memory stores program code that, when executed by one or more processors, carries out the techniques described herein.

[0196] Figure 16 illustrates a network node 600 according to an embodiment for training and / or implementing a service quality model as herein described. The network node 600 comprises communication circuitry 620, processing circuitry 630, and memory 640 and can be configured by program instructions to perform the methods the anomaly detection and training methods herein described. The network node 600 may implement all or part of the GPT anomaly detection 140 shown in Figure 1.

[0197] Communication circuitry 620 comprises network interface circuitry for communicating with other core network nodes over a communication network, such as an Internet Protocol (IP) network.

[0198] Processing circuitry 630 controls the overall operation of the network node 300 and is configured to perform one or more of the methods as herein described. The processing circuitry 630 may comprise one or more microprocessors, hardware, firmware, or a combination thereof. The processing circuitry 630 can be configured by software to perform all or part of the methods described herein.

[0199] Memory 640 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 630 for operation. Memory 640 may comprise any tangible, non-transitory computer- readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. Memory 640 stores a computer program 650 comprising executable instructions that configure the processing circuitry 630 to implement one or more of the methods described herein. A computer program in this regard may comprise one or more code modules corresponding to the means or units described above. In general, computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM). In some embodiments, computer program 650 for configuring the processing circuitry 630 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media. The computer program 640 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0200] Figure 17 is a block diagram illustrating a virtualization environment 1100 in which functions implemented by some embodiments may be virtualized. In the present context, virtualizing means creating virtual versions of apparatuses or devices which may include virtualizing hardware platforms, storage devices and networking resources. As used herein, virtualization can be applied to any device described herein, or components thereof, and relates to an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1100 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node 5 may be entirely virtualized.

[0201] Applications 1102 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in the virtualization environment to implement some of the features, functions, and / or benefits of some of the 10 embodiments disclosed herein.

[0202] Hardware 1104 includes processing circuitry, memory that stores software and / or instructions executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 15 1106 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 1108a and 1108b (one or more of which may be generally referred to as VMs 1108), and / or perform any of the functions, features and / or benefits described in relation with some embodiments described herein. The virtualization layer 1106 may present a virtual operating platform that appears like networking hardware to the VMs 1108. 20

[0203] The VMs 1108 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer 1106. Different embodiments of the instance of a virtual appliance 1102 may be implemented on one or more of VMs 1108, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV 25 may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment.

[0204] In the context of NFV, a VM 1108 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, nonvirtualized machine. Each of 30 the VMs 1108, and that part of hardware 1104 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs 1108 on top of the hardware 1104 and corresponds to the application 1102. P104471W001 29 Hardware 1104 may be implemented in a standalone network node with generic or specific components. Hardware 1104 may implement some functions via virtualization. Alternatively, hardware 1104 may be part of a larger cluster of hardware (e.g., such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration 1110, which, among others, oversees lifecycle management 5 of applications 1102.

[0205] Those skilled in the art will also appreciate that embodiments herein further include corresponding computer programs. A computer program comprises instructions which, when executed on at least one processor of an apparatus, cause the apparatus to carry out any of the respective processing described above. A computer program in this regard may comprise one or more code modules corresponding to the means or units described above.

[0206] Embodiments further include a carrier containing such a computer program. This carrier may comprise one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0207] In this regard, embodiments herein also include a computer program product stored on a non-transitory computer readable (storage or recording) medium and comprising instructions that, when executed by a processor of an apparatus, cause the apparatus to perform as described above.

[0208] Embodiments further include a computer program product comprising program code portions for performing the steps of any of the embodiments herein when the computer program product is executed by a computing device. This computer program product may be stored on a computer readable recording medium.

Claims

CLAIMSWhat Is claimed is:

1. A method (500) of detecting irregularities in the performance of a wireless communication network, the method (500) comprising: transforming (510) a set of key performance indicators (KPIs) for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval; transforming (520) the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; and detect (530) irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

2. The method (500) of claim 1, wherein transforming a set of KPIs for a network entity collected over a time interval into an intensity map comprises: ordering time series of KPI values for one or more performance attributes into a two-dimensional KPI matrix, where each row of the KPI matrix corresponds to one performance attribute and each column corresponds to a time slot in the time interval; and transforming the KPI matrix into a KPI intensity map.

3. The method (500) of claim 2, wherein ordering time series of KPI samples for one or more performance attributes into a two-dimensional KPI matrix comprises ordering the rows of the KPI matrix to optimize context correlation between KPIs according to an optimization criteria.

4. The method (500) of claim 2 or 3, wherein transforming the KPI matrix into a KPI intensity map comprises converting the KPI values in the KPI matrix into grayscale values to generate the KPI intensity map.

5. The method (500) of any one of claims 1 — 4, further comprising normalizing the KPI values prior to transforming the KPI values into a KPI intensity map.

6. The method (500) of any one of claims 1 - 5, wherein transforming the KPI intensity map into one or more KPI token sequences comprises applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences.

7. The method (500) of claim 6, wherein applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences comprises: applying the convolutional transformer to extract tokens from the KPI intensity maps; and enumerating the extracted tokens column-wise to generate the KPI token sequences.

8. The method (500) of any one of claims 1 - 7, wherein detecting irregularities in network performance comprises evaluating the KPI token sequences using a generative pre-trained transformer.

9. The method (500) of claim 8, wherein evaluating the KPI token sequences using a generative pre-trained transformer comprises, for one or more time slots in the time interval:calculating a perplexity for each KPI in the time slot based on the token language model; and detecting irregularities based on the calculated perplexities.

10. The method (500) of claim 9, wherein detecting irregularities based on the calculated perplexities comprises detecting irregularities based on a set of token sequences representing less than the entire time interval.

11. The method (500) of claim 9 or 10, further comprising, generating a KPI heat map for at least a portion of the time interval based on the calculated perplexities.

12. The method (500) of any one of claims 1 - 11, wherein the set of KPIs for the network entity comprises data session records generated at predetermined time intervals.

13. The method (500) of any one of claims 1 - 12, wherein the network entity comprises one of: a subscriber; a cell; or a network function.

14. A method (550) of training a token language model to detect irregularities in the performance of a wireless communication network, the method (550) comprising:Transforming (560) a set of key performance indicators (KPIs) for a network entity collected over a time interval into an intensity map, wherein each pixel in the intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval; transforming (570) the intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; andtraining (580) the token language model using the KPI token sequences generated from the intensity maps.

15. The method (550) of claim 14, wherein transforming a set of KPIs for a network entity collected over a time interval into an intensity map comprises: ordering time series of KPI values for one or more performance attributes into a two-dimensional KPI matrix, where each row of the KPI matrix corresponds to one performance attribute and each column corresponds to a time slot in the time interval; and transforming the KPI matrix into a KPI intensity map.

16. The method (550) of claim 15, wherein ordering time series of KPI samples for one or more performance attributes into a two-dimensional KPI matrix comprises ordering the rows of the KPI matrix to optimize context correlation between KPIs according to an optimization criteria.

17. The method (550) of claim 15 or 16, wherein transforming the KPI matrix into a KPI intensity map comprises converting the KPI values in the KPI matrix into grayscale values to generate the KPI intensity map.

18. The method (550) of any one of claims 14 - 17, further comprising normalizing the KPI values prior to transforming the KPI values into a KPI intensity map.

19. The method (550) of any one of claims 14 - 18, wherein transforming the KPI intensity map into one or more KPI token sequences comprises applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences.

20. The method (550) of claim 19, wherein applying a convolutional transformer to transform the KPI intensity maps into KPI token sequences comprises: applying a convolutional transformer to extract from the KPI intensity map; and enumerating the extracted tokens column-wise to generate the KPI token sequences.

21. The method (550) of any one of claim 13 - 20, wherein training the token language model using the KPI token sequences generated from the intensity maps comprises: performing unsupervised training with a first maximum likelihood target using the KPI token sequences; and performing supervised training with a second maximum likelihood target using labeled KPI token sequences.

22. The method (550) of any one of claims 14 - 21, further comprising: detecting irregularities in network performance by applying the trained token language model to KPI token sequences generated from network data collected during network operation.

23. The method (550) of claim 22, wherein detecting irregularities in network performance comprises evaluating the KPI token sequences using a generative pretrained transformer.

24. The method (550) of claim 23, wherein evaluating the KPI token sequences using a generative pre-trained transformer comprises, for each of one or more time slots in the time interval:calculating a perplexity for each KPI in the time slot based on the token language model; and detecting irregularities based on the calculated perplexities.

25. The method (550) of claim 24, wherein detecting irregularities based on the calculated perplexities comprises detecting irregularities based on a set of token sequences representing less than the entire time interval.

26. The method (550) of claim 24 or 25, further comprising, generating a KPI heat map for at least a portion of the time interval based on the calculated perplexities.

27. The method (550) of any one of claims 14 - 26, wherein the set of KPIs for the network entity comprises data session records generated at predetermined time intervals.

28. The method (550) of any one of claims 14 - 27, wherein the network entity comprises one of: a subscriber; a cell; or a network function.

29. An anomaly detection system (140, 600) for detecting irregularities in the performance of a wireless communication network, the anomaly detection system (140, 600) being configured to: transform a set of key performance indicators (KPIs) for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval;transform the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; and detect irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

30. The anomaly detection system (140, 600) of claim 29, wherein the anomaly detection system (140, 600) is further configured to perform the method of any one of claims 2 - 13.

31. An anomaly detection system (140, 600) for detecting irregularities in the performance of a wireless communication network, the anomaly detection system (140, 600) comprising processing circuitry (630) and memory (640) cooperatively coupled to the processing circuitry (630), said memory (640) storing program instructions that when executed by the processing circuitry (630) cause the anomaly detection system (140, 600) to: transform a set of key performance indicators (KPIs) for a network entity collected over a time interval into a KPI intensity map, wherein each pixel in the KPI intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval; transform the KPI intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; and detect irregularities in network performance by evaluating the KPI token sequences using a token language model generated from sample training data.

32. The anomaly detection system (140, 600) of claim 31, wherein the anomaly detection system (140, 600) is further configured to by the program instructions to perform the method of any one of claims 2 - 13.

33. A computer program product (650) comprising program instructions that, when executed by processing circuity in an anomaly detection system (140, 600) for a wireless communication network, causes the anomaly detection system (140, 600) to perform the method of any one of claims 1 - 13.

34. A carrier containing a computer program (650) of claim 29, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

35. A non-transitory computer-readable storage medium containing a computer program (650) comprising executable instructions that, when executed by processing circuitry (630) in an anomaly detection system (140, 600) for a wireless communication network causes the anomaly detection system (140, 600) to perform the method of any one of claims 1 - 13.

36. An anomaly detection system (140, 600) for detecting irregularities in the performance of a wireless communication network, the anomaly detection system (140, 600) being configured to: transform a set of key performance indicators (KPIs) for a network entity collected over a time interval into an intensity map, wherein each pixel in the intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval;transform the intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; and train the token language model using the KPI token sequences generated from the intensity maps.

37. The anomaly detection system (140, 600) of claim 36, wherein the anomaly detection system (140, 600) is further configured to perform the method of any one of claims 15 - 28.

38. An anomaly detection system (140, 600) for detecting irregularities in the performance of a wireless communication network, the anomaly detection system (140, 600) comprising processing circuitry (630) and memory (640) cooperatively coupled to the processing circuitry (630), said memory (640) storing program instructions that when executed by the processing circuitry (630) cause the anomaly detection system (140, 600) to: transform a set of key performance indicators (KPIs) for a network entity collected over a time interval into an intensity map, wherein each pixel in the intensity map corresponds to a KPI value associated with a performance attribute at a given time slot in the time interval; transform the intensity map into one or more KPI token sequences, where each KPI token sequence corresponds a set of KPIs for a time slot in the time interval; and train the token language model using the KPI token sequences generated from the intensity maps.

39. The anomaly detection system (140, 600) of claim 38, wherein the anomaly detection system (140, 600) is further configured to by the program instructions to perform the method of any one of claims 15 - 28.

40. A computer program product (650) comprising program instructions that, when executed by processing circuity in an anomaly detection system for a wireless communication network, causes the anomaly detection system to perform the method of any one of claims 14 - 28.

41. A carrier containing a computer program (650) of claim 40, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

42. A non-transitory computer-readable storage medium containing a computer program (650) comprising executable instructions that, when executed by processing circuitry (630) in an anomaly detection system for a wireless communication network causes the anomaly detection system to perform the method of any one of claims 14- 28.

Citation Information

Patent Citations

  • Early detection of irregular patterns in mobile networks

    WO2024018257A1

  • Unified service quality model for mobile networks

    WO2024042346A1

  • Fingerprinting root cause analysis in cellular systems

    US20170201897A1

  • Pattern detection in time-series data

    US20190379589A1

Cited By

  • Pet equipment voice control method and system based on intelligent switching

    CN120808775A