Evaluation of the operation of a communication network

The method improves anomaly detection in telecommunications networks by using distinct time windows to predict future metrics, enhancing accuracy and responsiveness for proactive management.

FR3168476A1Pending Publication Date: 2026-05-15ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
ORANGE SA
Filing Date
2024-11-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing anomaly detection methods in telecommunications networks rely on fixed thresholds, simple statistical analyses, or supervised learning requiring labeled data, which are not proactive and lack real-time adaptability.

Method used

A method involving the extraction of network performance metrics over distinct time windows, predicting future values, and evaluating network operation using these predictions to enable proactive anomaly detection without labeled data.

Benefits of technology

Enhances the accuracy and responsiveness of anomaly detection by capturing different trends in network metrics, allowing for timely proactive management and improved resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for evaluating the operation of a telecommunications network is proposed, comprising: obtaining measurement data in the form of network performance metrics; extracting initial metric values ​​over a first time window (t) with a duration of t and second metric values ​​over a second time window (t) with a duration less than t; determining future metric values ​​from the first and second values ​​using a prediction module; and evaluating network operation by processing the predicted future values. Abstract Figure: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Evaluation of the operation of a communication network technical field

[0001] This disclosure falls within the field of telecommunications. More specifically, it relates to a method for evaluating the operation of a telecommunications network, as well as a corresponding computer device, telecommunications network management system, and computer program. Previous technique

[0002] In the prior art, various anomaly detection methods are used in telecommunication networks, including reactive detection techniques and static threshold systems applied to network performance metrics. These methods also include algorithms based on simple statistical analyses or classical machine learning approaches.

[0003] These existing methods notably propose the use of fixed thresholds to signal anomalies or simple moving averages and standard deviation calculations to track the variation of network metrics. Some more advanced techniques use supervised learning, but require labeled data for training anomaly detection models.

[0004] In this context, there is a need for a proactive anomaly detection solution that allows for the capture and analysis of network variations in real time without requiring labeled data. Summary

[0005] This disclosure improves the situation.

[0006] A method for evaluating the operation of a telecommunications network is proposed, comprising: obtaining measurement data in the form of network performance metrics; an extraction of first metric values ​​over a first time window (T) with a duration ATj and second metric values ​​over a second time window (T') with a duration AT2 less than ATp a determination, by a prediction module, of future values ​​of the metrics based on the first and second values; and an evaluation of the network's operation using a processing of predicted future values.

[0007] Predicting future values ​​of metrics, combined with processing predicted future values, enables proactive and efficient anomaly detection in the telecommunications network.

[0008] Extracting the first and second values ​​of metrics over time windows of distinct durations makes it possible to capture two distinct types of trends in the relevant network metrics. For example, the first time window can capture long-term trends, as well as variations over a longer time interval, and the second time window can capture recent variations and variations over a shorter interval, which is more suitable for detecting peaks or, more generally, rapid variations in traffic. Using the extracted values ​​to predict future metric values ​​can help improve the accuracy of the predicted future values ​​and, consequently, the network's performance, particularly compared to nominal operation, and therefore the effectiveness of anomaly detection.

[0009] According to another aspect, a computer device for evaluating the operation of a telecommunications network is proposed, the device comprising: a module for obtaining measurement data in the form of network performance metrics; a module for extracting initial metric values ​​over a first time window (T) with a duration AT1 and second metric values ​​over a second time window (T') with a duration AT2 less than AT1 a prediction module configured to determine future metric values ​​from the first and second values; and a module for evaluating network operation using a processing of predicted future values.

[0010] According to another aspect, a telecommunications network management system is proposed, comprising: such a device; and a notification module configured to issue an alert based on the result of an evaluation of network operation by the device; and / or a decision-making module configured to dynamically adapt the operation of the telecommunications network based on the result of an evaluation of network operation by the device.

[0011] The notification module enables proactive management of a malfunction following its detection and, potentially, before it actually occurs. This can reduce response time to critical problems and facilitate rapid interventions.

[0012] The decision-making module can contribute to improving resource allocation and network configuration and thus improve quality of service and / or quality of experience.

[0013] According to another aspect, a computer program is proposed comprising instructions which, when the program is implemented by a processor, lead to the implementation of the process as defined herein. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.

[0014] The features described in the following paragraphs may optionally be implemented independently of each other or in combination with each other.

[0015] In one example, the duration of at least one of the first time window and the second time window is dynamically adjusted taking into account the telemetry data obtained and / or taking into account a previous result of the evaluation of the operation of the network.

[0016] Such dynamic adjustment can allow for better adaptation to real-time operating conditions and increased responsiveness in anomaly detection. For example, if the data shows stability, durations can be extended to optimize resource utilization.

[0017] In one example, the determination of future values ​​takes into account a temporal and multivariate correlation between the extracted values.

[0018] This consideration can improve the accuracy of forecasts by incorporating the complex interactions, over time, between the values ​​of different metrics. This can make it possible to identify behavioral patterns that may indicate anomalies, thereby strengthening the robustness of the network performance assessment.

[0019] In one example, the evaluation of the network's operation takes into account only the extracted values.

[0020] This simplifies the analysis process while being based on relevant data and can thus increase the speed of proactive detection of anomalies and reaction to these anomalies.

[0021] In an example, AT2 > 50%

[0022] Such a large second time window allows for the observation of significant trends while remaining responsive enough to capture rapid variations. This enhances the anomaly detection capability while maintaining an acceptable level of accuracy.

[0023] In one example, the method includes an improvement of a determination process by the prediction module using a result of the evaluation of the network's operation.

[0024] The integration of the results of the performance evaluation to adjust the prediction process allows for continuous adaptation of the prediction module, thereby improving its ability to anticipate future variations and proactively detect anomalies.

[0025] In one example, the processing of predicted future values ​​includes a classification of predicted future values ​​into normal values ​​corresponding to nominal or expected behavior of data routing in the network and into abnormal values ​​corresponding to degraded behavior of data routing in the network, the classification being implemented by an anomaly detection algorithm on unlabeled predicted future values.

[0026] This classification can facilitate a clear understanding of the network's behavior. This can enable immediate, and potentially targeted, corrective actions to be taken if anomalies are detected.

[0027] In one example, the network performance metric values ​​include aggregated values ​​or distributions of at least one element from the following list: response times to requests, terminal registration times with a core network, PDU session establishment times, data transmission error rates, communication throughput, use of at least one computing resource, and use of at least one memory resource. Brief description of the drawings

[0028] Other features, details and advantages will become apparent from reading the detailed description below and from analyzing the accompanying drawings, in which: Fig. 1

[0029] [Fig.1] shows, in an example of an embodiment, a device for evaluating the operation of a system under study. Fig. 2

[0030] [Fig.2] shows, in an example of an embodiment, a management system for the operation of a system under study. Fig. 3

[0031] [Fig.3] shows, in an example of an embodiment, an influence of a preparation of measurement data on a prediction result. Fig. 4

[0032] [Fig.4] shows, in an example of an embodiment, an evaluation of a function of a 5G communication network. Description of the implementation methods

[0033] In the description that follows, identical reference numerals designate identical elements or elements having similar functions.

[0034] Some specific terms are now clarified for a better understanding of the proposed technique.

[0035] The term "system under study" refers to any organized set of elements of any kind whose behavior and interactions can be monitored and analyzed to evaluate their performance or operation. Such a system may vary depending on the field of application and may include components, interconnected devices, or even biological entities or financial data. Depending on the field and the nature of the system under study, the nature of the measurement data used to evaluate its operation may vary considerably.

[0036] In the field of telecommunications, the system under study may correspond to a communication network, such as a 5G, 4G, or wired network, integrating infrastructure equipment and user devices. The operation of this system can be evaluated by measuring performance indicators such as throughput, latency, or communication reliability, or more generally by measuring any metric related to one or more components of the system under study and / or to one or more interactions between such components. More specifically, obtaining and analyzing protocol data and possibly the variability of this data can make it possible to determine the operating mode of the 5G system.

[0037] In the field of health, the system under study may be a living organism (for example, a human being or an animal) or a group of people. The functioning of this system may then correspond to biological and health parameters, such as the results of medical examinations (for example, blood tests) or data from medical sensors, such as heart rate, oxygenation level, temperature, etc.

[0038] In the financial sector, the system under study may comprise one or more financial assets. The operation of the system may then correspond to the evolution of prices and associated financial indicators, allowing for an analysis of market fluctuations, volatility, and trends.

[0039] In the field of the Internet of Things (IoT), a system under study may include a network of interconnected sensors for monitoring environmental or industrial parameters, such as temperature, pressure, or humidity in a facility. The operation of this system can be evaluated by analyzing the variations of these parameters to ensure that the environment remains in optimal or safe conditions or that its operation complies with nominal conditions defined for this loT system.

[0040] The "evaluation" of the operation of the system under study consists of analyzing the behavior of the system to determine whether it conforms to expectations (in which case the behavior is described as normal) or whether it presents deviations that may indicate a malfunction (in which case the behavior is described as abnormal).

[0041] Measurement data is data obtained, directly or indirectly, by observing or recording a phenomenon, process, or activity. It can be qualitative or quantitative depending on the nature of the information captured. It can represent a measurement at a specific moment or an aggregation over a given period. Measurement data can be a raw measured value or a result derived from one or more processing steps applied to one or more measured values ​​(for example, an average, a percentile, a normalized value, a time variation, or any other result of statistical or non-statistical transformation). Thus, measurement data can include simple values ​​or combined results from several measurements of the same metric or from several correlated metrics. For example, in the field of communication networks, measurement data could be a response time or bandwidth consumption.In some cases, several metrics can be used together to characterize a complex measurement. For example, when measuring latency, it's possible to use a metric corresponding to the ICMP (Internet Control Message Protocol), which determines response times via "pings," and correlate it with data obtained using a specific response time calculator or application data from other protocols, such as the loading time of a web page via HTTP (Hypertext Transfer Protocol). This combination of metrics provides a more complete and reliable view of the observed latency, taking into account the variations inherent to each protocol or measurement tool. For example, in the healthcare field, a measurement could be a heart rate or a blood glucose level.For example, in the field of finance, a measurable data point could be a stock price. For example, in the field of the Internet of Things, a measurable data point could be a value measured by a sensor on a connected device, such as temperature or humidity level.

[0042] A measurement data type represents a specific aspect or dimension that is measured (for example, the temperature of a living organism, atmospheric humidity, the number of requests per second in a communication network). A measurement data type is data that is generally collected and stored from distinctly, but can include multiple dimensions. For example, a three-dimensional spatial position (x, y, z) is a single type of measurement data, although it consists of three distinct values ​​for its coordinates.

[0043] A measurement data category can group together types of measurement data sharing common characteristics related to their method of acquisition. For example, a distinction can be made between measurement data collected instantaneously (raw data obtained in real time) and data obtained over a defined period (aggregated over a time interval via statistical processing or transformations). For example, a temperature measurement taken continuously without aggregation is instantaneous data, while a temperature measurement averaged over a given interval is aggregated data. In communication networks, a continuous measurement of the RTT (round-trip time, which refers to the time it takes a signal to travel through a closed circuit) is instantaneous data, while an average RTT value over an interval is aggregated data.

[0044] A time window is a time interval during which measurement data are collected and analyzed. In the case of periodically collected measurements, a time window T of duration AT] contains X measurements, where the relationship between ATb X and the measurement step P is defined by X = AT j ! P J.

[0045] A time series is a sequence of measurement data values ​​associated with successive time windows. The time series allows for tracking changes, identifying trends, or detecting anomalies over an extended period. The windows in a time series can be arranged contiguously, partially overlap, or be distinct to meet the analytical objectives.

[0046] Reference is now made to [Fig. 1], which represents a possible example of an evaluation device 1 of the operation of a system studied according to the proposed technique.

[0047] Evaluating the operation of a system under study consists of detecting whether the network is likely to function normally or malfunction in the near future.

[0048] Device 1 comprises different logic modules, each defined by a specific function: a retrieval module 11, an extraction module 12, a prediction module 13, and an evaluation module 14.

[0049] The acquisition module 11 is configured to obtain measurement data. For example, a measurement data server may be provided to collect the measurement data, store it, and transmit it to the acquisition module 11.

[0050] Measurement data can be of one or more categories and of one or more types.

[0051] In the following notations, it is considered that: Measurement data allows us to study one or more interactions within a given system under study. The total number of entities involved in the interactions studied is denoted Vc; the Vc entities include P entities internal to the system and L entities external to the system, and the interactions studied implement K modes of interaction.

[0052] The extraction module 12 is configured to extract at least the following from the measurement data obtained: of the first values ​​of the measurement data, and of the second values ​​of the measurement data.

[0053] The first values ​​of the measurement data are obtained for at least a first time window having a first duration. The first values ​​of the measurement data can, for example, be obtained in the form of a time series for each type of measurement data, with the first value of such a time series being associated with a corresponding first time window.

[0054] The second values ​​of the measurement data are obtained for at least a second time window having a second duration shorter than the first duration. The second values ​​of the measurement data can, for example, be obtained in the form of a time series for each type of measurement data, with a second value of such a time series being associated with a corresponding second time window.

[0055] The prediction module 13 is configured to predict one or more future values ​​of the measurement data. The future values ​​of the measurement data can, for example, be obtained in the form of a time series by type of measurement data, a future value of such a time series being associated with a corresponding future time window.

[0056] The evaluation module 14 is configured to proactively evaluate a predictive operation of the system under study based, at least, on the future values ​​of the measurement data predicted by the prediction module 13.

[0057] In a possible application context, the measurement data are telemetry data from a telecommunications network (or system).

[0058] Telemetry data from such a network can include three main categories: logs, performance metrics, and distributed traces. These categories of measurement data are complementary and each serves distinct network monitoring and analysis purposes.

[0059] Logs represent a category of time-series measurement data, collected on an event-driven basis. They are often triggered by specific events or changes in network state, such as a user connection, disconnection, error, or warning. Logs contain precise information about each event, such as user IDs, IP addresses, timestamps, the type of event (e.g., error, warning, or information), and messages detailing the event or action.

[0060] Logs are primarily used for troubleshooting and event tracing. By analyzing logs, it is possible to trace the history of events in the network, diagnose incidents, and monitor interactions between network components.

[0061] In some systems, logs can be transformed into real-time metrics to provide an aggregated view of the network state (for example, by extracting error rates or connection throughput). However, this transformation requires considerable resources and depends on the quality of the logs generated by each application component.

[0062] Metrics represent continuous or periodic quantitative data, providing structured information on network performance and health. They include measures such as communication throughput, error rates, response times, and resource utilization (such as CPU, memory, or bandwidth).

[0063] Due to their digital nature, metrics are ideal for representing a system over time and maintaining a history in the form of time series. Metrics make it possible to detect anomalies by monitoring thresholds or trends; for example, a sudden increase in latency or a CPU overload may indicate performance degradation or a risk of failure.

[0064] Metrics are often preferred because they are easy to store in time series databases and their real-time monitoring is relatively inexpensive in terms of resources compared to other forms of telemetry. Thus, a use case in which measurement or telemetry data are directly obtained in the form of a time series, for example as metrics, is a particularly suitable use case for implementing the proposed technique.

[0065] Distributed traces follow the detailed communication paths and interactions between different network components during transactions or sessions, in particularly in distributed architectures based on microservices. Traces capture each step of a message or transaction's journey through the network, recording session IDs, components traversed, and arrival and departure times for each step.

[0066] Traces facilitate debugging and understanding the operation of a distributed application. They make it possible to reconstruct the path followed by a request in complex systems, identifying specific interactions and the propagation of events across different services.

[0067] Although extremely useful, traces have certain limitations. Their generation requires the involvement of all microservices to propagate context between components, which complicates their implementation. Furthermore, traces can generally only be examined once the request has been fully processed, making them less suitable for proactive, real-time actions. In addition, their large size often necessitates sampling to reduce collection and storage costs, which can lead to a loss of accuracy.

[0068] By way of illustration, a specific example of implementation is now detailed.

[0069] In this example embodiment, the system studied is a 5G telecommunications network.

[0070] The interactions studied are communications between V microservices internal to the 5G core network and U components external to the 5G core network, for example a 5G-AN (“5G-Access Network”) component, with VC = V+ U. The communications are implemented by K possible interaction modes, for example communication protocols or communication interfaces, for example HTTP / 2 “Hypertext Transfer Protocol / 2”, NGAP “Next Generation Application Protocol”, NAS “Non-Access Stratum”, PFCP “Packet Forwarding Control Protocol”, and GTP-U “General Packet Radio System Tunneling Protocol User Plane”, or end-to-end or E2E “end-to-end” procedures, for example the registration of a terminal, or “UE” for “user equipment”, with the core network.

[0071] The following description contains notations referring to a detailed model of 5G network metrics and protocols. These notations specifically take into account each of the network component Vc and each of the interaction modes considered. Such detailed modeling allows for a more nuanced analysis and a better understanding of the interactions between the network components.

[0072] As is known, in a 5G telecommunications network, the core network is structured into two functional planes. The control plane is responsible for managing signaling, authorizations, The user plane handles authentication and user session management. It regulates communications and defines data routing rules for each user. The user plane is responsible for the actual routing of data packets between users and end services. It processes data flows according to the rules defined by the control plane. Each functional plane includes network functions that are often deployed as containerized microservices. These microservices are often deployed and managed in an environment orchestrated by a platform such as Kubemetes. Such a platform facilitates the deployment, scaling, and management of containers running the network microservices.

[0073] In this example embodiment, the measurement data are telemetry data from the 5G telecommunications network, which are obtained in the form of performance metrics.

[0074] A first category of telemetry data comprises raw data corresponding to the initial values ​​obtained directly by observing or recording phenomena in the 5G network. This data, collected in real time, is instantaneous and captures conditions at the exact moment of measurement. It is useful for reflecting the current / current states of the network and providing detailed information without prior processing. In this context, "real time" means that the data is captured immediately, without any significant delay between its production and recording. This data can be collected continuously at defined intervals. When a fixed time step is defined, it can be associated with a collection rate, such as a measurement every second or a measurement every fifteen seconds.Real-time data can also be captured based on specific triggering events (e.g., each network request or each state change of a microservice).

[0075] In an example embodiment, a type of metric belonging to the first category of measurement data can be defined, namely: a set of MC memory resource consumption metrics.

[0076] The MC memory resource consumption metric set represents the memory resource usage of the 5G core network. Memory consumption refers to the use of RAM by each container hosting a microservice in the 5G network. In Kubernetes, each container is associated with a defined amount of memory it can use, with configurable limits to prevent overconsumption. This memory consumption can be monitored using tools built into Kubernetes (such as cAdvisor or Prometheus), which track the memory used by each container in real time. Kubernetes collects metrics such as Residential Storage (RSS) and cache memory consumed. by each container. MC can be formally defined as MC = {mc{i)}, where iG V. mc(i) is a value expressed, for example, in megabytes (MB), which quantifies the memory resource usage by one microservice * among the V microservices considered. Memory consumption is quantified instantaneously by measuring the size of memory reserved by each microservice at any given time.

[0077] A second category of telemetry data includes the results of statistical processing of a set of measurements taken not instantaneously, but over a time interval. These statistical processing results are aggregated values, for example, means, maximums, statistical distributions, etc.

[0078] In an example embodiment, four types of metrics belonging to the second category of measurement data can be defined, namely: a set of response time percentile metrics RTpj, a set of request rate metrics RRt, a set of error rate metrics sRt, and a set of computing resource consumption metrics CCT.

[0079] Response times for specific requests, measured between two points in the network, represent raw data. Distributions of these response times are then obtained by applying statistical processing (such as percentile determination) to the raw data. For example, for the HTTP / 2 protocol, the distribution of response times including one or more percentiles is calculated from the raw response time values ​​captured in real time for each request. The set of RTpj response time percentile metrics represents the p-th percentile of the response times observed in communications involving the 5G core network during the period T. This is the value below which a specific portion p, expressed for example as a percentage, of the observed response times lies. For example, RTso? can represent the set of median response times.RTpj can be formally defined as RTpt = {rTp:r(i, where k&K, G y2 and ie V. rTpT(i, j, k) is a value, expressed in units of time, which quantifies the ^-th percentile of response time to requests exchanged during the period T between a pair consisting of a microservice 1 among the V microservices considered and one j among the Vc microservices and external components considered according to one k among the K communication protocols and E2E procedures considered. .

[0080] For example, RTpr may include a metric denoting the distribution of response times to requests sent by a first microservice to a second microservice using the HTTP / 2 protocol. In this context, the response times are measured at the level of the first microservice, for example in the form of the median response time and the 90th, 95th and 99th percentiles of the response time.

[0081] For example, RTpr may include a metric T denoting the distribution of registration times of a terminal with respect to the core network. In this context, registration times are measured at the terminal level, for example in the form of the median registration time and the 99th, 99.9th and 99.99th percentiles of the registration time.

[0082] For example, RTpr may include a TPDU metric denoting the distribution of establishment times of a PDU (Packet Data Unit) session. In this context, establishment times are measured at the terminal level, for example in the form of the median establishment time and the 99th, 99.9th and 99.99th percentiles of the establishment time.

[0083] The set of request rate metrics RR? represents the communication rates involving the 5G core network during the period T. RRT can be formally defined as RR? = {rrT{i, j,k)},oùkeK, £ y2 and ie V. rrT( i, j, k) is a unitless value, expressed for example as a percentage, which quantifies the rate of requests exchanged during the period T between a pair consisting of a microservice 1 among the V microservices considered and one j among the Vc microservices and external components considered according to one k among the K communication protocols and E2E procedures considered.

[0084] The set of error rate metrics tRr represents the error rates observed in communications involving the 5G core network during the period T. eRt can be formally defined as eRt = [err(i, j,k)} , where k^K, ( i, j) G ct *e V- err(^ is a unitless value, expressed for example as a percentage, which quantifies the error rate observed during the period T between a pair consisting of a microservice * among the V microservices considered and one J among the Vc microservices and external components considered according to one k among the K communication protocols and E2E procedures considered.

[0085] The CCT compute resource consumption metric set represents the compute resource consumption of the 5G core microservices during period T. It can be, for example, the utilization rate of one or more 5G core processors. In containerization, CPU resources for each microservice can be precisely allocated in Kubernetes, which allows CPU limits to be defined (e.g., in fractions of vCPUs or cores). CCT can be formally defined as CCT = [ccT(i)] where ccr(i) is a value expressed, for example, in milliCPUs, which quantifies the utilization rate of computing resources observed during period T by a microservice i among the V microservices considered.

[0086] The set XT of telemetry data of the first category and the second category can be defined as follows: XT = RTp? U RRT U eR? U CCT.

[0087] The set X of telemetry data of the first category and the second category can be defined as follows: X = XT U MC.

[0088] In the example considered, the extraction of measurement data may include a transformation of measurement data into matrices.

[0089] The measurement data can, for example, be transformed, by the extraction module 12, into a first measurement data matrix X^it) ct a second measurement data matrix ,w x (0 these two matrices then provided as input to the prediction module.

[0090] (t) cst a first matrix of measurement data where each type of data The measurement xi is determined over a time window T of duration ATb

[0091] Thus, at any given moment, Xw(t) = x^-W+l) x2( / -W+l) X 1( / -1) ^i( / )l' x2(t -1) x2( / ) ^-1) ^)' where W represents the time interval covered by the matrix X^(f), Pæ" example of on the order of a few hours or days.

[0092] In this representation of XW(t), the time step separating two consecutive instants is set to ATj = 5 min. This choice of time step is dictated in the example embodiment considered by the need to meet the performance requirements commonly specified in Service Level Agreements (SLAs). A five-minute time window allows for the evaluation and monitoring of network performance frequently enough to ensure compliance with SLAs, while offering a granularity compatible with real-time monitoring objectives. Other time steps can be considered depending on the type of requirement.

[0093] The first measurement data matrix (t) is an example of a set of first values ​​of the measurement data over at least a first time window T having a duration ATb

[0094] x,w'(t) is a measurement data matrix where each type of measurement data x is determined over a time window T' of duration AT2 less than ATb

[0095] Thus, at any time. x'^-W'+l) x^-W'+l) .x'^-W'+l) where W' represents the time interval covered by the matrix ■ VF' can for example be less than or equal to VF.

[0096] In this representation of x1^ (t), the time step separating two consecutive instants t- 1 et' is equal to AT2.

[0097] The durations VF and / or VF' of the time intervals considered may be dynamic. For example, it may be expected that the measurement data will accumulate over time in the matrix XW(t) and / or in the matrix '

[0098] The durations VF and / or VF' of the time intervals considered may be fixed. For example, it may be provided that the matrix ( / ) and / or the matrix ) represent respectively a rolling measurement data history, so that as more recent measurement data is obtained, older measurement data is excluded.

[0099] The second measurement data matrix X^Xt) is an example of a set of second values ​​of the measurement data over at least a second window temporal T' having a duration AT2 < ATb

[0100] In general, the first and second values ​​of the measurement data can be organized in any format allowing the prediction module 13 to capture dependencies between the first values ​​of the measurement data and the second values ​​of the measurement data, without necessarily imposing a matrix structure.

[0101] The role of the prediction module is to determine, at time\, a prediction, denoted, of future measurement data at one or more future times t + S.

[0102] The prediction module 13 uses an artificial intelligence f} capable of receiving the measurement data provided as input and predicting future measurement data.

[0103] Artificial intelligence can for example be a neural network model, for example a recurrent neural network such as a closed recurrent unit (GRU for "Gated Recurrent Unit" in English).

[0104] Artificial intelligence is designed to recognize dependencies (or relationships, or correlations) in the measurement data provided as input.

[0105] These dependencies may include: strictly temporal dependencies, that is, between values ​​of the same variable at different times, for example between a value x£f- and a value a# and / or strictly spatial (or multivariate) dependencies, that is, between values ​​of distinct variables at the same instant, for example between a value and a value x20' and / or u spatio-temporal dependencies, that is, between values ​​of distinct variables at distinct times, for example between a value x^t- 1) and a value

[0106] Dependencies may include: dependencies between initial values ​​of the measurement data, and / or dependencies between second values ​​of the measurement data, and / or dependencies between at least one first value of the measurement data and at least one second value of the measurement data.

[0107] To illustrate this point, when the input data are provided in the form of matrices, for example the matrix XW(f) and the matrix ' ' these dependencies may include intra-matrix dependencies, that is to say between values ​​of the same matrix, for example a spatio-temporal dependency between a value X^t - 1) and a value x / t) of the first matrix ( / ) , that is to say a dependency between first values ​​of the measurement data.

[0108] Dependencies can also include inter-matrix dependencies, that is, between distinct matrix values, for example a spatio-temporal dependency between a value x 0 of the first matrix (t) ct a value x'^ _ । j of the second matrix ^,w'( / ), that is, a dependency between a first value of the measurement data and a second value of the measurement data.

[0109] Artificial intelligence uses the measurement data provided and the dependencies recognized within this measurement data provided to determine the prediction x(t + of the measurement data at at least one time t + S.

[0110] Formally, / S ' x^t+s) X^(t + S) ,xw(t+s)

[0111] Artificial intelligence f { can be trained with the objective of minimizing the mean squared error (MSE) between actual future values ​​X^Z + s) and predicted future values ​​x(tf + s)-

[0112] The training can be based on deep learning methods and / or machine learning methods

[0113] For example, such training can use training data sets formed from previously obtained measurement data. This previously obtained measurement data includes the first and second measurement data values. Furthermore, this previously obtained measurement data also includes the corresponding actual future values ​​as the training target. The mean squared error is used as the loss function L to measure the difference between the predicted future values ​​x(f + sj ct 'cs actual future values ​​X^ + s). The loss function is defined as follows: L = );

[0114] An optimization algorithm, such as stochastic gradient descent (SGD) or Adam, can be used to progressively adjust the artificial intelligence parameters during training, based on the gradients of the loss function calculated for each batch of training data, so as to minimize the loss function L.

[0115] The role of evaluation module 14 is to qualify the operation of the system under study, that is to say to detect one or more anomalies or an absence of anomaly.

[0116] The evaluation module 14 relies for this purpose at least on the data of future measurements determined by the prediction module. For example, the evaluation module may rely solely on the data of future measurements.

[0117] This detection can, for example, be based on one or more criteria of normality.

[0118] Normal values ​​are defined as those corresponding to an expected or nominal behavior of the system. These values ​​may be based on predefined or dynamic thresholds, established according to performance parameters such as expected service levels or other reference criteria.

[0119] To avoid relying solely on fixed thresholds, the evaluation module can use an artificial intelligence f2 capable of receiving predicted future measurement data and producing an anomaly vector + «)•

[0120] Formally, / t + S " +s) âï(t + s)

[0121] Artificial intelligence can, for example, be a K-means type algorithm, which allows for unsupervised classification of values ​​without imposing predefined criteria. This type of algorithm has the advantage of not requiring a labeled history of measurement data, previously classified into categories of normal operation and anomaly.

[0122] The algorithm partitions the predicted future measurement data into at least two clusters, comprising: at least one initial cluster identified as containing normal values ​​indicative of an expected behavior of the system under study, and at least one second cluster that can be identified as containing abnormal values ​​indicative of deviant, or unexpected, behavior of the system under study.

[0123] For example, and according to one possible convention, a K-means type algorithm can be designed to determine whether a given predicted future value xJ-(t + s) belongs to a first cluster or a second cluster cf and to fix the value of ci-(t + s) to: 0 if x\( t + S ) belongs to the first cluster (ÿ?, or 1 if x^(t + S) belongs to the second cluster cf-

[0124] The algorithm can be configured to partition the predicted future measurement data into more than two clusters, particularly when the system under study admits several distinct types of normal operation and / or several distinct types of anomalies.

[0125] The result of the evaluation of the functioning of the system studied can be exploited in several ways.

[0126] For example, the result of the evaluation can be used to continuously improve the process of determining future values ​​by the prediction tool. For example, in the case of frequent detection of anomalies for specific values ​​of certain parameters, the prediction tool can use the patterns of detected anomalies to adjust its future predictions.

[0127] For example, and as illustrated with reference to [Fig.2], device 1 can be integrated into a management system for the operation of a system under study comprising, in addition to device 1, a notification module 15 and / or a decision-making module 16.

[0128] The notification module 15 is configured to issue alerts based on the network operation assessment results provided by device 1. These alerts may be sent to maintenance teams, network administrators, or automated response systems. For example, it may be planned An alert should be issued when an anomaly is detected, and no alert should be issued when the system is functioning normally. Several alert levels (e.g., warning, critical) can be defined based on the type, severity, frequency, and expected duration of the detected anomalies. A single threshold breach can trigger a warning, while a series of significant anomalies across multiple metrics can generate a critical alert requiring immediate action.

[0129] The decision-making module 16 is configured to dynamically adapt the operation of the telecommunications network based on the evaluation results. It can react to detected anomalies to maintain the system under study in a normal operating state.

[0130] For example, continuing with the embodiment where the system under study is a 5G communication network, if an overload is detected on certain microservices, module 16 can decide to increase processing capacity by allocating more resources (CPU or memory) to the affected components. When an anomaly is detected in a specific part of the network, the module can redirect traffic to alternative routes to avoid congestion. After an anomaly is detected, an additional root cause analysis module can be activated to identify the source of the anomaly.This module can leverage a microservices dependency graph, which maps the relationships between different network services, to identify the services or components causing the anomaly, analyze interdependencies to understand how an anomaly in one service can affect other connected services, and provide a detailed diagnosis to facilitate proactive correction of the anomaly and limit its impact on the entire network.

[0131] In the example embodiment considered, the artificial intelligences and were tested in a real context.

[0132] The measurement data used were extracted from a network antenna (gNB) located at Lyon-Saint Exupéry Airport, which also covers a TGV station. The data represent the arrival and departure times of user equipment within the coverage area of ​​this antenna over a 10-day period, selected during school holidays, a period during which the airport and the station experience high traffic. The telemetry data collected during the first eight days were used to train the artificial intelligences f and / 2. Subsequently, the artificial intelligences f and / 2 were tested in an online environment, with a simulation reproducing the traffic patterns of the last two days of the period.

[0133] The performance of artificial intelligence f was evaluated using the following indicators.

[0134] MSE (Mean Squared Error), RMSE (Root Mean Squared Error), and MAE (Mean Absolute Error) respectively denote the mean squared error, its root, and the mean absolute error between the predicted future measurement data and the actual future measurement data. Ttrain denotes the training time, expressed for example in seconds.

[0135] The results of the evaluation are provided in Table 1. [Tables 1] MSE RMSE MAE T train h 0.0005 0.013 10.48 1283.48

[0136] The same indicators were used to evaluate the performance of known predictive modules for proactive anomaly detection, as a reference. According to all the indicators considered, the performance of artificial intelligence / 1 is superior to that of the state of the art.

[0137] These superior performances are due to the fact that fv takes into account not only first measurement data over a first time window (T) having a duration AT^ but also second measurement data over a second time window (T') having a duration AT2 less than AT(.

[0138] Figure 3 shows, by means of a histogram 2, the impact of data preparation by the extraction module on the mean squared error (MSE).

[0139] Bar 21 indicates the value of the mean squared error when f} takes into account first measurement data over first time windows (T) each having a duration ATj = 5 min but does not take into account any second measurement data.

[0140] Bar 22 indicates the value of the mean squared error when f takes into account first measurement data on first time windows (T) each having a duration ATj = 5 min and second measurement data on second time windows (T') each having a duration AT2- 1 min.

[0141] Bar 23 indicates the value of the mean squared error when f { takes into account first measurement data on first time windows (T) each having a duration AT} = 5 min and second measurement data on second time windows (T') each having a duration AT2 = 2 min.

[0142] Bar 24 indicates the value of the mean squared error when f} takes into account first measurement data on first time windows (T) each having a duration ATj = 5 min and second measurement data on second time windows (T') each having a duration AT2 = 3 min.

[0143] Bar 25 indicates the value of the mean squared error when taking into account first measurement data on first time windows (T) each having a duration AT! = 5 min and second measurement data on second time windows (T') each having a duration AT2 = 4 min.

[0144] It can be seen from the results in Figure 3 that, in the embodiment considered, taking into account the second set of measurement data is sufficient, regardless of the value of AT2, to reduce the mean squared error. It can also be seen that, in the embodiment considered, when AT2 > 50% AT1 (in particular when AT2 = 3 min or AT2 - 4 min), the mean squared error is reduced even further. Furthermore, it can be seen that, in the embodiment considered, when AT2 > 70% AT1 (in particular when AT2 = 4 min), the mean squared error is reduced even further.

[0145] Although in the example embodiment considered, the value of AT! is fixed at 5 minutes and the tested values ​​of AT2 vary in steps of one minute, these choices do not have general scope.

[0146] The values ​​of Y and AT2 can be adapted according to the specific characteristics of the system under study and the type of analysis desired. These durations can be either fixed or dynamically adjusted to meet the needs of evaluating the operation of the system under study.

[0147] The values ​​of ATj and / or AT2 can, for example, be chosen based on prior processing of the measurement data obtained. For instance, a criterion of stability of the measurement data obtained can be applied. If the measurements show prolonged stability (i.e., few variations), then a longer duration for ATj and / or AT2 can be preferred to avoid overly frequent processing and save resources. Conversely, if the measurement data vary significantly over a short period, shorter values ​​for ATj and / or AT2 can be applied to capture the changes more accurately.

[0148] The values ​​of ATj and / or AT2 can, for example, be chosen based on a previous evaluation of the network's operation. Shorter values ​​can be adopted for ATj and / or AT2 to facilitate enhanced monitoring after an anomaly has been detected. Conversely, shorter values ​​can be adopted for ATj and / or AT2 if the operation of the system under study is considered to be stable and normal.

[0149] Regarding artificial intelligence / 2, the performance of several candidate algorithms, all from the literature, was evaluated using the following indicators.

[0150] SC denotes a silhouette coefficient, CH denotes the Calinski-Harabasz index, and DB denotes the Davies-Bouldin index. These three quantities are related to clusters. trained by the evaluation module and allow to evaluate, according to distinct criteria, the performance of the proposed method of evaluating the operation of the 5G network. train designates the duration of training, expressed for example in seconds.

[0151] SC indicates the ratio between inter-cluster dispersion and intra-cluster dispersion, allowing for the measurement of cohesion and separation between clusters. SC can vary from ■°° to 1. Values ​​close to 1 indicate better cluster structuring.

[0152] CH allows for the evaluation of average similarity within clusters as well as dissimilarity between clusters. This index measures the compactness and separation of clusters, with higher values ​​indicating better-defined clusters.

[0153] DB evaluates clustering quality by taking into account intra-cluster similarity and inter-cluster dissimilarity. Lower values ​​of this index indicate more compact and better-separated clusters.

[0154] The evaluation results are provided in Table 2. According to the SC, CH and DB indicators, the performance of the K-Means algorithm is significantly superior to that of the other algorithms tested. [Tables 2] SC CH DB T tram HDBSCAN 0.07 28.93 171.55 221.30 Isolation Forest 0.50 1.75 13015.82 15.16 OCSVM 0.62 0.41 57674.06 0.53 Z-Score 0.70 0.42 83661.56 0.02 K-Means 0.79 0.28 124010.67 1.43

[0155] Figure 4 illustrates, in an example, the proactive capability of the method for evaluating the operation of the 5G network before the occurrence of the anomaly. The performance metric considered is a time series corresponding to the 99th percentile of the recording time in milliseconds of terminals (UEs) at the 5G core network.

[0156] Curve 31 shows the actual terminal recording times over time. Curve 32 shows the future values ​​successively predicted by the prediction module at successive current times, and points 33 indicate, among the predicted future values, those classified as anomalies by the evaluation module. It can be seen that the beginning 34 of the recording time plateau, occurring shortly after 3:30 p.m., is correctly and proactively detected at 3:29 p.m. as an anomaly 35. It can also be seen that the end 36 of the recording time plateau, occurring shortly after 3:35 p.m., is correctly and proactively detected at 3:35 p.m. as the absence of an anomaly 37.

[0157] This disclosure is not limited to the examples described above, only by way of example, but encompasses all the variants that a person skilled in the art may consider in the context of the protection sought.

[0158] Some examples of variants, applied to a communication network as a system under study, are specified by way of illustration.

[0159] The extraction module can be configured to extract first and / or second values ​​from traces, instead of, or in addition to, performance metrics. When used in conjunction with metrics, traces provide an additional level of detail that enriches performance analysis. Metrics can provide an overview of global indicators (e.g., average response time, error rate), while traces allow for deeper analysis by identifying the specific steps where anomalies occur. For example, traces can reveal critical dependencies between microservices that, if one component is overloaded, affect other parts of the system. If abnormal latency is detected via metrics, traces can be used to retrace the request path and isolate the microservice(s) responsible for the delay.Traces can also be analyzed to identify abnormal communication patterns or bottlenecks on specific network routes, which would not be apparent from aggregated metrics alone.

[0160] The artificial intelligence (AI)-based prediction module can be replaced by a prediction module based on a statistical model approach, such as an ARIMA (Auto-Regressive Integrated Moving Average) model, for cases where linear trend predictions are sufficient. This substitution can be useful for reducing the complexity of processing measurement data.

[0161] The evaluation module can be configured to rely on data from multiple external systems, and not only on internal network measurement data, thus enabling multi-channel analysis.

[0162] Feedback from the decision module to the prediction module may be provided to allow for dynamic adaptation of the prediction parameters (for example, an adjustment of the time window duration) according to operational conditions. Such feedback can thus improve the accuracy of the forecasts.

[0163] A telecommunications network management system may include several devices similar to the device described, each configured to monitor and evaluate different sections of the network. The management system may further include a centralized coordination module to aggregate the performance evaluations of each section.

Claims

Demands

1. A method for evaluating the operation of a telecommunications network, comprising: obtaining measurement data in the form of network performance metrics; extracting first values ​​of the metrics over a first time window (T) having a duration AT; and second values ​​of the metrics over a second time window (T') having a duration AT2 less than ATp; determining, by a prediction module, future values ​​of the metrics from the first and second values; and evaluating the operation of the network using a processing of the predicted future values.

2. A method according to claim 1, wherein the duration of at least one of the first time window and the second time window is dynamically adjusted taking into account the telemetry data obtained and / or taking into account a previous result of the network operation evaluation.

3. A method according to any one of the preceding claims, wherein the determination of future values ​​takes into account a temporal and multivariate correlation between the extracted values.

4. A method according to any one of the preceding claims, wherein the evaluation of the network's operation takes into account only the extracted values.

5. Method according to any one of the preceding claims, wherein AT2 > 50% AT],

6. A method according to any one of the preceding claims, comprising an improvement of a determination process by the prediction module using a result of the evaluation of the network's operation.

7. A method according to any one of the preceding claims, wherein the processing of predicted future values ​​comprises classifying the predicted future values ​​into normal values ​​corresponding to nominal or expected data routing behavior in the network and abnormal values ​​corresponding to degraded data routing behavior in the network, the classification being implemented by a anomaly detection algorithm on unlabeled predicted future values.

8. A method according to any one of the preceding claims, wherein the network performance metric values ​​comprise aggregated values ​​or distributions of at least one of the following: request response times, terminal registration times with a network core, PDU session establishment times, data transmission error rates, communication throughput, utilization of at least one computing resource, and utilization of at least one memory resource.

9. Computer device for evaluating the operation of a telecommunications network, the device comprising: a module for obtaining measurement data in the form of network performance metrics; a module for extracting first values ​​of the metrics over a first time window (T) having a duration ATj and second values ​​of the metrics over a second time window (T') having a duration AT2 less than ATp; a prediction module configured to determine future values ​​of the metrics from the first and second values; and a module for evaluating the operation of the network using a processing of the predicted future values.

10. A telecommunications network management system, comprising: a device according to claim 9; and a notification module configured to issue an alert based on the result of an evaluation of network operation by the device; and / or a decision-making module configured to dynamically adapt the operation of the telecommunications network based on the result of an evaluation of network operation by the device.

11. A computer program comprising instructions which, when the program is implemented by a processor, lead to the implementation of the method according to any one of claims 1 to 8.