System and method of determining set of alerts by runtime network monitoring
The method and system improve network observability by using service and network metrics to detect degradation, generate relevant alerts, and incorporate automated anomaly detection, enhancing accuracy and reducing false positives for proactive issue resolution.
Patent Information
- Application Number
- PCT/EP2024/066971
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
Existing network observability systems struggle to accurately identify and diagnose network incidents, lack comprehensive solutions incorporating external information, and fail to clearly connect symptoms to affected services, making it difficult for network engineers to understand and address service degradations.
A method and system that uses service level indicator metrics and network level metrics to detect service health degradation, generate candidate symptom expressions, and determine a set of alerts, incorporating automated anomaly detection and intent-driven networking to enhance monitoring accuracy and relevance.
The system reduces false positives, enables proactive issue resolution, and optimizes network reliability by generating alerts based on meaningful patterns and relationships, minimizing downtime and manual effort.
Smart Images

Figure EP2024066971_26122025_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD OF DETERMINING SET OF ALERTS BY RUNTIME NETWORK MONITORING
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to a field of intent-driven networking (IDN) and network observability and, more specifically, to a system and a method of determining a set of alerts by runtime network monitoring.
[0004] BACKGROUND
[0005] Advancements in the field of network observability systems have gained popularity over the years due to a plethora of applications, such as improving service reliability and ensuring network performance. The network observability systems are capable of monitoring Service Level Indicators (SLIs) metrics and other telemetry metrics that directly or indirectly describe the degradation of service health. However, the complexity of large-scale networks and the deployment of numerous services have made it challenging for network engineers to understand what constitutes normal and desirable states of network entities, such as devices and routing protocols, in order to fulfil the identified intents.
[0006] Currently, certain attempts have been made in the domain of the network observability systems, one of the key challenges is the ability to accurately identify and diagnose network incidents, which are confirmed network issues that impact customers of the network service and are related to Service Level Agreements (SLAs). Another problem is the lack of comprehensive solutions that incorporate external non-public technical information, ideas, or major designs to improve network performance. Additionally, the existing solution in this domain often lacks a clear connection between symptoms and services potentially affected by them, making it difficult for Network Reliability Engineers (NREs) to effectively understand and address service degradations.
[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with conventional systems and conventional methods of alert determination by runtime network monitoring.
[0008] SUMMARY
[0009] The present disclosure provides a system, a method, and a computer program of determining a set of alerts by a runtime network monitoring. The present disclosure provides a solution to the existing problem of how to automate the process of defining and managing alerts in the system by runtime network monitoring. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides an improved system and an improved method of determining a set of alerts by runtime network monitoring.
[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.
[0011] In one aspect, the present disclosure provides a method of determining the set of alerts by runtime network monitoring. The method includes using a plurality of service level indicator metrics, which describe a behaviour of at least one service being implemented on the network, to detect at least one time interval during which the at least one service being implemented on the network is demonstrating a degradation of service health. The method further includes using the detected at least one-time interval and a plurality of network level metrics to determine a plurality of candidate symptom expressions correlated with the at least one-time interval. The method further includes using the determined plurality of candidate symptom expressions to determine a set of alerts. Furthermore, the method includes activating the determined set of alerts and deploying the determined set of alerts to a runtime network monitoring system.
[0012] Advantageously, the method uses a comprehensive approach for the generation of the set of alerts. The combination of the service level indicator metrics and the network level metrics allows the system to find among the network level metrics any related behaviour that is likely impacting the service performance, allowing for a more thorough monitoring process. The method enables the detection of time intervals during which the service being implemented on the network demonstrates degradation of network service health. The early detection capability of the method allows for proactive intervention to address potential issues before they escalate, minimizing downtime and optimizing service reliability. By extracting candidate symptom expressions from detected time intervals and the network level metrics, the method ensures that the set of alerts eventually generated by the system are based on meaningful patterns and relationships within the network environment. This enhances the accuracy and relevance of the generated set of alerts, reducing false positives and alert fatigue. The method allows for the determination of the set of alerts tailored to the specific characteristics and requirements of the network environment. The customized approach ensures that the set of alerts are aligned with the unique needs of the network, enhancing their effectiveness in detecting and addressing potential issues. By basing the set of alerts on candidate symptom expressions that have been correlated with actual service degradation intervals, the method ensures that the alerts are highly relevant. This precision reduces the likelihood of false positives, ensuring that NREs are notified only of genuinely concerning behaviours that are likely to impact service health. Further, the automated deployment streamlines the implementation process, reducing manual effort and ensuring that the runtime network monitoring is equipped with the necessary alerts to effectively monitor network health.
[0013] In an implementation form, the method further includes using the at least one-time interval and the plurality of network level metrics to monitor the deployed set of alerts. In another implementation form, the method further includes using the at least one-time interval and the plurality of network level metrics to determine a set of symptom expressions which are used by a runtime network monitoring system to determine alerts.
[0014] Advantageously, by exposing the deployed set of alert expressions, expressed on the plurality of network level metrics and correlated to the time intervals, it becomes possible to identify the possible problem behind a network incident during troubleshooting or post-mortem analysis. The identification of the problem helps reduce the network expertise required to generate symptom expressions and enables the generation of network-specific alerts without the need for network engineers to possess detailed knowledge of the network.
[0015] By analyzing specific time intervals and a range of network-level metrics, the method can pinpoint exact conditions and patterns associated with service degradation. This precision enables the runtime network monitoring system to generate alerts that accurately reflect the current state of the network, facilitating prompt identification and resolution of issues.
[0016] In an implementation form, the result of the monitoring is used to determine an updated set of alerts with impact scores related to the deployed set of alerts, and wherein the monitoring tracks a relevance, with respect to service degradation events, of the deployed set of alerts over a time period. In another implementation form, the result of the monitoring is used to determine an updated set of symptom expressions with related impact scores, and wherein the monitoring tracks a relevance, with respect to service degradation events, of the deployed set of symptom expressions over a time period.
[0017] The method tracks the relevance of the set of alerts and the underlying symptom expressions over time. The tracking ensures that the system for the runtime monitoring is equipped with the most relevant and up-to-date symptom expressions. The up-to- date symptom expressions lead to a reduction in false alarms and enable quicker identification and resolution of service degradation issues. The regular review and removal of irrelevant symptom expressions also contribute to the overall effectiveness and adaptability of the system for runtime network monitoring, allowing it to remain reliable and efficient in detecting and addressing service degradation events. Furthermore, tracking the relevance of symptom expressions over time allows network operators to identify trends and patterns in service degradation events. By continuously refining the set of symptom expressions based on monitoring results and assigning impact scores, the method ensures that the runtime network monitoring system prioritizes alerts effectively. The optimization enhances the ability of the system to detect and respond to service degradation events promptly, improving network reliability and minimizing downtime.
[0018] In an implementation form, the plurality of service level indicator metrics include delay, packet drop or flow count.
[0019] Advantageously, the combined analysis of delay, packet drop, and flow count metrics allows the network operators to gain comprehensive insights into the network performance. The correlation in delay with packet drop metrics helps the network operators to detect congestion points and potential issues. The flow count metrics provide visibility into network traffic patterns, aiding in proactive issue resolution and capacity planning. The synergistic approach empowers the network operators to address network issues proactively, optimize data transmission, and enhance overall performance.
[0020] In an implementation form, the runtime network monitoring system receives service level indicator metrics, compares them against thresholds and generates alerts as a result of the comparing.
[0021] By continuously monitoring key performance indicators against predefined thresholds, the runtime network monitoring system may promptly identify deviations from normal network behaviour. The proactive approach of the runtime network monitoring system allows network administrators to address potential issues before they escalate into more significant problems, minimizing downtime and ensuring optimal network performance.
[0022] In an implementation form, the method further includes to determine the at least one-time interval using automated anomaly detection.
[0023] Advantageously, the automated anomaly detection efficiently analyses large data volumes, reducing manual monitoring efforts. The automated anomaly detection offers high accuracy, minimizing false alerts, and are scalable for varying network sizes. The continuous monitoring in real-time helps the automated anomaly detection promptly identify issues, enabling faster response times. Moreover, the automated anomaly detection adapts to changing network conditions over time, ensuring ongoing effectiveness.
[0024] In such an implementation form, the automated anomaly detection is heuristic-based.
[0025] The advantage of the heuristic-based automated anomaly detection is its adaptability to evolving network conditions. The adaptability helps in the detection of anomalies that may not conform to strict threshold criteria, improving the effectiveness of the method in identifying emerging issues and reducing false positives. Additionally, heuristic-based detection may uncover complex patterns and anomalies that may not be captured by simple threshold-based methods, enhancing the overall accuracy and reliability of anomaly detection in diverse network environments.
[0026] In another implementation form, the automated detection is machine learning based.
[0027] Advantageously, the machine learning-based automated detection enables the method to detect unexpected behaviours by leveraging anomaly detection techniques. The detection of unexpected behaviour allows for the identification of abnormal patterns that may indicate concerning behaviour. Further, the use of machine learning-based automated detection improves the overall detection accuracy by considering the importance of different features in the set of alerts, leading to a more precise identification of the behaviours that are concerning. In an implementation form, the detected at least one-time interval and a plurality of network level metrics to determine a plurality of candidate symptom expressions correlated with the at least one-time interval uses an artificial intelligence algorithm.
[0028] The use of artificial intelligence algorithm allows the method to efficiently process large amounts of data and identify relevant patterns or correlations that may not be easily detectable through traditional methods. Further, the artificial intelligence algorithm enhances the accuracy and efficiency of identifying the candidate symptom expressions. The artificial intelligence algorithm allows for automated analysis of the input data, reducing the reliance on manual inspection and potentially uncovering hidden patterns or relationships.
[0029] In such an implementation, the artificial intelligence algorithm is rule-based, template-based, or knowledge-based.
[0030] The utilization of rule-based, template-based, and knowledge-based algorithms in the method enables the generation of symptoms that are both accurate and understandable. The generation of accurate and understandable symptoms allows for improved identification and understanding of concerning behaviours in the network. The use of different types of the artificial intelligence algorithms, the method enhances the overall effectiveness of symptom generation, leading to more efficient and reliable detection of issues or anomalies in the network.
[0031] In an implementation form, the method uses root cause analysis algorithms to determine a set of alerts.
[0032] Advantageously, the root cause analysis algorithms pinpoint the precise underlying issues that lead to concerning behaviour within the network. The precision of the root cause analysis algorithms ensures that the set of alerts generated is highly accurate and directly addresses the root causes of any network anomalies. By automating the process of identifying these root causes, the method streamlines the symptom expression generation process, significantly reducing the time and effort required for manual investigation. The increased efficiency allows network engineers to allocate their resources more effectively, focusing on implementing solutions rather than spending valuable time diagnosing problems. Additionally, the root cause analysis algorithms continuously learn and adapt based on new data and feedback, optimizing the alert generation of the system over time.
[0033] In an implementation form, the runtime network monitoring system monitors a network which uses intent driven networking.
[0034] The intent-driven networking ensures that network operations align with specified intents and objectives, verifying that network services meet desired criteria. The runtime monitoring system dynamically adjusts parameters and thresholds based on changes in network intents or business requirements, ensuring effective monitoring. The intent-driven networking optimizes network performance by identifying resource allocation adjustments or policy refinements. Additionally, intent-driven networking enhances visibility into network behaviour and performance trends, empowering network administrators with deeper insights.
[0035] In another aspect, the present disclosure provides a system comprising means adapted for carrying out all the steps of the method.
[0036] The system achieves all the advantages and technical effects of the method of the present disclosure.
[0037] In another aspect, the present disclosure provides a computer program comprising instructions for carrying out all the steps of the method when said computer program is executed on a computer system.
[0038] It is to be appreciated that all the aforementioned implementation forms can be combined. It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.
[0039] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.
[0040] BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.
[0042] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:
[0043] FIG. 1 is a block diagram that depicts a runtime network monitoring system, in accordance with an embodiment of the present disclosure;
[0044] FIG. 2 is a flowchart depicting a method of determining a set of symptom expressions to be deployed to a runtime network monitoring system to generate alerts, in accordance with an embodiment of the present disclosure;
[0045] FIG. 3 is an exemplary diagram depicting execution of a method of determining a set of alerts, in accordance with an embodiment of the present disclosure;
[0046] FIG. 4 is an exemplary diagram depicting a graphical chart of the detection of concerning behaviors, in accordance with another embodiment of the present disclosure;
[0047] FIG. 5 is an exemplary diagram depicting a graphical chart of the generation of symptom patterns, in accordance with an embodiment of the present disclosure;
[0048] FIG. 6 is an exemplary diagram depicting a graphical chart of the alert definition, in accordance with an embodiment of the present disclosure; and
[0049] FIG. 7 is an exemplary diagram depicting a graphical chart of the dynamic alert, in accordance with an embodiment of the present disclosure. In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0050] DETAILED DESCRIPTION OF EMBODIMENTS
[0051] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0052] FIG. 1 is a block diagram that depicts a system configured to determine a set of alerts by runtime network monitoring, in accordance with an embodiment of the present disclosure. With reference to FIG.1, there is shown a block diagram 100 that includes a runtime network monitoring system 102 (hereinafter referred to as the system 102). The system 102 includes a processor 104, and a memory 106. In an implementation, the system 102 further includes an Artificial Intelligence (Al) Engine 110 and a network interface 112. The processor 104 is communicatively coupled with the memory 106, the Al engine 110 and the network interface 112.
[0053] The network interface 112 is configured to generate a set of alerts based on a historical data 108. The historical data 108 comprises service level indicators (SLI) metrics and network level metrics (NLM). An infrastructure diagram 114 comprises a SLI metrics block and a network level metrics (NLM) block. The infrastructure diagram 114 provides an overview of how the SLI metrics and the NLM are collected and derived from various network devices, protocols, and data flows for network observability purposes. The SLI metrics are detected between a first end point 116 of a virtual private network (VPN) and a second end point 118 of VPN. The segment routing network 120 enables the flow of the SLI metrics between the first end point 116 of VPN and the second end point 118 of VPN. In an implementation, the plurality of the SLI metrics include delay, packet drop or flow count. The delay also referred to as latency, measures the time it takes for data packets to travel from the source to the destination across the network. High delay may indicate congestion, network bottlenecks, or inefficient routing, which may degrade the user experience, especially for real-time applications like voice calls and video calls. Advantageously, monitoring the delay helps ensure that network performance meets the required standards and user expectations.
[0054] Further, the packet drop refers to the loss of data packets during transmission across the network. The packet drop may occur due to various reasons, such as network congestion, buffer overflow, or hardware failures. Monitoring the packet drop rates helps identify the network issues and ensures reliable data delivery. The flow count represents the number of active data flows or connections within the network. Monitoring the flow count provides insights into network utilization, traffic patterns, and resource allocation. By tracking the flow count, the network operators may optimize network performance, allocate bandwidth efficiently, and detect abnormal traffic behaviour.
[0055] Further, the network level metrics (NLM) block includes a base station 122, a network device 1, a network device 2, and a network device N. In an implementation, the network devices may represent routers, switches, or other networking equipment that are part of a core network infrastructure. The core network block may represent the core networking infrastructure that interconnects the network devices and enables communication between them and the VPN endpoints. The network devices may provide operational data such as interface Counters, CPU / Memory / OS, process information (e.g., Border Gateway Protocol or Multiprotocol Label Switching), control plane configuration, device logs, device configuration, and topological changes. The present disclosure provides the system 102 designed to determine a set of alerts by runtime network monitoring. The system 102 uses the processor 104 to analyse the historical data 108 i.e., the SLI metrics and the NLM. The processor 104 employs a plurality of the SLI metrics to identify time intervals indicating degraded service health for at least one network service. The analysis by the processor 104 may be heuristic-based or machine learning-based anomaly detection methodologies by using the Al engine 110. Furthermore, the processor 104 utilizes the Al engine 110 in the determination of candidate symptom expressions, which may be rule-based, template-based, or knowledge-based, enabling the generation of diverse candidate symptom expressions. Root cause analysis algorithms are also employed by the processor 104 to identify the underlying causes of service degradation events. The processor 104 determines the candidate symptom expressions correlated with degraded service health by utilizing the detected time intervals and the NLM and subsequently derives a set of alerts based on this analysis. The set of alerts are made accessible through the network interface 112 of the system 102 and are activated for the use by the system 102 for runtime network monitoring. The synergy between the processor 104 and the Al engine 110 enables effective processing and analysis of the historical data 108, ultimately providing the set of alerts for runtime network monitoring.
[0056] The processor 104 refers to a computational element that is operable to respond to and process instructions that drive the system 102. The processor 104 may refer to one or more individual processors, processing devices, and various elements associated with a processing device that may be shared by other processing devices. Additionally, the one or more individual processors, processing devices, and elements are arranged in various architectures for responding to and processing the instructions that drive the system 102.
[0057] Examples of the processor 104 may include but are not limited to, a hardware processor, a digital signal processor (DSP), a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, a data processing unit, a graphics processing unit (GPU), and other processors or control circuitry.
[0058] The memory 106 refers to a volatile or persistent medium, such as an electrical circuit, magnetic disk, virtual memory, or optical disk, in which a computer may store data or software for any duration. Optionally, the memory 106 is a non-volatile mass storage, such as a physical storage media. Examples of implementation of the memory 106 may include, but are not limited to, an Electrically Erasable Programmable Read-Only Memory (EEPROM), Dynamic Random-Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory.
[0059] The historical data 108 comprises past records, including the SLI metrics indicating service performance and health over previous time periods, along with the NLM capturing network behaviour and performance metrics during the same time frames. The historical data 108 undergoes retrospective analysis to pinpoint previous time intervals when network services exhibited concerning behaviours or suffered degradations. The analysis of the historical data 108 targets the discovery of patterns within the NLM that correlate with the identified time windows of concerning behaviour at the service level.
[0060] The Al engine 110 comprises AI / ML models, algorithms, and processing components. The Al engine 110 uses advanced techniques such as machine learning, deep learning, and natural language processing to automate aspects of alert definition and refinement. By analysing the network data, the Al engine 110 detects patterns, anomalies, and correlations that signal concerning network behaviours or service degradations. The Al engine 110, through machine learning, improves its ability to accurately detect and predict network issues by learning from the historical data 108. The network interface 112 may include a hardware or software that is configured to establish communication among the processor 104, the Al engine 110, and the memory 106. Examples of the network interface 112 may include, but are not limited to, a computer port, a network socket, a network interface controller (NIC), and any other network interface.
[0061] The set of alerts refers to a group of warnings generated by the system 102 to highlight particular network problems. The set of alerts may be established using predetermined thresholds and are sent for the runtime network monitoring. Each alert in the set of alerts may come with different attributes, like severity levels, affected services, and suggested actions, aimed at aiding in efficient response and resolution.
[0062] In operation, the processor 104 is configured to use the plurality of the SLI metrics, which describe a behaviour of at least one service being implemented on the network, to detect at least one-time interval during which the at least one service being implemented on the network is demonstrating the degradation of service health. The processor 104 uses the plurality of SLI metrics to capture the behaviour of one or more services operating within the network. Initially, the processor 104 selects the set of SLI metrics that are relevant to the specific network service being monitored. The SLI metrics are chosen based on the requirements and objectives of the service. The SLI metrics, which are time series metrics that directly reflect the behaviour of a service, identify time intervals in which the network service exhibits concerning behaviours. For example, the number of flows carried over the network to support a specific customer VPN may be used as the SLI metric. By monitoring Service Level Objectives (SLOs), which are predefined thresholds on predetermined behaviours that cause service degradations, the processor 104 labels the time intervals in which any SLI metrics deviate from their desired value as concerning. The monitoring ensures that the identified time intervals reflect real service impact and degradation of service health. Beneficially, by utilizing the SLI metrics and monitoring the SLOs, the processor 104 ensures that the identified time intervals reflect real service impact and degradation of the service health. The identification by the processor 104 enables the timely detection of service degradation and the deployment of targeted alerts, enhancing the overall performance and effectiveness of the system 102.
[0063] In accordance with an embodiment, the processor 104 is further configured to determine the at least one-time interval using automated anomaly detection. Using statistical methods, machine learning algorithms, or a combination of both, the processor 104 identifies anomalies in the SLI metrics that indicate potential service degradation. Anomalies may signify issues such as network congestion, hardware failures, software errors, or cyber-attacks.
[0064] In accordance with an embodiment, the automated anomaly detection is heuristic-based. Advantageously, by relying on heuristics, the processor 104 may quickly analyse large volumes of the SLI metrics and identify anomalies in real time. The analysis help in detecting potential threats or abnormalities promptly, allowing for timely intervention and mitigation. Additionally, the use of heuristics reduces the computational demand compared to more complex machine learning-based approaches, making it a practical solution for implementing the automated anomaly detection in the system 102.
[0065] In accordance with an embodiment, the automated detection is machine learning based. For example, the network is experiencing intermittent packet loss, which is causing degradation in the quality of a video streaming service. The traditional rule-based methods may struggle to detect the degradation in the quality of a video streaming service because the packet loss occurs sporadically and may not exceed predefined thresholds consistently. With the machine learning-based detection, the processor 104 may analyse the historical data 108 of the SLI metrics, such as packet loss rates, latency, and throughput, along with corresponding service performance indicators. By identifying complex patterns and correlations within these datasets, the machine learning-based automated detection may learn to recognize the subtle signs of packet loss-related service degradation.
[0066] In accordance with an embodiment, the processor 104 is further configured to use the detected at least one-time interval and a plurality of network level metrics to determine a plurality of candidate symptom expressions correlated with the at least onetime interval. The term "candidate symptom expressions" refers to observable patterns that may indicate the presence of a particular issue or problem within the network, serving as potential indicators or symptoms that may be used for troubleshooting, diagnosis, and resolution purposes. The processor 104 collects the plurality of NLM from various sources within the network infrastructure. Using the SLI metrics and other relevant information, the processor 104 identifies at least one-time interval during which the service being implemented on the network demonstrates degradation. The time interval serves as the focus for analysing the NLM. The processor 104 then correlates the time interval with the collected network level metrics to understand the network's behaviour during that specific period. The processor 104 looks for patterns or anomalies in the network level metrics that coincide with the detected degradation in service health. Based on the correlation analysis, the processor 104 generates the candidate symptom expressions. The candidate symptom expressions may be based on thresholds, patterns, or other expressions applied to the network level metrics that indicate potential issues or abnormalities during the identified time interval. The generated candidate symptom expressions undergo evaluation and refinement to ensure their relevance and accuracy in detecting service degradation. After refinement, the processor 104 finalizes the set of candidate symptom expressions that are highly correlated with the identified time interval of service degradation. By determining the candidate symptom expressions correlated with specific time intervals and the network level metrics, the processor 104 may better identify and address network issues, reducing false alarms and enabling more effective troubleshooting and analysis for further monitoring and alerting purposes within the system 102.
[0067] In accordance with an embodiment, the processor 104 uses an artificial intelligence algorithm to detect at least one-time interval and the plurality of NLM to determine the plurality of candidate symptom expressions correlated with the at least one-time interval. The Al algorithms may swiftly analyse vast amounts of network data, spotting potential issues without human input. They unearth subtle anomalies, enhancing accuracy beyond human capability. Additionally, they effortlessly scale to manage extensive networks, which is ideal for modem production system setups.
[0068] In accordance with an embodiment, the artificial intelligence algorithm is rule-based, template-based, or knowledge based. The rule-based Al algorithms are beneficial in alert determination because they allow for the creation of specific rules that encode domain knowledge about network behaviours and anomalies. The rules may be tailored to detect various conditions or events that may indicate the presence of network issues or abnormalities. By evaluating input data against the rules, the Al algorithms may quickly identify and trigger alerts for potential issues, enabling timely response and resolution. The template-based Al algorithms helps in alert determination by using predefined templates or patterns to recognize and classify different types of alerts. The templates capture common patterns or structures found in the data, allowing for efficient processing and identification of relevant alerts. By matching incoming data against predefined templates, the template-based Al algorithm may accurately classify alerts and trigger appropriate responses based on the detected patterns. The knowledge-based Al algorithms contribute to alert determination by leveraging a knowledge base containing facts, rules, and relationships about network behaviour. The knowledge-based Al algorithms use symbolic reasoning and inference mechanisms to derive new knowledge or make decisions based on existing information. The knowledge-based algorithms may understand the context of network events and make informed decisions about alert generation by encoding domain-specific knowledge. The knowledge-based algorithms may also break down complex alert scenarios into smaller, more manageable tasks, facilitating effective alert prioritization and response.
[0069] In accordance with an embodiment, the processor 104 is further configured to use the determined plurality of candidate symptom expressions, the at least one-time interval and the plurality of network level metrics, to determine a set of alerts. The processor 104 begins by reviewing candidate symptoms identified by network experts. The plurality of candidate symptom expressions are validated and translated into alerts as necessary. Further, the processor 104 selects and analyses the NLM to detect symptoms in the identified window. The meaningful patterns correlating with the plurality of candidate symptom expressions symptoms are then expressed using expression definitions such as thresholds, boundaries, and generic patterns. Further, the processor 104 uses root cause analysis algorithms to generate the set of alerts. The root cause analysis is a method used to identify the underlying causes of issues within a network. The root cause analysis aims to uncover the primary factors that contribute to undesirable outcomes rather than just addressing the symptoms. With the use of the root cause analysis, the processor 104 first clearly defines the issue that needs to be investigated. The issue identification may involve identifying the symptoms or manifestations of the issue and understanding its impact on the system 102. Further, the relevant SLI metrics and the NLM are collected to provide insight into the issue. Further, identification of potential causes is done that could be responsible for the problem. The identification is done by analysing historical data 108. Based on the identified possible causes, hypotheses are formulated to test the validity of each potential root cause. After testing the hypotheses, the processor 104 verifies whether it aligns with the observed symptoms and behaviours. After verification, the processor 104 tries to develop corrective actions to address the underlying issues and prevent the recurrence of the problem. The identified solutions are implemented within the system 102, and their effectiveness is monitored over time to ensure that the problem has been adequately addressed. The root cause analysis algorithms help to automate the process of identifying the root cause of network problems, saving time and effort for network administrators. Instead of manually investigating issues, administrators may rely on automated analysis to quickly diagnose problems and take appropriate actions.
[0070] In accordance with an embodiment, the processor 104 is further configured to use the at least one-time interval and the plurality of network level metrics to monitor the deployed set of alerts. For example, network engineers want to set alerts for abnormal patterns of network activity that could indicate a security threat. The system 102 provides identification of the alert criteria to be used (e.g. CPU usage above 90%, bandwidth usage above 95%, etc.), and these need to be provided as specific rules within the system 102. The rules define the conditions under which an alert should be triggered. In an example, if one wants to set a rule to generate an alert if CPU usage exceeds 90% for more than five minutes or if there are more than one hundred failed login attempts within a one-hour period. The process involves integrating the alerting system with the existing network infrastructure. The process may require installing monitoring agents on network devices, configuring switches to send syslog system, or using APIs to collect data from cloud-based services. The goal is to ensure that the alerting system has access to the necessary data to monitor network activity effectively.
[0071] Before deploying the alert rules into the system 102, it is essential to test the set of alerts thoroughly to ensure they work as intended. The testing phase might involve simulating different scenarios to validate that the system 102 may accurately detect and respond to the specified conditions. The simulation helps identify any potential issues or false positives that need to be addressed before deployment. After the alert rules have been tested and validated, they may be deployed into the system 102. The deployment involves activating the rules within the system 102 and ensuring that they are actively monitoring network traffic.
[0072] The symptom expressions serve as indicators of network health or performance issues. The symptom expressions are used to define the criteria for generating alerts when certain conditions or patterns are detected. The symptom expressions may encompass a wide range of factors, including variations in traffic patterns, changes in latency or packet loss rates, deviations from expected behaviour, and other relevant metrics. For example, consider a condition where the system 102 tracks various network-level metrics, such as latency, packet loss, and bandwidth utilization, over different time intervals. The system 102 is configured to determine symptom expressions based on these metrics to detect and alert on potential network issues. For example, a symptom expression for a latency spike may be defined as a significant increase in latency observed consistently over a short period. Specifically, this involves the metric of latency, measured in milliseconds, tracked over a 5-minute interval. The expression might state that if latency exceeds 100 milliseconds continuously for more than 5 consecutive minutes within any given hour, it indicates a potential performance degradation or issue, such as network congestion or a hardware malfunction. When the system 102 detects the symptom expression, it triggers an alert to notify network administrators of the latency spike. The prompt alert enables administrators to investigate and address the underlying cause swiftly, thereby maintaining network performance and reliability. In accordance with an embodiment, the processor 104 is further configured to use results of the monitoring to determine an updated set of alerts with impact scores related to the deployed set of alerts, and wherein the monitoring tracks a relevance, with respect to service degradation events of the deployed set of alerts over a time period. After monitoring network activity, the processor 104 uses the collected data to evaluate the effectiveness of the currently deployed set of alerts. Based on this evaluation, the processor 104 decides whether any adjustments or updates are necessary to improve the performance of the system 102.
[0073] In an example, the processor 104 detects a sudden increase in web traffic. The existing system 102 triggers alerts based on predefined thresholds for bandwidth usage. However, upon analysis, the processor 104 determines that the alerts do not adequately capture the severity of the situation. The processor 104 updates the alert criteria to include additional parameters, such as the number of concurrent connections or the geographical distribution of the traffic, to better identify and respond to potential issues. The processor 104 assigns impact scores to the alerts based on their significance and potential consequences. The impact scores help prioritize alerts and determine the appropriate response actions.
[0074] In an example, the system 102 detects a critical security breach. The processor 104 assigns a high impact score to this alert, indicating its severity and the need for immediate action. Meanwhile, alerts for minor performance fluctuations might receive lower impact scores, signalling that they may be addressed with lower priority. The system 102 tracks the relevance of deployed alerts with respect to service degradation events over a period. The tracking helps assess the effectiveness of the alerting system and identify any trends or patterns in network behaviour.
[0075] Over time, the processor 104 analyses how well the deployed alerts correlate with actual service degradation events. The processor 104 tracks metrics such as false positive rates, detection accuracy, and response time to evaluate the alerts’ relevance. If certain alerts consistently fail to align with service degradation events or if they become less relevant over time, the processor 104 may adjust or retire those alerts accordingly.
[0076] Advantageously, by using monitoring results to update the set of alerts, the system 102 may adapt to evolving network conditions and security threats, leading to more accurate detection and fewer false positives. Assigning the impact scores to alerts helps prioritize them based on their severity, allowing the network administrators to focus on the most critical issues first and respond promptly to potential threats or service disruptions. Tracking the relevance of deployed alerts over time enables the system 102 to identify areas for improvement and refine alerting criteria to better align with actual network events, ensuring ongoing effectiveness and reliability.
[0077] The system 102 uses the network-level metrics, such as latency, packet loss, and bandwidth utilization, over different time intervals metrics to determine symptom expressions, which are patterns indicating potential network issues. Initially, the system 102 begins by monitoring the network and identifying initial symptom expressions, such as a latency spike defined as latency exceeding 100 milliseconds for more than 5 consecutive minutes within an hour. When the system detects this symptom expression, it generates an alert, notifying network administrators to investigate and resolve the issue. The system 102 continues to monitor the network, collecting data on the performance and relevance of the deployed symptom expressions, tracking how often these expressions lead to alerts and whether these alerts correspond to actual service degradation events. Over time, the system 102 analyzes the collected data, assessing the impact of each symptom expression by assigning impact scores based on the severity and frequency of the corresponding alerts. If the latency spike expression frequently triggers alerts corresponding to significant performance issues, its impact score would be high, while expressions leading to false positives would be reevaluated. The system 102 tracks the relevance of each symptom expression, evaluating how well each predicts actual network issues and adjusting for changes in network behaviour or conditions. Based on impact scores and relevance tracking, the system 102 updates the symptom expressions. For instance, if a specific latency threshold is too sensitive, it may be adjusted to reduce false positives. New symptom expressions may be added based on emerging patterns, while less relevant or outdated expressions are modified or removed. For example, an initial latency spike expression might be adjusted to latency exceeding 120 milliseconds for 5 consecutive minutes within an hour to reduce false positives, and a new packet loss expression might be added for packet loss exceeding 3% for 10 consecutive minutes. The continuous updating leads to more accurate alerts, allowing network administrators to respond more effectively to actual issues, thereby reducing downtime and maintaining high service quality. By using monitoring results to refine symptom expressions with impact scores and relevance tracking, the system 102 remains accurate and responsive to real network conditions. The adaptive approach enhances ability of the system 102 to detect genuine issues while minimizing false positives, ultimately improving network reliability and performance. In accordance with an embodiment, runtime network monitoring system 102 receives service level indicator metrics, compares them against thresholds and generates alerts as a result of the comparing. The continuous monitoring of the SLI metrics and comparing them against predefined thresholds, the system 102 may detect issues in real time as they occur, allowing for prompt response and mitigation. Alerting based on threshold comparison enables proactive maintenance and problem resolution before issues escalate and impact network performance or user experience. Prioritizing alerts based on severity levels helps network administrators allocate resources effectively, focusing their attention on critical issues that require immediate action while optimizing resource utilization. The identification of deviations from expected service levels helps the system 102 to maintain and improve overall service quality, enhancing user satisfaction and trust in the network infrastructure.
[0078] In accordance with an embodiment, the runtime network monitoring system 102 monitors a network which uses intent driven networking. The runtime network monitoring system 102 constantly gathers data from network devices like routers, switches, and servers, as well as applications. The metrics related to performance, traffic patterns, resource usage, and service availability. The system 102 correlates its activities with the intended goals and policies of the network, which administrators set. The intentions may include maintaining high availability, optimizing performance for specific applications, enforcing security measures, or prioritizing certain types of traffic.
[0079] For example, if the intent of the network is to prioritize VoIP (Voice over Internet Protocol) traffic to ensure high call quality, the system 102 may continuously monitor network traffic and prioritize packets tagged as VoIP, ensuring they receive sufficient bandwidth and minimal latency.
[0080] Advantageously, by the continuously monitoring network behaviour against predefined intentions, the system 102 may proactively identify deviations or issues that may impact network performance. With real-time insights into network behaviour, the system 102 may automatically adjust network configurations to maintain alignment with intended objectives. By ensuring that the network operates in accordance with its intended design and objectives, the system 102 contributes to improved performance, reliability, and user experience.
[0081] FIG. 2 is a flowchart depicting a method of determining a set of alerts to be deployed to a runtime network monitoring system, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a flowchart of a method 200 of determining the set of alerts to be deployed to the system 102. The method 200 includes a series of operations 202 to 208.
[0082] At operation 202, the method 200 includes using the plurality of service level indicator metrics which describe the behaviour of at least one service being implemented on the network, to detect at least one-time interval during which the at least one service being implemented on the network is demonstrating a degradation of service health. For example, the processor 104 continuously collects data on these SLI metrics from various points in the network. The processor 104 analyses the SLI metrics over time to establish baseline performance levels. For example, it may find that under normal conditions, latency stays below 100 milliseconds, throughput is consistently above 5 Mbps, the error rate is less than 1%, and availability is maintained at 99.9%. During peak usage hours, the processor 104 notices a sudden increase in latency, with some users experiencing delays of up to 500 milliseconds. Additionally, throughput drops below 2 Mbps, and error rates spike to 5%. The deviations from the baseline indicate a degradation in service health. To confirm the degradation, the processor 104 cross-references the SLI metrics. The processor 104 detects that the increase in latency coincides with the decrease in throughput and the spike in error rates, suggesting the systemic issue rather than isolated incidents.
[0083] At operation 204, the method 200 includes using the detected at least one-time interval and the plurality of network level metrics to determine the plurality of candidate symptom expressions correlated with the at least one-time interval. For example, processor 104 detected a time interval during which there was a significant increase in latency and a decrease in throughput, indicating a degradation in service performance. The time interval spans from 8:00 PM to 10:00 PM. The processor 104 collects various NLM during this time interval, such as bandwidth utilization, packet loss rate, CPU and memory usage of network devices, number of concurrent connections, and network congestion level. Using the network level metrics, the processor 104 identifies patterns and correlations that may indicate the root cause of the performance degradation. For example, bandwidth utilization consistently exceeds 90% during peak hours, packet loss rate increases when bandwidth utilization is high, CPU and memory usage of core network devices spike during the same time interval, and network congestion levels are significantly higher during peak usage hours. Based on these observations, the processor 104 determines the candidate symptom expressions correlated with the degraded time interval. For example, high bandwidth utilization coupled with increased packet loss rate and CPU / memory usage suggests network congestion as a potential cause. The correlation between the network congestion levels and the degraded service performance during peak hours further supports this hypothesis.
[0084] At step 206, the method 200 includes using the determined plurality of candidate symptom expressions, the at least one-time interval and the plurality of network level metrics, to determine the set of alerts. The processor 104 integrates the candidate symptom expressions identified during the analysis phase. The candidate symptom expressions are indicators of potential issues or abnormalities within the network. The processor 104 correlates the identified symptom expressions with the specific time intervals during which they occurred. The correlation helps establish temporal context and understand when the issues or anomalies occurred. Alongside the candidate symptom expressions and time intervals, the processor 104 analyses the plurality of the network level metrics. The network level metrics provide additional contextual information about the overall health and performance of the network. Based on the integrated symptom expressions, time intervals, and the network level metrics, the processor 104 establishes conditions for triggering alerts. The conditions typically involve patterns indicative of potential issues or degradation in service health. Upon meeting the conditions, alerts are generated by the system 102 to notify network administrators or operators about the detected issues or anomalies.
[0085] By generating targeted alerts based on specific conditions, the method 200 helps prioritize resources and efforts towards addressing critical issues, optimizing operational efficiency. Timely alerts facilitate prompt action, reducing downtime and minimizing the impact of network disruptions on users and business operations.
[0086] At step 208, the method 200 includes activating the determined set of alerts and deploying the determined set of alerts to a runtime network monitoring system. When the processor 104 identifies anomalies from expected network behaviour based on predefined criteria, corresponding alerts are activated. Activation involves setting triggers or thresholds for specific conditions or events that warrant attention. Upon activation of the alerts, they need to be deployed to the system 102. The deployment involves integrating the set of alerts into the framework of the system 102 so that they may be effectively monitored and managed. The integration includes configuring the system 102 to recognize and respond to the activated set of alerts. For an example, the detection of anomalies in network traffic patterns is done by the system 102. The predefined alerts are triggered when the network traffic exceeds 90% of its maximum capacity for more than five minutes. The spike in network traffic that surpasses the predefined threshold is detected. It activates the alert for high network utilization. The activated alert is deployed to the system 102. Additionally, the alert may be logged for further analysis. Activating and deploying alerts promptly ensures that network issues are identified and addressed in a timely manner, reducing potential downtime. Further, integration of the set of alerts into the system 102 helps administrators to efficiently track network health and performance, allowing for proactive management and mitigation of issues.
[0087] Advantageously, an effective alert activation and deployment contribute to the overall reliability of the network by enabling quick responses to potential anomalies. The determination is done automatically, eliminating the need for manual intervention. The alerts are specifically tailored to the network being monitored, considering factors such as network topology, vendors, and protocols. The process does not require network engineers to possess in-depth knowledge of these network-specific details. By automating the alert generation process, the method 200 reduces the reliance on network engineers and their associated costs. It also saves time when designing, troubleshooting, and analysing network incidents.
[0088] Advantageously, the method 200 uses a comprehensive approach for the generation of the set of alerts. The combination of the SLI metrics and the NLM ensures that potential issues affecting the service's health are detected at both the service and network levels, allowing for a more thorough monitoring process. The method 200 enables the detection of time intervals during which the service being implemented on the network demonstrates degradation of network service health. The early detection capability of the method 200 allows for proactive intervention to address potential issues before they escalate, minimizing downtime and optimizing service reliability. By correlating the candidate symptom expressions with detected time intervals and network level metrics, the method ensures that the set of alerts are based on meaningful patterns and relationships within the network environment. The correlation enhances the accuracy and relevance of the generated set of alerts, reducing false positives and alert fatigue. The method 200 allows for the determination of the set of alerts tailored to the specific characteristics and requirements of the network environment. The customized approach ensures that alerts are aligned with the unique needs of the network, enhancing their effectiveness in detecting and addressing potential issues. Further, the automated deployment streamlines the implementation process, reducing manual effort and ensuring that the runtime network monitoring is equipped with the necessary alerts to monitor network health effectively.
[0089] The steps 202 to 208 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.
[0090] There is provided a computer program comprising instructions that, when executed by a computer system, cause the computer system to implement the method 200. In an example, the instructions are implemented on the computer-readable media, which include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory. In an example, the instructions are generated by a computer program, which is implemented in view of the method 200 for generating the set of alerts for the system 102 by runtime network monitoring.
[0091] FIG. 3 is an exemplary diagram depicting the execution of the method of determining the set of alerts, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIGs 1 and 2. With reference to FIG. 3, there is shown an exemplary diagram 300 that depicts the use of the system 102 (of FIG. 1) to generate the set of alerts by runtime network monitoring. There is shown two phases for generation of the set of alerts to be deployed into the system 102. The first phase involves analyses of the historical data 108 and a second phase involves the deployment of the set of alerts to the system 102 for runtime network monitoring. In an exemplary scenario, first phase involves the historical data 108 analysis. The historical data 108 is represented in the form of a SLI telemetry represented by a first graph 302B and network level metrics represented by a second graph 302C. Further, Service Level Objective (SLO) definitions are represented by a third graph 302A.
[0092] At operation 302, detection of concerning behaviours is performed. The processor 104 takes the SLI metrics as input to detect and produce the list of time intervals in which the network service is exhibiting concerning behaviours. The SLI telemetry uses the SLO definitions for detecting time intervals that are exhibiting concerning behaviours. The SLO definitions are predefined thresholds on predetermined behaviours that are clearly generating service degradations.
[0093] At operation 304, the generation of symptoms patterns takes place. The processor 104 combines the time intervals of concerning behaviour 302D identified at operation 302 with the network level metrics for the generation of symptoms patterns. The system pattern are candidate symptom expressions that are combinations of metrics that indicate potential network issues. The candidate symptom expressions help to pinpoint specific areas of concern within the network. For instance, if packet loss is consistently high during specific time intervals and correlated with increased latency, this may be the candidate symptom expression indicating network congestion. Further, definition of identified symptoms 304A are obtained using the candidate symptom expressions.
[0094] At operation 306, the processor 104 generates the alert definition using the candidate symptom expressions and network level metrics to produce a final set of alerts to be deployed 306A. The process involves filtering out low-quality symptoms and aggregating high-quality ones to create alerts that are accurate and actionable. It ensures that the alerts are understandable and relevant to network administrators. For example, the processor 104 may group similar symptoms together and prioritize alerts based on their impact on network performance.
[0095] At operation 308, the dynamic alert evaluation takes place. The dynamic alert evaluation is part of the second phase, i.e. realtime network monitoring. In the second phase, runtime data 310, which includes the real-time Service Level Indicator (SLI) telemetry and real time network level metrics. The real-time Service Level Indicator (SLI) telemetry is represented by a fourth graph 310A and the real time network level metrics are represented by a fifth graph 310B. The runtime data 310 is continuously ingested into the system 102 to update the impact scores of deployed alerts.
[0096] Additionally, Service Level Objective SLO definitions used by the system 102 are represented by a sixth graph 312. It is to be noted that functionality of the sixth graph 312 is same as the third graph 302A i.e. both graphs represent the same parameters. The SLO definitions are also ingested into the system 102. The runtime data 310 refers to the information collected during the operation of the system 102. The continuous ingestion of the runtime data 310 along with the SLO definitions allows the system 102 to dynamically update the impact scores of deployed alerts. Impact scores indicate the severity or significance of an alert in relation to network performance and service health. By incorporating the runtime data 310, the system 102 may accurately assess the relevance and effectiveness of alerts in identifying potential issues or service degradations. Furthermore, the ingestion of the SLO definitions enables the system 102 to adapt to changes in service level objectives, which specify the desired performance or quality standards for network services. By integrating the SLO definitions, the system 102 may align alerting criteria with evolving service requirements, ensuring that alerts remain relevant and actionable in maintaining service reliability and quality. By utilizing the runtime data 310, the system 102 may promptly detect and respond to changes or anomalies in network performance, minimizing downtime and service disruptions. Furthermore, continuous ingestion of real time SLI telemetry and the real time network level metrics enables the system 102 to generate alerts based on up-to-date information, ensuring that alerts are relevant and indicative of actual service conditions. The alert definitions are the ones to which the impact score is assigned. Incorporating the SLO definitions allows the system 102 to adjust alerting criteria to meet evolving service objectives and requirements, ensuring that alerts align with business priorities and objectives. After generation of alerts the system 102 generates new time intervals of concerning behaviours 314 which raises alerts. The system 102 after analysing the concerning behaviour intervals updates the impact score assigned to the related alert definition 316. Advantageously, by analysing the runtime data 310, new time intervals of concerning behaviours 314, deployed alerts, and the network level metrics refines the set of alerts that no longer reflect the current state of the network or have a high false alarm rate are automatically disabled. The automatic disablement ensures that the system 102 adapts to changing network conditions and maintains the relevance of its alerts. For instance, if a previously deployed alert for high latency becomes irrelevant due to network optimization, the system 102 automatically disables it to prevent unnecessary notifications.
[0097] FIG. 4 is an exemplary diagram depicting a graphical chart of the detection of concerning behaviors, in accordance with another embodiment of the present disclosure. FIG. 4 is described in conjunction with elements from FIGs. 1, 2 and 3. With reference to FIG. 4, there is shown an exemplary diagram 400 that depicts architectural diagram of the system in form of graphical chart for detection of the concerning behaviours. The exemplary diagram 400 consists of four graphs and two components. At the top, there are two graphs i.e. the first graph 302B and the third graph 302A. A central part of the diagram comprises two components i.e. a first component 406 and a second component 408. The first component 406 represents detection of concerning behaviours. The detection of concerning behaviours may be performed using "Anomaly Detection" 402A and "SLO-based Service Degradation Detection." 404A. Further, the second component 408 represents the time intervals of concerning behaviours 302D.
[0098] The bottom part of the exemplary diagram 400 displays two graphs, i.e. a seventh graph and an eighth graph. The seventh graph 402B represents the effect of "Anomaly Detection” 402 A. The eighth graph 404B, represents the effect of the "SLO-based Service Degradation Detection. " 404A.
[0099] FIG. 5 is an exemplary diagram depicting a graphical chart of the generation of symptom patterns, in accordance with an embodiment of the present disclosure. FIG. 5 is described in conjunction with elements from FIGs. 1 , 2, 3, and 4. With reference to FIG. 5, there is shown an exemplary diagram 500 that illustrates an architecture for generating symptom patterns using the Al algorithms.
[0100] The top part of the exemplary diagram 500 comprises the second graph 302C and time intervals of concerning behaviours 302D, and both of them are used for generation of symptoms patterns.
[0101] The central part of the exemplary diagram 500 includes "Generation of Symptom patterns”. The generation of symptom patterns is done using either one or more of the three modules the knowledge based 504A, the rule based 504B and the template based 504C. The modules employ different techniques for identifying and generating symptom patterns based on either domain knowledge or predefined templates. In effect the generation of symptom patterns generates definition of identified symptoms 304A.
[0102] The bottom part of the exemplary diagram 500 depicts a ninth graph 506 having two portions. A first portion 506A of the ninth graph 506 and a second portion 506B of the ninth graph 506 shows instances of detected symptom patterns over time. In the ninth graph 506 the x-axis represents the time range, and the y-axis represents the packet rate. The first portion 506A and the second portion 506B represents visualization of a network metric called "send-unicast-packet-rate" over a period from October 28, 2022, to November 1, 2022.
[0103] The first portion 506A highlights time intervals of concerning behaviours. The time intervals of concerning behaviours are identified by a rule-based algorithm that analyses the end-to-end throughput values, which in this context represents an SLI, for a given service and flags intervals where the values exceed certain predetermined thresholds or meet specific conditions. The rule-based algorithm looks for intervals where the end-to-end throughput deviates significantly from the expected or normal range (SLO), either by exceeding an upper limit or dropping below a lower limit. The deviations may indicate potential network issues, such as congestion, overload, or other abnormal conditions. The second portion 506B shows the symptom expression that was identified by the system 102, showing the actual values of the send-unicast-packet-rate metric, which exhibits a cyclical pattern with periodic peaks and valleys. In the second portion 506B, it is reported the threshold assigned by the system 102, which indicates the condition in which the NLM (send-unicast-packet-rate) correlates the intervals in which the SLI (i.e. end-to-end throughput) expresses violations of the SLOs (i.e. end-to-end throughput is greater than desired level).
[0104] By using a rule-based approach, the Al algorithm may automatically identify and highlight these intervals of interest, allowing network administrators or engineers to quickly pinpoint and investigate periods that require further attention or troubleshooting.
[0105] Advantageously, the exemplary diagram 500 represents an architecture that combines different techniques, such as knowledgebased, template-based, and rule-based approaches, to generate symptom patterns from various data sources or signals. These symptom patterns could then be used for monitoring, diagnostics, or anomaly detection purposes in various domains.
[0106] FIG. 6 is an exemplary diagram depicting a graphical chart of the alert definition, in accordance with an embodiment of the present disclosure. FIG. 6 is described in conjunction with elements from FIGs. 1, 2, 3, 4, and 5. With reference to FIG. 6, there is shown an exemplary diagram 600 depicting the process of generation of the set of alerts by the system 102 (FIG. 1). The processor 104 takes the network level metrics represented by the second graph 302C, the time intervals of concerning behaviour 302D, the definition of identified symptoms 304A as inputs. The goal of the processor 104 is to produce the final set of alerts to be deployed by removing low-quality symptoms and grouping high-quality ones, providing aggregated views of the candidates to make it easier for experts to review and generate alerts. The process of obtaining alert definition may be heuristic based 602, human validation 604 based, symptoms aggregation 606 based. In an example, symptom aggregated commonly refer to the same concerning behaviours, allowing the selection of the most appropriate symptoms representing the incident. For example, selecting the top ‘N’ symptoms ranked by a balance of accuracy and understandability. Further network engineer validation may be additionally performed to discard non-interesting symptoms based on domain knowledge.
[0107] The processor 104 further removes low-quality symptoms based on key indicators like accuracy and understandability (e.g., the number of conditions in a rule) and groups high-quality symptoms and provides aggregated views for easier review and alert generation. Flexibility to use different components or algorithms for symptom validation and selection, if the symptoms are solely used for incident detection without a need to constrain the number of monitored symptoms. The goal is to refine and curate the final set of alerts by leveraging both automated techniques and human expertise, ensuring that only the most relevant and high-quality symptoms are selected for deployment and monitoring.
[0108] FIG. 7 is an exemplary diagram depicting a graphical chart of the dynamic alert in accordance with an embodiment of the present disclosure. FIG. 7 is described in conjunction with elements from FIGs. 1, 2, 3, 4, 5 and 6. With reference to FIG. 7, there is shown an exemplary diagram 700 that depicts an architecture for a dynamic alert evaluation that combines monitoring data, impact score updates, and alert refinement to provide intelligent alerting and analysis. The exemplary diagram 700 illustrates a dynamic alert evaluation system that combines monitoring data, impact score updates, and alert refinement to provide intelligent alerting and analysis. The use of the real-time network level metrics represented by the fifth graph 310B, the new time intervals of concerning behaviours 314 which raises alerts, and the final set of deployed alerts 306B helps the processor 104 for the dynamic alert evaluation of the system 102. The dynamic alert evaluation comprises impact score update 702 and alert refinement 704.
[0109] Impact score update 702 is responsible for generating updated impact score for the alert definitions 316 based on the monitored data. The impact score likely quantifies the severity or potential impact of any detected anomalies or issues. The alert refinement 704 takes the impact score and potentially other factors into account to refine and contextualize alert definitions. It likely applies rules, filters, or machine learning models to prioritize, categorize, or enrich the alerts, reducing false positives and providing more meaningful insights. Further, the updated impact score alerts 316 refer to the parameters generated by the processor 104 based on the dynamic alert evaluation. The dynamic alert evaluation is an iterative process. As new monitoring data comes in, the impact score is updated by the processor 104, which then triggers the alert refinement process to re-evaluate and adjust the alerts accordingly. The graph 706 visualizes the refined alerts or their impact scores over time. This may help users quickly identify and understand the most critical issues or events that require attention.
[0110] The exemplary diagram 700 represents an alert management system that dynamically evaluates monitoring data, calculates impact scores, and refines alerts through an iterative process. This approach aims to provide more accurate, actionable, and contextualized alerts, improving incident response and decision-making processes.
[0111] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe, and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A method (200) of determining a set of alerts to be deployed to a runtime network monitoring system, comprising steps of: a) using a plurality of service level indicator metrics which describe a behaviour of at least one service being implemented on a network, to detect at least one time interval during which the at least one service being implemented on the network is demonstrating a degradation of service health; b) using the detected at least one time interval and a plurality of network level metrics to determine a plurality of candidate symptom expressions correlated with the at least one-time interval; c) using the determined plurality of candidate symptom expressions, the at least one-time interval, and the plurality of network level metrics, to determine a set of alerts; and d) activating the determined set of alerts and deploying the determined set of alerts to a runtime network monitoring system.
2. The method (200) of claim 1, further comprising a step of e) using the at least one-time interval and the plurality of network level metrics to monitor the deployed set of alerts.
3. The method (200) of claim 2, wherein the results of the monitoring is used to determine an updated set of alerts with impact scores related to the deployed set of alerts, and wherein the monitoring tracks a relevance, with respect to service degradation events, of the deployed set of alerts over a time period.
4. The method (200) of claim 1 , wherein the plurality of service level indicator metrics include delay, packet drop or flow count.
5. The method of (200) claim 1, wherein the runtime network monitoring system receives service level indicator metrics, compares them against thresholds and generates alerts as a result of the comparing.
6. The method (200) of claim 1, wherein step a) determines the at least one-time interval using automated anomaly detection.
7. The method (200) of claim 6, wherein the automated anomaly detection is heuristic based.
8. The method (200) of claim 6, wherein the automated detection is machine learning based.
9. The method (200) of claim 1 , wherein the step b) uses an artificial intelligence algorithm.
10. The method (200) of claim 9, wherein the artificial intelligence algorithm is rule based, template based, or knowledge based.
11. The method (200) of claim 1 , wherein the step c) uses root cause analysis algorithms.
12. The method (200) of claim 1, wherein the runtime network monitoring system monitors a network which uses intent driven networking.
13. A system (102) comprising means adapted for carrying out all the steps of the method (200) according to any preceding method claim.
14. A computer program comprising instructions for carrying out all the steps of the method (200) according to any preceding method claim, when said computer program is executed on a computer system.
Citation Information
Patent Citations
Threshold selection for KPI candidacy in root cause analysis of network issues
US11616682B2
Using machine learning based on cross-signal correlation for root cause analysis in a network assurance service
US20190356533A1
Automated processes and systems for managing and troubleshooting services in a distributed computing system
US20230108819A1