Automatic detection and remediation of high processor usage at network device
By automatically detecting and remediating high processor usage in network devices through a cloud-based network management system, the inefficiency of manual detection and remediation in existing technologies is solved, thereby improving network performance and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JUNIPER NETWORKS INC
- Filing Date
- 2024-08-23
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, high processor usage in network devices is difficult to detect and remedy effectively, leading to network performance degradation and failures. Manual solutions are time-consuming and impractical.
A cloud-based network management system (NMS) monitors the processor usage of network devices, automatically detects high processor usage, analyzes the root cause, and takes remedial measures, such as terminating abnormal processes or generating recommended remedial measures.
It improves network performance and reliability, reduces the time of network performance degradation caused by high CPU usage, and enables automated fault recovery.
Smart Images

Figure CN121890048A_ABST
Abstract
Description
[0001] This application claims the benefit of U.S. Patent Application No. 18 / 774,745, filed July 16, 2024, and U.S. Provisional Patent Application No. 202341056781, filed August 24, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates generally to computer networks, and more specifically to the monitoring and troubleshooting of computer networks. Background Technology
[0003] Commercial locations or sites (such as offices, hospitals, airports, stadiums, or retail stores) typically install complex wireless network systems (including networks of wireless access points (APs)) throughout the premises to provide wireless network services to one or more wireless client devices (or simply "clients"). An AP is a physical electronic device that enables other devices to wirelessly connect to a wired network using various wireless network protocols and technologies, such as IEEE 802.11-compliant (i.e., "WiFi"), one or more wireless LAN protocols including Bluetooth / Bluetooth Low Energy (BLE), mesh networking protocols such as ZigBee, or other wireless network technologies. Many different types of wireless client devices (such as laptops, smartphones, tablets, wearables, appliances, and Internet of Things (IoT) devices) incorporate wireless communication technologies and can be configured to connect to a compatible wireless access point to access a wired network when the device is within range of the access point. When a client device is running a cloud-based application (such as a Voice over Internet Protocol (VoIP) application, a streaming video application, a gaming application, or a video conferencing application), data is exchanged from the client device to the cloud-based application server during the application session via one or more access points and one or more wired network devices (e.g., switches, routers, and / or gateway devices). Summary of the Invention
[0004] In general, this invention describes techniques for detecting high processor usage (e.g., high central processing unit (CPU) usage) at a network device and for remedying the detected high processor usage at the network device. High processor usage can affect the routing efficiency of a network device. For example, high processor usage can reduce the expected execution of routing system processes by the processor (e.g., by delaying the execution of routing system processes or by not executing routing system processes). When routing system processes are delayed or not executed by the network device's processor while being executed by the network device's processor, the network device, as well as other network devices directly connected to it, can react as if a network problem exists, and this can cause failover or even catastrophic failure of network sites.
[0005] Network devices can include switches, routers, gateways, or other suitable network devices capable of sending and receiving network traffic. A network can include tens of thousands of network devices, and at any given time, hundreds or thousands of these devices may exhibit problems routing network traffic. Consequently, it may be time-consuming or even impractical for network administrators to manually identify which of these devices are experiencing high processor usage and to manually implement remediation measures to address the high processor usage on those devices.
[0006] According to various aspects of this disclosure, a cloud-based network management system (NMS) can monitor processor usage statistics of network devices in a network, including processor usage statistics of processes executing on network devices in the network, to detect high processor usage at one or more network devices. The NMS can use the collected processor usage statistics at each network device exhibiting high processor usage to determine whether the high processor usage is caused by anomalous behavior, such as high processor usage of one or more processes executing on the processor. The NMS can invoke one or more remedial measures to address the anomalous behavior at each of the one or more network devices exhibiting high processor usage caused by anomalous behavior. Such remedial measures can be assigned based on root cause analysis of the processes causing the anomalous behavior. For example, the NMS can automatically terminate one or more processes executing on the processor that are the root cause of the high processor usage. In some embodiments, the NMS can also recommend one or more remedial measures to a network administrator to address the anomalous behavior leading to the high processor usage. In some embodiments, a severity score is also assigned to the remedial measures based on the duration and magnitude of the high processor usage.
[0007] The technology disclosed herein provides one or more technical advantages and practical applications. This technology enables cloud-based NMS to systematically detect high processor usage in network devices within a network, which may be caused by anomalous behavior of the network devices, determine the root cause of the high processor usage, and automatically take action to remedy the anomalous behavior of the network devices. Therefore, this technology can reduce the amount of time that network devices experiencing high CPU usage can degrade network performance, thereby improving network performance and reliability.
[0008] In some aspects, the techniques described herein relate to a network management system comprising: a memory; and one or more processors coupled to the memory and configured to: obtain processor usage statistics for one or more network devices; for a given network device among the one or more network devices, determine aggregated processor usage statistics across a time window based on the processor usage statistics; based on the aggregated overall processor usage of the given network device exceeding a baseline threshold, analyze the aggregated per-process processor usage of the given network device to determine one or more processes as the root cause of the anomalous behavior of the given network device; and generate remedial measures to remedy the root cause.
[0009] In some aspects, the technology described herein relates to a method comprising: obtaining processor usage statistics for one or more network devices by one or more processors of a network management system; determining, by the one or more processors and for a given network device among the one or more network devices, aggregated processor usage statistics across a time window based on the processor usage statistics; analyzing, by the one or more processors, the aggregated per-process processor usage of the given network device based on the aggregated overall processor usage of the given network device exceeding a baseline threshold to determine one or more processes as the root cause of the anomalous behavior of the given network device; and generating remedial measures by the one or more processors to remedy the root cause.
[0010] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium containing instructions that, when executed by one or more processors of a network management system, cause one or more processors to: obtain processor usage statistics for one or more network devices; for a given network device among the one or more network devices, determine aggregated processor usage statistics across a time window based on the processor usage statistics; based on the aggregated overall processor usage of the given network device exceeding a baseline threshold, analyze the aggregated per-process processor usage of the given network device to determine one or more processes as the root cause of the anomalous behavior of the given network device; and generate remedial measures to remedy the root cause.
[0011] Details of one or more embodiments of the technology of the present invention are set forth in the following drawings and description. Other features, objects, and advantages of these technologies will be apparent from the specification, drawings, and claims. Attached Figure Description
[0012] Figure 1A This is a block diagram of an example network system including a network management system, based on one or more technologies according to this disclosure.
[0013] Figure 1B It is shown Figure 1A A block diagram detailing further embodiments of the network system.
[0014] Figure 2 This is a block diagram illustrating a network device according to an embodiment of the technology disclosed herein.
[0015] Figure 3 This is a block diagram of a network management system according to one or more embodiments of the technologies disclosed herein.
[0016] Figure 4 An exemplary graphical user interface is shown, providing a view of processor usage on a network device.
[0017] Figure 5 An exemplary graphical user interface is shown, providing a view of the root cause of a network problem.
[0018] Figure 6 An exemplary graphical user interface shows a view providing details of the root cause of a network problem and recommended actions for remedying the problem.
[0019] Figure 7 An exemplary graphical user interface is shown, providing a view of processor usage at a network device.
[0020] Figure 8 This is a flowchart illustrating example operations performed by the network management system of the embodiment to detect and remedy high processor usage.
[0021] Figure 9 This is a flowchart illustrating exemplary operations performed by an exemplary network management system.
[0022] Throughout the accompanying drawings and description, the same reference numerals refer to the same elements. Detailed Implementation
[0023] Figure 1AThis is a block diagram of an example network system 100 including a network management system (NMS) 130 according to one or more technologies of this disclosure. The example network system 100 includes multiple sites 102A-102N, at which a network service provider manages one or more wireless networks 106A-106N, respectively. Although in Figure 1A In this document, each site 102A-102N is shown as including a single wireless network 106A-106N, but in some examples, each site 102A-102N may include multiple wireless networks, and this disclosure is not limited to this aspect.
[0024] Each site 102A to 102N includes multiple network access server (NAS) devices, such as access points (APs) 142, switches 146, or routers (not shown). For example, site 102A includes multiple APs 142A-1 to 142A-M. Similarly, site 102N includes multiple APs 142N-1 to 142N-M. Each AP 142 can be any type of wireless access point, including but not limited to commercial or enterprise APs, routers, or any other device connected to a wired network and capable of providing wireless network access to client devices within the site.
[0025] Each site 102A-102N also includes multiple client devices, also referred to as User Equipment (UE), representing various wireless-enabled devices within each site, and typically referred to as UE or client device 148. For example, multiple UEs 148A-1 to 148A-K are currently located at site 102A. Similarly, multiple UEs 148N-1 to 148N-K are currently located at site 102N. Each UE 148 can be any type of wireless client device, including but not limited to mobile devices (such as smartphones, tablets, or laptops), personal digital assistants (PDAs), wireless terminals, smartwatches, smart rings, or other wearable devices. UE 148 may also include wired client-side devices, such as IoT devices (such as printers), security devices, environmental sensors, or any other device connected to a wired network and configured to communicate over one or more wireless networks 106.
[0026] To provide wireless network services to UE 148 and / or communicate via wireless network 106, AP 142 and other wired client-side devices at site 102 are directly or indirectly connected to one or more network devices (e.g., switches, routers, etc.) via physical cables (e.g., Ethernet cables). Figure 1AIn one embodiment, site 102A includes switch 146A, and each of APs 142A-1 to 142A-M at site 102A is connected to switch 146A. Similarly, site 102N includes switch 146N, and each of APs 142N-1 to 142N-M at site 102N is connected to switch 146N. Although in Figure 1A The illustration appears to show each site 102 comprising a single switch 146 and all APs 142 of a given site 102 connected to that single switch 146. However, in other embodiments, each site 102 may include more or fewer switches and / or routers. Furthermore, APs and other wired client-side devices at a given site may connect to two or more switches and / or routers. Additionally, two or more switches at a site may be interconnected with each other and / or connected to two or more routers, for example, via a mesh or partial mesh topology in a central branch architecture. In some embodiments, the interconnected switches and routers comprise a wired local area network (LAN) at site 102 hosting the wireless network 106.
[0027] Example network system 100 also includes various networking components for providing networking services within a wired network, including (as an example) an authentication, authorization, and accounting (AAA) server 110 for authenticating users and / or UE 148, a dynamic host configuration protocol (DHCP) server 116 for dynamically assigning network addresses (e.g., IP addresses) to UE 148 during authentication, a domain name system (DNS) server 122 for resolving domain names to network addresses, multiple servers 128A-128X (collectively referred to as "Server 128") (e.g., web servers, database servers, file servers, etc.), and a network management system (NMS) 130. Figure 1A As shown, different devices and systems of network system 100 are coupled together via one or more networks 134 (e.g., the Internet and / or corporate intranets).
[0028] exist Figure 1AIn some embodiments, NMS 130 is a cloud-based computing platform for managing wireless networks 106A-106N at one or more sites 102A-102N. As further described herein, NMS 130 provides an integrated suite of management tools and implements various technologies disclosed herein. Typically, NMS 130 can provide a cloud-based platform for wireless network data acquisition, monitoring, activity logging, reporting, predictive analytics, network anomaly identification, and alarm generation. In some embodiments, NMS 130 outputs notifications (such as warnings, alarms, graphical indicators on dashboards, log messages, text / SMS messages, email messages, etc.) and / or recommendations regarding wireless network issues to site or network administrators (“administrators”) who interact with and / or operate management device 111. Furthermore, in some embodiments, NMS 130 operates in response to configuration input received from administrators who interact with and / or operate management device 111.
[0029] The administrator and management device 111 may include IT personnel and administrator computing devices associated with one or more sites 102. The management device 111 may be implemented as any suitable device for presenting output and / or accepting user input. For example, the management device 111 may include a display. The management device 111 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or an administrator. The management device 111 may, for example, represent a workstation, laptop or notebook computer, desktop computer, tablet computer, or any other computing device that can be operated by a user and / or present a user interface according to one or more aspects of this disclosure. The management device 111 may be physically separate from the NMS 130 and / or located in a different location from the NMS 130, such that the management device 111 can communicate with the NMS 130 via a network 134 or other communication methods.
[0030] In some embodiments, one or more NAS devices (e.g., AP 142, switch 146, or router) may be connected to edge devices 150A-150N via physical cables (e.g., Ethernet cables). Edge device 150 includes a cloud-managed wireless local area network (LAN) controller. Each edge device 150 may include an on-premises device at site 102 that communicates with NMS 130 to extend certain microservices from NMS 130 to the on-premises NAS device while using NMS 130 and its distributed software architecture for scalable and resilient operation, management, troubleshooting, and analysis.
[0031] Each network device of network system 100 (e.g., servers 110, 116, 122 and / or 128, AP 142, UE 148, switch 146, and any other server or device attached to or forming part of network system 100) may include a system log or error log module, wherein each of these network devices records the status of the network device, including normal operating status and error conditions. Throughout this disclosure, one or more of the network devices of network system 100 (e.g., servers 110, 116, 122 and / or 128, AP 142, UE 148, and switch 146) may be considered “third-party” network devices when owned by and / or associated with an entity different from NMS 130, such that NMS 130 does not receive, collect, or otherwise access the recorded status and other data of the third-party network devices. In some embodiments, edge device 150 may provide a proxy through which the recorded status and other data of the third-party network devices may be reported to NMS 130.
[0032] In some embodiments, NMS 130 monitors network data 137 (e.g., one or more Service Level Expectation (SLE) metrics) received from wireless networks 106A-106N at each site 102A-102N, and manages network resources (such as APs 142 at each site) to deliver a high-quality wireless experience to end users, IoT devices, and clients at the sites. For example, NMS 130 may include a Virtual Network Assistant (VNA) 133 that implements an event processing platform for providing real-time insights into IT operations and simplifying troubleshooting, and automatically taking remedial actions or providing recommendations to proactively resolve wireless network issues. VNA 133 may, for example, include an event processing platform configured to handle concurrent streams of hundreds or thousands of network data 137 from sensors and / or agents associated with AP 142 and / or nodes within network 134. For example, VNA 133 of NMS 130 may include underlying analytics and network error detection engines and alerting systems according to different examples described herein. The underlying analytics engine of VNA 133 can apply historical data and models to inbound event streams to calculate assertions, such as the predicted occurrence of identified anomalies or events constituting network error conditions. Furthermore, VNA 133 can provide real-time alerts and reports to notify site or network administrators of any predicted events, anomalies, or trends via management device 111, and can perform root cause analysis and automatic or assisted error remediation. In some embodiments, VNA 133 of NMS 130 can apply machine learning techniques to identify the root causes of error conditions detected or predicted from the flow of network data 137. If the root cause can be resolved automatically, VNA 133 can invoke one or more corrective actions to correct the root cause of the error condition, thereby automatically improving underlying SLE metrics and also automatically improving the user experience.
[0033] Further examples of the operation implemented by the VNA133 of the NMS130 are described in detail in the following U.S. patents: U.S. Patent No. 9,832,082, entitled "Monitoring Wireless Access Point Events," issued November 28, 2017; U.S. Publication No. US 2021 / 0306201, entitled "Network System Fault Resolution Using a Machine Learning Model," published September 30, 2021; U.S. Patent No. 10,985,969, entitled "Systems and Methods for a Virtual Network Assistant," issued April 20, 2021; U.S. Patent No. 10,958,585, entitled "Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault Detection," issued March 23, 2021; and U.S. Patent No. 10,958,585, entitled "Method for Spatio-Temporal...", issued March 23, 2021. U.S. Patent No. 10,958,537, entitled “Modeling”; and U.S. Patent No. 10,862,742, entitled “Method for Conveying AP Error Codes Over BLE Advertisements”, granted on December 8, 2020. Each of these patents and patent applications is incorporated herein by reference in its entirety.
[0034] In operation, NMS 130 observes, collects, and / or receives network data 137, which may take the form of data extracted from messages, counters, and statistics, for example. Depending on one implementation, a computing device is part of NMS 130. Depending on other implementations, NMS 130 may include one or more computing devices, dedicated servers, virtual machines, containers, services, or other forms of environments for performing the techniques described herein. Similarly, computing resources and components implementing VNA 133 may be part of NMS 130, may run on other servers or execution environments, or may be distributed across nodes within network 134 (e.g., routers, switches, controllers, gateways, etc.).
[0035] According to one or more techniques of the present invention, NMS 130 is configured to monitor the processor usage, such as central processing unit (CPU) usage, of each of one or more network devices, such as switch 146, in network 134. NMS 130 is configured to periodically (e.g., every minute, every 3 minutes, etc.) collect processor usage statistics, such as CPU usage statistics, from each of the one or more switches 146. NMS 130 is configured to use the collected processor usage statistics to detect high processor usage (e.g., high CPU usage) at one or more switches 146 and remedy the detected high processor usage at one or more switches 146.
[0036] NMS 130 is configured to collect overall processor usage statistics (e.g., overall CPU usage statistics) and per-process CPU usage statistics (e.g., per-process CPU usage statistics) for each of one or more switches 146. Collecting per-process CPU usage statistics for each of one or more switches 146 enables NMS 130 to determine the root cause of high processor usage on a particular network switch and to determine remedies to address the high processor usage on that particular network switch.
[0037] The operating system of a network switch can track the processor usage of the network switch, which can include the overall processor usage and the per-process processor usage. The operating system can determine processor usage statistics based on this data and can expose the tracked statistics. For example, the operating system can determine overall processor usage statistics, which may be the overall percentage utilization of the processor, and can determine per-process processor usage statistics, which may be the percentage utilization of the processor by each process executing on the processor. An agent executing on the network switch can periodically read the processor usage statistics from the operating system and periodically send these statistics to the NMS 130 for storage in the database 136. The agent can also periodically collect other statistics, such as the amount of network traffic routed through the network switch, and periodically send these collected statistics to the NMS 130 for storage in the database 136.
[0038] NMS 130 is configured to determine, at least in part, whether a network switch is experiencing high CPU usage for each of one or more switches 146 based on processor usage statistics collected by NMS 130. In an example where each of the switches 146 may comprise multiple modules and / or chassis, NMS 130 is configured to determine, at least in part, whether a network switch is experiencing high processor usage for each of the boot modules and / or chassis of the switches 146 based on processor usage statistics collected by NMS 130.
[0039] NMS 130 is configured to determine aggregate processor usage statistics across a time window for each of one or more switches 146. The time window can be the previous 20 minutes, 30 minutes, one hour, etc. The aggregate processor usage statistics across the time window for the network switches can include a count of the number of times the overall processor usage of the network switches exceeded a specified high processor usage threshold (e.g., 90% utilization) within the time window, the average (e.g., mean) overall processor usage of the network switches within the time window, and the average processor usage per process executed on the processor during the time window.
[0040] NMS 130 is configured to, for each of one or more switches 146, determine whether the overall processor usage of the network switch across a time window exceeds a baseline processor usage threshold. In some embodiments, the baseline threshold may be a specified percentage of processor utilization, such as 80% CPU utilization. In some embodiments, the baseline processor usage threshold may be a long-term learning threshold specific to a particular network switch and may be based on tracking the historical processor usage of a particular network switch. NMS 130 is also configured to, for each of one or more switches 146, determine whether the number of times the overall processor usage of the network switch exceeds a specified high processor usage threshold within a time window is greater than a high processor usage frequency threshold, which may be 2, 3, etc.
[0041] The NMS 130 is configured to analyze aggregated processor usage statistics of network switches across a time window to detect anomalous behavior, based on the overall processor usage of these network switches exceeding a baseline threshold. In other words, if the NMS 130 determines that processor usage of network switches across a time window is high, it can determine whether this high processor usage is caused by anomalous behavior. For example, the NMS 130 can be configured to detect anomalous behavior for each network switch that could be the root cause of the high processor usage, provided that the network switch has an overall processor usage exceeding a baseline processor usage threshold across the time window and has an overall processor usage exceeding a specified high processor usage threshold more than a high processor usage frequency threshold within that time window.
[0042] To detect abnormal behavior of a network switch, NMS 130 can determine the total network traffic routed through the network switch during a time window and the per-process processor usage of processes at the network switch across the time window. NMS 130 can be configured to retrieve network traffic statistics collected and stored in database 136 of the network switch and determine the total network traffic routed through the network switch during the time window based on these network traffic statistics.
[0043] To determine per-process processor usage at a network switch across a time window, the NMS 130 is configured to determine the processor usage of each process executing at the processor during the time window. Determining per-process processor usage at the network switch across a time window allows the NMS 130 to identify which processes contribute to high processor usage at the network switch. It also allows the NMS 130 to identify process-level processor usage anomalies based on normal process usage and to use mutual information to determine the frequency of these anomalies in each process.
[0044] NMS 130 can be configured to perform heuristic-based detection of anomalous behavior that is the root cause of high processor usage by a network switch using an anomaly detection model 135. In some embodiments, the anomaly detection model 135 can be trained via machine learning to perform heuristic-based detection of anomalous behavior that is the root cause of high processor usage. The anomalous behavior that is the root cause of high processor usage can be the behavior of the network switch itself, rather than high network traffic routed through the network switch, that is the cause of the high processor usage. Such anomalous behavior may include one or more processes executing at the processor that have high processor usage or whose network switches are not deployed in a recommended manner.
[0045] Anomaly detection model 135 can be a statistical model that analyzes long-term statistics of overall processor usage at the network switch and processor usage of individual processes executing at the processor to detect anomalies at the network switch. Mutual information and anomaly detection can be used to fine-tune anomaly detection model 135 to find commonalities among processes causing problems at the network switch. Anomaly detection model 135 can also be programmed or trained to determine which problems are likely caused by certain processes that consume more processor cycles than normal processes. Anomaly detection model 135 can therefore be able to determine which anomalies among the detected anomalies are true locations and / or false alarms for certain problems, and can determine which remedial measures can be performed for certain problems.
[0046] NMS 130 can input processor usage statistics and / or network traffic statistics of the network switch into anomaly detection model 135, and anomaly detection model 135 can determine and output an indication of whether the anomalous behavior is the root cause of high processor usage and / or one or more features most relevant to the anomalous behavior based on the input data. In some embodiments, NMS 130 can input features including processor usage statistics and / or network traffic statistics of the network switch into anomaly detection model 135, and output an indication of whether the anomalous behavior is the root cause of high processor usage and one or more features most relevant to the anomalous behavior. Processor usage statistics can include aggregated processor usage statistics across time windows, such as a count of the number of times the overall processor usage of the network switch exceeds a specified high processor usage threshold within a time window, the average overall processor usage of the network switch within the time window, and / or the average processor usage per process executed at the processor during the time window. Processor usage statistics can also include per-process processor usage of processes at the network switch across time windows. Network traffic statistics of the network switch can include the total network traffic routed via the network switch during the time window.
[0047] Anomaly detection model 135 can output an indication of whether the anomalous behavior is the root cause of high processor usage and / or one or more features most relevant to the anomalous behavior. For example, anomaly detection model 135 can determine an anomaly score (which can be between 0 and 1) based on the input features, which can correspond to the probability that the high processor usage of the network switch is caused by the anomalous behavior detected by anomaly detection model 135. If the anomaly score is higher than an anomaly score threshold, such as 0.6 in the example where the anomaly score is between 0 and 1, then NMS 130 can determine that the high processor usage of the network switch is caused by the anomalous behavior detected by anomaly detection model 135.
[0048] If NMS 130 determines that the high processor usage of the network switch is caused by abnormal behavior detected by the anomaly detection model 135, NMS 130 can be configured to store the determined processor usage statistics of the network switch and the determined network traffic statistics in database 136. NMS 130 can use these processor usage statistics and network traffic statistics of the network switch in future time windows to more accurately detect abnormal behavior of the network switch, and / or use exponential averaging to correlate the determined processor usage statistics and network traffic statistics of the network switch with those in future time windows.
[0049] If NMS 130 determines that the high processor usage of the network switch is caused by anomalous behavior detected by the anomaly detection model, then the anomaly detection model 135 may also output one or more features identified as most relevant to the anomalous behavior detected at the network switch. For example, the anomaly detection model 135 may be able to detect whether the network switch has been deployed in the recommended manner. A network switch not deployed in the recommended manner (e.g., by using uncertified optical connectors or other physical components) may result in suboptimal use of the network switch, and this could be the cause of the high processor usage of the network switch. Accordingly, if the anomaly detection model 135 detects that the network switch is not deployed in the recommended manner, then the anomaly detection model 135 may output an indication that the anomalous behavior is caused by the network switch not being deployed in the recommended manner.
[0050] In some embodiments, NMS 130 is configured to use an anomaly detection model 135 to determine one or more processes as the root cause of anomalous behavior of the network switch. In some embodiments, the anomaly detection model 135 is capable of detecting that high processor usage of the network switch is caused by one or more processes executing on the processor of the network switch, and in response, can output an indication of one or more processes executing on the processor as the root cause of the anomalous behavior.
[0051] Processes utilizing the processor of a network switch can include user-space processes and system-space processes. Accordingly, the anomaly detection model 135 can output indications of one or more user-space processes and / or one or more system-space processes that are the root cause of the anomalous behavior.
[0052] In some embodiments, user-space processes that do not appear to have high processor usage can still be the root cause of high processor usage by system-space processes, for example, if the user-space process causes a large number of system-space processes to start and execute on the processor. Anomaly detection model 135 can be programmed and / or trained to associate user-space processes of the network switch with system-space processes that the user-space process can cause to start, and thus may be able to detect and output an indication that a user-space process is the root cause of high processor usage by the network switch, even if the user-space process does not have high processor usage.
[0053] NMS 130 can generate remedial measures in response to determining that one or more processes are the root cause of anomalous behavior in the network switch. In some embodiments, NMS 130 can be configured to automatically invoke one or more remedial measures to address the root cause of the anomalous behavior. For example, if the anomaly detection model detects that a user-space process executing on the network switch's processor is the root cause of high processor usage in the network switch, NMS 130 can be configured to automatically terminate or restart the user-space process to resolve the high processor usage.
[0054] In some embodiments, NMS 130 can be configured to generate and, for example, output notifications to a network administrator of the WAN, suggesting one or more remedial actions to address the anomalous behavior. For example, NMS 130 can generate data representing a user interface for display on a user interface device operated, for example, by a network administrator of the enterprise network, which presents recommendations for performing one or more remedial actions. In some embodiments, NMS 130 can output instructions for remedial actions and recommended actions in the form of a chatbot searchable by a user, such as a system administrator of the WAN.
[0055] For example, if NMS 130 determines that the anomalous behavior is caused by the network switch not being deployed in a recommended manner (e.g., by using uncertified optical connectors or other physical components), NMS 130 can be configured to generate and output recommended remedies to use certified optical connectors or physical components. In another embodiment, if the anomaly detection model detects that a system space process executing on the network switch's processor is the root cause of high processor usage on the network switch, NMS 130 can be configured to generate and output recommended remedies to terminate or restart the system space process. In some embodiments, if such a system space process is whitelisted for NMS 130 to terminate or restart, NMS 130 can be configured to automatically terminate or restart the system space process to resolve the high processor usage.
[0056] Although the techniques of this invention are described in relation to detecting high CPU usage of a network switch, the techniques described herein can be similarly applied to detecting high memory usage and / or high temperatures of a network switch. Furthermore, although the techniques of this invention are described in relation to a network switch, the techniques described herein can be similarly applied to routers, access points, and any other suitable network devices in a network.
[0057] The technology disclosed herein provides one or more technical advantages and practical applications. This technology enables a cloud-based NMS 130 to systematically detect high processor usage of switch 146 in the network that may be caused by abnormal behavior of switch 146, and to take action to remedy the abnormal behavior of the network switch.
[0058] Furthermore, NMS 130 can provide user visibility into the anomalous behavior of network switch 146 in the network. For example, NMS 130 can generate data representing a user interface for display on a user interface device operated, for example, by a network administrator of an enterprise network. The user interface can present indications of the anomalous behavior of switch 146 as the root cause of high processor usage. NMS 130 can further generate and, for example, output notifications to the network administrator of the enterprise network, suggesting the implementation of one or more remedial measures to address the root cause of the high processor usage of switch 146. In other embodiments, NMS 130 can instead automatically invoke one or more remedial measures to address the anomalous behavior, such as automatically terminating one or more processes at switch 146 that are causing the high processor usage.
[0059] While the technology of the present invention is described in this embodiment as being performed by NMS 130, the technology described herein can be performed by any other computing device, system, and / or server, and the invention is not limited thereto. For example, one or more computing devices configured to perform the functions of the technology of the present invention may reside in a dedicated server, or be contained in any other server besides NMS 130, or may be distributed throughout network 100, and may or may not be part of NMS 130.
[0060] Figure 1B It is shown Figure 1A A block diagram detailing further embodiments of the network system. In this embodiment, Figure 1B NMS 130 is shown, and NMS 130 is configured to provide services from a "client" (e.g., user device 148 connected to wireless network 106 and wired LAN 175). Figure 1B (The far left) crosses over to the “cloud” (e.g., cloud-based application services 181 that can be hosted by computing resources within data center 179) Figure 1B The far right) is operated by an AI / machine learning-based computing platform that provides full automation, insights, and assurance (WiFi assurance, wired assurance, and WAN assurance).
[0061] As described herein, the NMS 130 provides an integrated suite of management tools and implements various technologies disclosed herein. Typically, the NMS 130 can provide a cloud-based platform for wireless network data acquisition, monitoring, activity logging, reporting, predictive analytics, network anomaly detection, and alarm generation. For example, the network management system 130 can be configured to proactively monitor and adaptively configure the network system 100 to provide self-driving capabilities. Furthermore, the VNA 133 includes a natural language processing engine for providing AI-driven support and troubleshooting, anomaly detection, AI-driven location services, and AI-driven radio frequency (RF) optimization with reinforcement learning.
[0062] like Figure 1B As shown in the example, the AI-driven NMS 130 also provides configuration management, monitoring, and automated supervision of a software-defined wide area network (SD-WAN) 177, which operates as an intermediate network to communicatively couple wireless network 106 and wired LAN 175 to data center 179 and application services 181. Typically, SD-WAN 177 provides a seamless, secure, traffic-designed connection between “branch” routers 187A of the wired network 175 hosting wireless network 106 (e.g., a hub or campus network) and a “central” router 187B further up the cloud stack of cloud-based application services 181. SD-WAN 177 typically operates and manages overlay networks on top of the underlying physical wide area network (WAN), providing connectivity to geographically separated customer networks. In other words, SD-WAN 177 extends software-defined networking (SDN) capabilities to the WAN and allows the network to decouple the underlying physical network infrastructure from virtualized network infrastructure and applications, enabling the network to be configured and managed in a flexible and scalable manner.
[0063] In some embodiments, the underlying routers of the SD-WAN 177 can implement a stateful, session-based routing scheme, where routers 187A and 187B dynamically modify the contents of the original packet headers initiated by client device 148 to direct traffic to application service 181 along a selected path (e.g., path 189) without using tunnels and / or additional labels. In this way, routers 187A and 187B can be more efficient and scalable for large networks because the use of tunnelless, session-based routing allows routers 187A and 187B to achieve significant network resource utilization by avoiding the need to perform encapsulation and decapsulation at tunnel endpoints. Furthermore, in some examples, each router 187A and 187B can independently perform path selection and traffic engineering to control the packet flow associated with each session without using a centralized SDN controller for path selection and label distribution. In some embodiments, routers 187A and 187B implement session-based routing as Secure Vector Routing (SVR) provided by Juniper Networks.
[0064] Additional information regarding session-based routing and SVRs is described in the following patents: U.S. Patent No. 9,729,439, published August 8, 2017, entitled “COMPUTER NETWORK PACKET FLOW CONTROLLER”; U.S. Patent No. 9,729,682, published August 8, 2017, entitled “NETWORK DEVICE AND METHOD FOR PROCESSINGA SESSION USING A PACKET SIGNATURE”; U.S. Patent No. 9,762,485, published September 12, 2017, entitled “NETWORK PACKET FLOW CONTROLLER WITH EXTENDED SESSIONMANAGEMENT”; and U.S. Patent No. 9,762,485, published January 16, 2018, entitled “ROUTER WITH OPTIMIZED STATISTICAL”. U.S. Patent No. 9,871,748, entitled “FUNCTIONALITY”; U.S. Patent No. 9,985,883, published May 29, 2018, entitled “NAME-BASED ROUTING SYSTEM AND METHOD”; U.S. Patent No. 10,200,264, published February 5, 2019, entitled “LINK STATUS MONITORING BASED ON PACKETLOSS DETECTION”; U.S. Patent No. 10,277,506, published April 30, 2019, entitled “STATEFUL LOAD BALANCING IN A STATELESS NETWORK”; and U.S. Patent No. 10,277,506, published October 1, 2019, entitled “NETWORK PACKET FLOW CONTROLLER WITH EXTENDEDSESSION”. The entire contents of each of the following patents are incorporated herein by reference: U.S. Patent No. 10,432,522 entitled “MANAGEMENT”; and U.S. Patent Application Publication No. 11,075,824 entitled “IN-LINE PERFORMANCE MONITORING”, published on July 27, 2021.
[0065] In some embodiments, the AI-driven NMS 130 can enable intent-based configuration and management of the network system 100, including enabling the construction, presentation, and execution of intent-driven workflows for configuring and managing devices associated with the wireless network 106, the wired LAN network 175, and / or the SD-WAN 177. For example, declarative requirements express the desired configuration of network components without specifying precise local device configurations and control flows. By utilizing declarative requirements, what should be achieved is specified, rather than how it should be achieved. Declarative requirements can contrast with the necessary instructions that describe the exact device configuration syntax and control flow required to achieve the configuration. By utilizing declarative requirements instead of imperative instructions, users and / or user systems are relieved of the burden of determining the exact device configurations required to achieve the desired results for the user / system. For example, when utilizing various types of devices from different vendors, specifying and managing the exact necessary instructions for configuring each device in the network is often difficult and cumbersome. As new devices are added and devices fail, the types and kinds of devices in the network can change dynamically. Managing various types of devices from different vendors to configure a cohesive network with different configuration protocols, syntaxes, and software versions is often difficult to achieve. Thus, by requiring only the user / system to specify declarative requirements that define the expected results applicable across a wide variety of device types, the management and configuration of network devices become more efficient. Further illustrative details and technical descriptions of intent-based network management systems are found in U.S. Patent No. 10,756,983, entitled "Intent-based Analytics," and U.S. Patent No. 10,992,543, entitled "Automatically generating anintent-based network model of an existing computer network," each of which is incorporated herein by reference.
[0066] According to the technology described in this disclosure, NMS 130 includes a virtual network assistant 133 configured to monitor CPU usage of network devices. The CPU usage agent can periodically collect CPU usage statistics from the network devices to determine high network usage at one or more of the network devices. NMS 130 may also include an anomaly detection model 135 configured to determine whether such high CPU usage is caused by anomalous behavior. NMS 130 can therefore determine and / or automatically execute remedial measures to improve such anomalous behavior.
[0067] Figure 2 This is a block diagram illustrating a network device 200 according to an embodiment of the technology disclosed herein. Generally, the network device 200 may be... Figure 1B One of the routers 187A and 187B, and Figure 1A One of the switches 146, or supporting Figure 1B An embodiment of one or more of a wireless network 106, a wired LAN 175 or SD-WAN 177, and a data center 179, and another network device (e.g., router 187). In this embodiment, network device 200 includes interface cards 226A to 226N (“IFC 226”) that receive packets via incoming links 228A to 228N (“Incoming Links 228”) and transmit packets via outgoing links 230A to 230N (“Outgoing Links 230”). IFC 226 is typically coupled to links 228, 230 via multiple interface ports. Network device 200 also includes a control unit 202 that determines the route of received packets and forwards packets accordingly via IFC 226.
[0068] Control unit 202 may include one or more processors 203, a routing engine 204, and a packet forwarding engine 222. Processor 203 may implement functions within network device 200 and / or execute instructions to perform functions of network device 200. For example, processor(s) 203 execute software instructions, such as software instructions for defining software or computer programs, which are stored in a computer-readable storage medium, such as a non-transitory computer-readable medium including storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, storing instructions for causing one or more processors 203 to perform the techniques described herein. Figure 2 In one embodiment, processor 203 may be referred to as the CPU of network device 200.
[0069] The routing engine 204 operates as a control plane for the network device 200 and includes an operating system that provides a multitasking operating environment for the execution of multiple parallel processes. The routing engine 204 interacts with other routers (e.g., such as...) Figure 1A The switch 146 communicates with the computer network (such as the switch 146) to establish and maintain the computer network (such as the switch 146) Figures 1A to 1BThe network system 100 is used to transmit network traffic between one or more client devices. The routing protocol daemon (RPD) 208 of the routing engine 204 executes software instructions to implement one or more control plane networking protocols 212. For example, protocol 212 may include one or more routing protocols, such as Internet Group Management Protocol (IGMP) 221 and / or Border Gateway Protocol (BGP) 220, for exchanging routing information with other routing devices and for updating the Routing Information Base (RIB) 206, Multiprotocol Label Switching (MPLS) protocol 214, and other routing protocols. Protocol 212 may also include one or more communication session protocols 223, such as TCP, UDP, TLS, or ICMP. Protocol 212 may also include one or more performance monitoring protocols, such as BFD 225.
[0070] RIB 206 can describe the topology of the computer network in which network device 200 resides, and may also include routes through a shared tree in the computer network. RIB 206 describes various routes within the computer network and the appropriate next hop for each route, i.e., along the adjacent routing devices of each route. Routing engine 204 analyzes the information stored in RIB 206 and generates forwarding information for forwarding engine 222 stored in Forwarding Information Base (FIB) 224. FIB 224 can associate, for example, a network destination with a specific next hop and the corresponding IFC 226 and the physical output port for output link 230. FIB 224 can be a radix tree programmed into a dedicated forwarding chip, a series of tables, a complex database, a linked list, a radix tree, a database, a flat file, or various other data structures.
[0071] FIB 224 may also include a lookup structure. Given a key (such as an address), the lookup structure can provide one or more values. In some embodiments, the one or more values can be one or more next hops. A next hop can be implemented as microcode that performs one or more operations upon execution. The one or more next hops can be "chained," such that a set of chained next hops performs a set of operations for the corresponding different next hops upon execution. Examples of such operations may include applying one or more services to a packet, dropping a packet, and / or forwarding a packet using an interface and / or an interface identified by one or more next hops.
[0072] Session information 235 stores information used to identify sessions. In some embodiments, session information 235 takes the form of a session table. For example, service information 232 includes one or more entries specifying a session identifier. In some embodiments, the session identifier includes one or more of a source address, source port, destination address, destination port, or protocol associated with the forward and / or reverse flow of the session. As described above, when routing engine 204 receives data from a client device (e.g., ... Figure 1A The source device 112A) and the destination is another client device (e.g., Figure 1A When a packet in a packet stream is forwarded by the destination device 114, the routing engine 204 determines whether the packet belongs to a new session (e.g., whether it is the "first" or "leading" packet of the session). To determine whether a packet belongs to a new session, the routing engine 204 determines whether the session information 235 includes an entry corresponding to the source address, source port, destination address, destination port, and protocol of the first packet. If an entry exists, the session is not a new session. If no entry exists, the session is new, and the routing engine 204 generates a session identifier for the session and stores the session identifier in the session information 235. The routing engine 204 can then use the session identifier stored in the session information 235 to identify subsequent packets as belonging to the same session.
[0073] Service information 232 is stored by routing engine 204 to identify services associated with a session. In some embodiments, service information 232 is in the form of a service table. For example, service information 232 includes one or more entries specifying a service identifier and one or more of a source address, source port, destination address, destination port, or protocol associated with the service. In some embodiments, routing engine 204 may query service information 232 for a received packet using one or more of the session's source address, source port, destination address, destination port, or protocol to determine the service associated with the session. For example, routing engine 204 may determine the service identifier based on the correspondence between the source address, source port, destination address, destination port, or protocol in service information 232 and the source address, source port, destination address, destination port, or protocol specified by the session identifier. Routing engine 204 retrieves one or more service policies 234 corresponding to the identified service based on the service associated with the packet. Service policies may include, for example, path failover policies, Dynamic Host Configuration Protocol (DHCP) tagging policies, traffic engineering policies, and priority of network traffic associated with the session. The routing engine 204 applies one or more service policies 234 corresponding to the services associated with the group to the group.
[0074] In some embodiments, network device 200 may include a session-based router employing a stateful, session-based routing scheme that enables routing engine 204 to independently perform path selection and traffic engineering. Using session-based routing allows network device 200 to avoid using a centralized controller (such as an SDN controller) for path selection and traffic engineering, and to avoid tunneling. In some embodiments, network device 200 may implement session-based routing as a Secure Vector Router (SVR) provided by Juniper Networks. Where network device 200 includes a session-based router that serves as a network gateway for a site operating as an enterprise network, network device 200 may establish multiple peering paths over the underlying physical WAN with one or more other session-based routers that serve as network gateways for other sites operating as part of the enterprise network.
[0075] While primarily described herein as a session-based router, in other embodiments, network device 200 may include a network switch or a packet-based router, wherein routing engine 204 employs packet-based or flow-based routing schemes to forward packets according to network paths defined, for example, by a centralized controller performing path selection and traffic engineering. Where network device 200 includes a packet-based router operating as a network gateway for sites within an enterprise network, network device 200 may establish multiple tunnels over the underlying physical WAN with one or more other packet-based routers operating as network gateways for other sites within the enterprise network.
[0076] According to the technology disclosed herein, the processor usage agent 238 of the control unit 202 is configured to collect processor usage statistics of the network device 200 (e.g., usage statistics of the processor 203). The processor usage agent 238 can collect overall processor usage statistics of the processor 203 and per-process processor usage statistics of the processor 203, and send the collected processor usage statistics to the NMS 130. The processor 203 can be configured to execute processes that may include user-space processes and system-space processes, and the processor usage agent 238 can collect per-process processor usage statistics for both user-space processes and system-space processes.
[0077] The operating system of network device 200 can track the processor usage of processor 203, which may include the overall processor usage of processor 203 and the per-process processor usage of processor 203. The operating system of network device 200 can determine processor usage statistics for processor 203 based on the processor usage of processor 203, and can expose the tracked processor usage statistics for processor 203. For example, the operating system of network device 200 can determine the overall processor usage statistics, which may be the overall percentage utilization of processor 203, and can determine the per-process processor usage statistics, which may be the percentage utilization of processor 203 by that process for each process executing at processor 203. A processor usage agent 238 executing at network device 200 can periodically read processor usage statistics from the operating system and can periodically send the processor usage statistics of processor 203 to NMS 130. In some embodiments, the processor usage agent 238 can also periodically collect other statistics, such as the amount of network traffic routed via network device 200, and can periodically send such collected statistics to NMS 130.
[0078] Figure 3 This is a block diagram of a network management system (NMS) 300 according to one or more embodiments of the present disclosure. The NMS 300 can be used to implement, for example... Figures 1A to 1B The NMS 130 is used in this embodiment. In such an embodiment, the NMS 300 is responsible for monitoring and managing one or more wireless networks 106A-106N at sites 102A-102N.
[0079] The NMS 300 includes a communication interface 330, one or more processors 306, a user interface device 310, a memory 312, and a database 318. The various components are coupled together via a bus 314, through which they can exchange data and information. In some embodiments, the NMS 300 receives data from client devices 148, APs 142, switches 146, and other network nodes within the network 134 (e.g., ...). Figure 1B The NMS 300 receives data from one or more routers (187), which can be used to calculate one or more SLE metrics and / or update network data 316 in database 318. The NMS 300 analyzes this data for cloud-based management of wireless networks 106A-106N. In some embodiments, the NMS 300 may be... Figure 1A This refers to a portion of another server or any other server.
[0080] Processor 306 executes software instructions, such as software instructions for defining software or computer programs, which are stored in a computer-readable storage medium (such as memory 312), such as a non-transitory computer-readable medium including storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, the stored instructions causing one or more processors 306 to perform the techniques described herein.
[0081] The communication interface 330 may include, for example, an Ethernet interface. The communication interface 330 couples the NMS 300 to a network and / or the Internet, such as... Figure 1A This refers to any network 134 and / or any local area network shown. Communication interface 330 includes a receiver 332 and a transmitter 334, through which the NMS 300 communicates with / from client devices 148, AP 142, switch 146, servers 110, 116, 122, 128, and / or forms networks such as... Figure 1A Any other network node, device, or system within the network system 100 shown herein may receive / transmit data and information. In some scenarios described herein, where the network system 100 includes “third-party” network devices owned and / or associated with an entity different from the NMS 300, the NMS 300 may not receive, collect, or otherwise access network data from the third-party network devices.
[0082] The data and information received by the NMS 300 may include, for example, telemetry data, SLE-related data, or data from client devices AP148, AP142, switch 146, or other network nodes used by the NMS 300 for remotely monitoring the performance of the wireless network 106A-106N and application sessions from client devices to cloud-based application servers (e.g., Figure 1BThe NMS 300 receives event data from one or more routers (187). The data and information received by the NMS 300 may also include processor usage statistics collected by the switch 146, and the NMS 300 may store the collected processor usage statistics as processor usage data 317 in a database 318. The processor usage statistics may include both overall processor usage statistics for each of the one or more switches 146 and per-process processor usage statistics for each of the one or more switches 146. The overall processor usage statistics for the network switches may be the total percentage of processor utilization of the network devices, while the per-process processor usage statistics for the network switches may be the percentage of processor utilization by each process executing at the processor of the network switch. The NMS 300 uses the processor usage statistics to determine for each of the one or more switches 146 whether the network switch is experiencing high processor usage and whether abnormal behavior of the network switch is caused by high processor usage. The NMS 300 can also send data via communication interface 330 to any network device, such as client device 148, AP 142, switch 146, other network nodes within network 134, and management device 111, to remotely manage wireless networks 106A-106N and portions of wired networks.
[0083] Memory 312 includes one or more means configured to store programming modules and / or data associated with the operation of NMS 300. For example, memory 312 may include a computer-readable storage medium, such as a non-transitory computer-readable medium, containing storage devices (e.g., disk drives or optical drives) or memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory, the storage of which causes one or more processors 306 to execute instructions of the techniques described herein.
[0084] In this embodiment, memory 312 includes API 320, SLE module 322, Virtual Network Assistant (VNA) / AI engine 350, and Radio Resource Management (RRM) engine 360. According to the disclosed technology, VNA / AI engine 350 includes processor usage engine 352. NMS 300 may also include any other programming modules, software engines, and / or interfaces configured for remote monitoring and management of portions of wireless networks 106A-106N and wired networks, including AP 142 / 200, switch 146, or other network devices (e.g., Figure 1B Remote monitoring and management of any of the routers (187).
[0085] SLE module 322 enables the establishment and tracking of thresholds for SLE metrics for each network 106A-106N. SLE module 322 also analyzes SLE-related data collected by APs (such as any of AP 142) from UEs in each wireless network 106A-106N. For example, APs 142A-1 to 142A-N collect SLE-related data from UEs 148A-1 to 148A-N currently connected to wireless network 106A. This data is sent to NMS 300, which is executed by SLE module 322 to determine one or more SLE metrics for each UE 148A-1 to 148A-N currently connected to wireless network 106A. In addition to any network data collected by one or more APs 142A-1 to 142A-N in wireless network 106A, this data is also sent to NMS 300 and stored as network data 316 in, for example, database 318.
[0086] RRM Engine 360 monitors one or more metrics at locations 102A-102N to learn and optimize the RF environment at each location. For example, RRM Engine 360 can monitor coverage and capacity SLE metrics for wireless network 106 at site 102 to identify potential SLE coverage and / or capacity issues in wireless network 106 and adjust the radio settings of access points at each site to address the identified issues. For example, the RRM Engine can determine the channel and transmit power distribution among all APs 142 in each network 106A-106N. For example, RRM Engine 360 can monitor events, power, channel, bandwidth, and number of clients connected to each AP. RRM Engine 360 can also automatically change or update the configuration of one or more APs 142 at site 102 to improve coverage and capacity SLE metrics and thus provide users with an improved wireless experience.
[0087] The VNA / AI engine 350 analyzes data received from network devices as well as its own data to identify when an unexpected anomalous state is encountered at one of the network devices. For example, the VNA / AI engine 350 can identify the root cause of any unexpected or anomalous state, such as any poor SLE metric indicating connectivity problems at one or more network devices. Furthermore, the VNA / AI engine 350 can automatically invoke one or more corrective actions designed to resolve the identified root cause of one or more poor SLE metrics. Examples of corrective actions that can be automatically invoked by the VNA / AI engine 350 may include, but are not limited to, invoking the RRM engine 360 to reboot one or more APs, adjusting / modifying the transmission power of a specific radio in a specific AP, adding an SSID configuration to a specific AP, changing the channel on an AP or a group of APs, etc. Corrective actions may also include restarting switches and / or routers, invoking the download of new software to APs, switches, or routers, etc. These corrective actions are given for illustrative purposes only, and this disclosure is not limited to this aspect. If automatic remediation is unavailable or fails to adequately address the root cause, the VNA / AI engine 350 can proactively provide notifications, including recommended remediation measures to be taken by IT personnel (e.g., site or network administrators using management device 111), to resolve the network error.
[0088] VNA / AI engine 350 analyzes processor usage data 317 (including processor usage statistics received from switch 146) and its own data to detect high processor usage at one or more switches 146 and detect anomalous behavior of one or more switches 146 caused by high processor usage. For example, VNA / AI engine 350 can use processor usage engine 352 to determine whether anomalous behavior at a network device is caused by high CPU usage at the network device, and to determine the root cause of the anomalous behavior. In some embodiments, processor usage engine 352 utilizes artificial intelligence-based techniques to help determine whether anomalous behavior at a network device is caused by high processor usage at the network device. Furthermore, VNA / AI engine 350 can automatically invoke one or more corrective actions designed to resolve anomalous behavior at a network device caused by high processor usage. Examples of corrective actions that can be automatically invoked by VNA / AI engine 350 may include, but are not limited to, calling API 320 to terminate one or more processes identified as the root cause of the anomalous behavior at the network device. Corrective actions may also include restarting one or more network devices, invoking the download of new software to the network device, switch, or router, etc. These corrective actions are given for illustrative purposes only, and this disclosure is not limited to this. If automated remedies are unavailable or insufficient to address the root cause, the VNA / AI engine 350 may proactively provide a notification that includes recommended remedial actions to be taken by IT personnel to resolve the network error.
[0089] In some embodiments, the VNA / AI engine 350 may use supervised and / or unsupervised training to construct, train, apply, and retrain anomaly detection model 356 to determine whether a network switch's anomalous behavior is caused by high processor usage. The VNA / AI engine 350 may then apply the anomaly detection model 356 to data streams and / or logs of newly collected data (e.g., processor usage data 317) from the switch 146 to detect whether currently observed anomalous behavior of the network switch is caused by high processor usage. When applying the anomaly detection model 356 to the processor usage data 317 indicates that the anomalous behavior of the network device is caused by high processor usage, the VNA / AI engine 350 may invoke the processor usage engine 352 to trigger automatic or semi-automatic correction actions.
[0090] In some embodiments, the anomaly detection model 356 may include a supervised ML model trained using training data comprising pre-collected labeled network data received from network devices (e.g., client devices, APs, switches, and / or other network nodes) to identify anomalous behavior of network switches. The supervised ML model may include one of logistics regression, primitive Bayes, support vector machines (SVM), etc. In other embodiments, the anomaly detection model 356 may include an unsupervised ML model. Although not explicitly stated... Figure 3 As shown, in some embodiments, database 318 may store training data, and VNA / AI engine 350 or dedicated training module may be configured to train anomaly detection model 356 based on the training data to determine appropriate weights for one or more features across the training data.
[0091] According to the technology of the present invention, the processor usage engine 352 can monitor one or more network devices (e.g., by monitoring processor usage data 317 collected by the NMS 300 for each of the one or more switches 146). Figure 1A The processor usage engine 352 analyzes the processor usage of each of the switches 146 in the network system 100 to determine whether a network switch is experiencing high processor usage for each of the one or more switches 146. In embodiments where each switch 146 may comprise multiple modules and / or chassis, the processor usage engine 352 can determine whether a network switch is experiencing high processor usage for each of the lead modules and / or chassis of the switches 146. Although the technique is described with respect to switch 146, the technique is equally applicable to any other type of network device.
[0092] Processor usage engine 352 can determine aggregated processor usage statistics across a time window for each of one or more switches 146. The time window can be the previous 20 minutes, the previous 30 minutes, the previous hour, etc. The overall processor usage statistics for the network switches across the time window can include a count of the number of times the overall processor usage of the network switches exceeded a specified high processor usage threshold (e.g., 90% utilization) within the time window, the average (e.g., average) overall processor usage of the network switches within the time window, and the average processor usage per process executed on the processor during the time window.
[0093] Processor usage engine 352 may determine for each of one or more switches 146 whether the overall processor usage of network devices across a time window exceeds a baseline processor usage threshold. In some embodiments, the baseline threshold may be a specified percentage of processor utilization, such as 80% processor utilization. In some embodiments, the baseline processor usage threshold may be a long-term learning threshold specific to a particular network switch and may be based on tracking the historical processor usage of a particular network switch. In some embodiments, the baseline threshold may be determined based on statistics from a large sample set of network switches and may be a general baseline threshold for network devices. Processor usage engine 352 may also determine for each of one or more switches 146 whether the technique for determining the number of times the overall processor usage of network devices exceeds a specified high processor usage threshold within a time window is greater than a high processor usage frequency threshold, which may be 2, 3, etc.
[0094] The processor usage engine 352 can analyze aggregated processor usage statistics of network switches across a time window to detect anomalous behavior based on the overall processor usage of network switches exceeding a baseline threshold across a time window. That is, if the processor usage engine 352 determines that the processor usage of network switches across a time window is high, it can determine whether such high processor usage is caused by anomalous behavior. For example, the processor usage engine 352 can detect anomalous behavior for each network switch whose overall processor usage exceeds a baseline processor usage threshold across a time window and whose overall processor usage exceeds a specified high processor usage threshold within the time window by a count greater than a high processor usage frequency threshold; this anomalous behavior could be the root cause of the high processor usage.
[0095] To detect abnormal behavior of network switches, processor usage engine 352 can determine the total network traffic routed through the network switch during a time window and the per-process processor usage of processes at the network switch across the time window. Processor usage engine 352 can be configured to retrieve network traffic statistics of the network switch collected and stored in database 318, and determine the total network traffic routed through the network switch during the time window based on these statistics.
[0096] To determine per-process processor usage across a time window at a network switch, the processor usage engine 352 can determine the processor usage of processes executing at the processor during the time window and the processor usage of each process among those processes. Determining per-process processor usage across a time window at the network switch allows the NMS 300 to identify which processes contribute to high processor usage at the network switch. Furthermore, determining per-process processor usage across a time window at the network switch also allows the NMS 300 to identify process-level processor usage anomalies based on normal process usage and to use mutual information to determine the frequency of these anomalies in each process.
[0097] The processor usage engine 352 can use an anomaly detection model 356 to perform heuristic detection of anomalous behaviors that are the root cause of high processor usage in a network device. In some embodiments, the anomaly detection model 356 can be trained via machine learning to perform heuristic detection of anomalous behaviors that are the root cause of high processor usage. An anomalous behavior that is the root cause of high processor usage could be the behavior of a network switch, rather than high network traffic routed via the network switch, that is causing the high processor usage. Such anomalous behavior could include one or more processes executing at the processor that have high processor usage or whose network switches are not deployed in a recommended manner.
[0098] Anomaly detection model 356 may be a statistical model that analyzes long-term statistics of overall processor usage of the network switch and processor usage of individual processes executing at the processor to detect anomalies at the network switch. In some embodiments, anomaly detection model 356 may also track overall processor (e.g., CPU) usage and / or processor usage across individual processes in multiple deployments to determine the overall distribution of processor usage and firmware-related commonalities in anomalous behavior. Anomaly detection model 356 may be fine-tuned using mutual information and anomaly detection to find commonalities among processes causing problems at network switch 146. Anomaly detection model 356 may also be programmed or trained to determine which problems are likely caused by certain processes that consume more processor cycles than normal. Anomaly detection model 356 may therefore be able to determine which anomalies among the detected anomalies are true and / or false positives for certain problems and to determine remedies that can be performed for specific problems.
[0099] The processor usage engine 352 can input processor usage statistics and / or network traffic statistics of the network switch into the anomaly detection model 356. These statistics can be stored as processor usage data 317. The anomaly detection model 356 can determine and output an indication of whether the abnormal behavior of the network switch is the root cause of high processor usage and / or an indication of one or more features most relevant to the abnormal behavior based on the input data. In some embodiments, the processor usage engine 352 can input features including processor usage statistics and / or network traffic statistics of the network switch into the anomaly detection model 356, and the anomaly detection model 356 can output an indication of whether the abnormal behavior is the root cause of high processor usage and one or more features most relevant to the abnormal behavior.
[0100] Processor usage statistics input into the anomaly detection model 356 may include aggregated processor usage statistics across time windows, such as the count of the number of times the overall processor usage of the network switch exceeds a specified high processor usage threshold within the time window, the average overall processor usage of the network switch within the time window, and / or the average processor usage per process executing at the processor during the time window. Processor usage statistics may also include per-process processor usage at the network switch across the time window. Network traffic statistics for the network switch may include the total network traffic routed through the network switch during the time window.
[0101] Anomaly detection model 356 can output an indication of whether the anomalous behavior of the network switch is the root cause of high processor usage and / or one or more features most relevant to that anomalous behavior. For example, anomaly detection model 356 can determine an anomaly score between 0 and 1 based on the input features, which can correspond to the probability that the high processor usage of the network switch is caused by the anomalous behavior detected by the anomaly detection model. If the anomaly score is higher than an anomaly score threshold (e.g., 0.6 in an instance where the anomaly score is between 0 and 1), then processor usage engine 352 can determine that the high processor usage of the network switch is caused by the anomalous behavior detected by anomaly detection model 356.
[0102] If the processor usage engine 352 determines that the high processor usage of the network switch is caused by abnormal behavior detected by the anomaly detection model 356, the processor usage engine 352 can be configured to store the determined processor usage statistics and the determined network traffic statistics of the network switch in the database 318. The NMS 300 and the processor usage engine 352 can use these processor usage statistics and network traffic statistics of the network switch in future time windows to more accurately detect abnormal behavior of the network switch, and / or use exponential averaging to correlate the determined processor usage statistics and network traffic statistics of the network switch with those in future time windows.
[0103] If the processor usage engine 352 determines that the high processor usage of the network switch is caused by anomalous behavior detected by the anomaly detection model 356, the anomaly detection model 356 may also output one or more features identified as most relevant to the anomalous behavior detected at the network switch. For example, the anomaly detection model 356 may be able to detect whether the network switch has been deployed in the recommended manner. A network switch not deployed in the recommended manner (e.g., by using uncertified optical connectors or other physical components) may result in suboptimal use of the network switch and could be the cause of high processor usage. Therefore, if the anomaly detection model 356 detects that the network switch is not deployed in the recommended manner, the anomaly detection model 356 may output an indication that the anomalous behavior is caused by the network switch not being deployed in the recommended manner.
[0104] In some embodiments, the processor usage engine 352 may use the anomaly detection model 356 to identify one or more processes as the root cause of anomalous behavior of the network switch. In some embodiments, the anomaly detection model 356 may be able to detect that high processor usage of the network switch is caused by one or more processes executing on the processor of the network switch, and in response, may output an indication of one or more processes executing on the processor as the root cause of the anomalous behavior. Processes utilizing the processor of the network switch may include user-space processes and system-space processes. Accordingly, the anomaly detection model 356 may output an indication of one or more user-space processes and / or one or more system-space processes as the root cause of the anomalous behavior.
[0105] In some embodiments, user-space processes that do not appear to have high processor usage can still be the root cause of high processor usage by system-space processes, for example, if the user-space process causes a large number of system-space processes to start and execute on the processor. Anomaly detection model 356 can be programmed and / or trained to associate user-space processes of the network switch with system-space processes that the user-space process can cause to start, and thus may be able to detect and output an indication that a user-space process is the root cause of high processor usage by the network switch, even if the user-space process does not have high processor usage.
[0106] The processor usage engine 352 can generate remedial measures in response to determining high processor usage caused by one or more processes as the root cause of anomalous behavior of the network switch. In some embodiments, the processor usage engine 352 can automatically invoke one or more remedial measures to resolve the root cause of the anomalous behavior. For example, if the anomaly detection model 356 detects that a user-space process executing on the processor of the network switch is the root cause of high processor usage of the network switch, the processor usage engine 352 can automatically terminate or restart the user-space process to resolve the high processor usage.
[0107] In some embodiments, processor usage engine 352 may generate and output notifications to a network administrator of the WAN, recommending one or more remedial actions to address the anomalous behavior. For example, processor usage engine 352 may generate data representing a user interface for display on a user interface device operated by, for example, a network administrator of the enterprise network, presenting recommendations for performing one or more remedial actions. In some embodiments, processor usage engine 352 may output instructions for remedial actions and recommended actions in the form of a user-searchable chatbot, such as a system administrator of the WAN.
[0108] For example, if the processor usage engine 352 determines that the anomalous behavior is caused by the network switch not being deployed in a recommended manner (e.g., by using uncertified optical connectors or other physical components), the NMS 300 can generate and output recommended remedies to use certified optical connectors or physical components. In another embodiment, if the anomaly detection model 356 detects that a system space process executing at the network switch's processor is the root cause of high processor usage on the network switch, the NMS 300 can generate and output recommended remedies to terminate or restart the system space process. In some embodiments, if such a system space process is whitelisted for termination or restart by the NMS 300, the NMS 300 can automatically terminate or restart the system space process to resolve the high processor usage.
[0109] Examples of techniques performed by NMS 300, processor usage engine 352, and anomaly detection model 356 (as described herein) are presented below: 1. In the current batch, read the total CPU usage and CPU usage per process data from the cloud for each switch module.
[0110] 2. Aggregated CPU status for each module within a 20-minute time window: a. CPU count > threshold 1 (90%) = counter1; b. Average CPU usage per module = avg_cpu; and c. Average CPU usage per process per module = avg_process_cpu.
[0111] 3. If counter1 > 2 (2 high CPU points) and avg_cpu > threshold 2 (80% or the baseline obtained through long-term learning), then: d. Calculate the total network traffic routed via the switch within the current aggregation window; and e. Calculate the average CPU usage for each process corresponding to the high CPU point.
[0112] 4. Input the features from steps (2) and (3) into the anomaly detection model.
[0113] 5. The output of the anomaly detection model includes anomaly scores (between 0 and 1) for high CPU events and the most relevant features.
[0114] 6. If the anomaly score is greater than the threshold of 3 (0.6), then: f. Save counter1, avg_cpu, and avg_process_cpu to the cloud database; and g. Generate measures for user intervention or to automatically remedy the root cause by terminating the root cause process from the cloud.
[0115] 7. In the next batch, perform steps (1)-(6) and use the exponential average to correlate the results with the characteristics of the previous batch.
[0116] While the technology of the present invention is described in this embodiment as being performed by NMS 130, the technology described herein can be performed by any other computing device, system, and / or server, and the invention is not limited thereto. For example, one or more computing devices configured to perform the functions of the technology of the present invention may reside in a dedicated server, or be contained in any other server besides NMS 130, or may be distributed throughout network 100, and may or may not be part of NMS 130.
[0117] Figure 4 An exemplary graphical user interface is shown, providing a view of processor usage in a network device. (Referring to Figures 1 to...) Figure 3 Described Figure 4 A network management system (such as network management system 300) can generate data representing a graphical user interface for display on a user interface device (e.g., user interface device 310), which can be operated by the network administrator of the enterprise network.
[0118] like Figure 4 As shown, the NMS 300 can output a graphical user interface (GUI) 400 for display at, for example, a user interface device 310, which presents a visualization of the processor usage 402 and memory usage 404 of network devices (e.g., one of the switches 146) in the network system 100. The GUI 400 also presents notifications 406 that can be output by the NMS 300, such as notifications that the processor usage of a network device is higher than a predetermined maximum processor usage threshold.
[0119] Figure 5 An exemplary graphical user interface is shown, providing a view of the root cause of network problems. (Referring to Figures 1 to...) Figure 3 Described Figure 5 A network management system (such as network management system 300) can generate data representing a graphical user interface for display on a user interface device (e.g., user interface device 310), which can be operated by the network administrator of the enterprise network.
[0120] like Figure 5As shown, the NMS 300 can output a graphical user interface (GUI) 500 for display, for example, at a user interface device 310, to present indications of network problems (e.g., problems in network system 100) and indications of the root causes of network problems. For example, the GUI 500 can present a view of the root causes of problems with network switch 502, such as lost VLANs, bad cables, negotiation mismatches, loop detection, port flipping, port jamming, etc. Figure 5 (not shown in the image) or high processor usage ( Figure 5 (Not shown in the image).
[0121] Figure 6 An exemplary graphical user interface is shown, providing a view of the root cause of network problems. (Referring to Figures 1 to...) Figure 3 describe Figure 6 A network management system (such as network management system 300) can generate data representing a graphical user interface for display on a user interface device (e.g., user interface device 310), which can be operated by the network administrator of the enterprise network.
[0122] like Figure 6 As shown, the NMS 300 can output a graphical user interface (GUI) 600 for display, for example, at a user interface device 310, to present indications of network problems (e.g., problems in network system 100) and indications of the root causes of the network problems. For example, the GUI 600 can present a view of the root causes of a problem with network switch 602, and one of the root causes of the problem with network switch 602 is high CPU usage. The GUI 600 can present recommended actions 604 for remedying the high CPU usage of the network switch. For example, recommended action 604 can indicate that the CPU usage of the network switch is 95%.
[0123] If the user selects the option provided in GUI 600 to view more details about the network switch's high CPU usage, the NMS 300 can output GUI 606, which presents additional details about the network switch's high CPU usage. For example, GUI 606 can indicate that one or more processes executing on the network switch's CPU are consuming a large number of CPU cycles, and it can also indicate that the high CPU usage of one of the processes executing on the CPU is related to the use of an uncertified optical connector on the switch.
[0124] Figure 7 An exemplary graphical user interface is shown, providing a view of processor usage at a network device. (Referring to Figures 1 to...) Figure 3 describe Figure 7A network management system (such as network management system 300) can generate data representing a graphical user interface for display on a user interface device (e.g., user interface device 310), which can be operated by the network administrator of the enterprise network.
[0125] like Figure 7 As shown, the NMS 300 can output a graphical user interface (GUI) 700 for display at, for example, a user interface device 310, to present indications of network problems (e.g., problems in network system 100) and indications of the root causes of the network problems. For example, the GUI 700 can present a view of processes executing on one or more processors of the network switch during a time window and the processor utilization of each process. The GUI 700 can present indications of one or more processes having high utilization on one or more processors of the network switch. For example, the GUI 700 can indicate that process "sh" has an average processor utilization of 6% during the time window, which may be high compared to the historical processor utilization of process "sh". Therefore, the GUI 700 can instruct the user to remedy the high processor utilization of the network switch by terminating process "sh".
[0126] Figure 8 This is a flowchart illustrating example operations performed by a sample network management system to detect and remediate high processor usage. About Figure 3 Described Figure 8 .
[0127] like Figure 8 As shown, the processor usage engine 352 of the NMS 300 can access processor (e.g., CPU) usage statistics of network devices (e.g., one or more switches 146), and for each of the multiple network devices, can determine the overall processor usage statistics of the network device and the per-process processor usage statistics of the network device (802). For example, the processor usage engine 352 can access such processor usage statistics stored in processor usage data 317.
[0128] For each of the multiple network devices, the processor utilization engine 352 can aggregate processor utilization statistics over a time window (e.g., a 20-minute time window) based on the overall processor utilization statistics of the network device and the per-process processor utilization statistics of the network device (804). The processor utilization engine 352 can determine a count (referred to herein as counter 1) of the number of times the processor utilization of the network device exceeds a specified threshold (e.g., 90% of the total processor utilization) within the time window. The processor utilization engine 352 can also determine the average processor utilization of the network device within the time window (referred to herein as avg_cpu) and the average per-process processor utilization of the network device within the time window (referred to herein as avg_process_cpu).
[0129] Processor usage engine 352 can determine whether the count of the number of times the processor utilization of a network device exceeds a specified threshold within a time window is greater than a specified number (such as 2), and whether the average processor utilization of the network device within the time window is greater than a specified threshold (such as 80% or the baseline percentage of long-term learning) (806). If processor usage engine 352 determines that the count of the number of times the processor utilization of a network device exceeds the specified threshold within the time window is not greater than the specified number, or that the average processor utilization of the network device within the time window is not greater than the specified threshold ("No" at 806), then processor usage engine 352 can determine that no high processor utilization was detected at the network device during the time window (807). If the processor utilization engine 352 determines that the count of the number of times the processor utilization of the network device is higher than a specified threshold within a time window is greater than a specified number and the average processor utilization of the network device within the time window is greater than the specified threshold ("Yes" at 806), the processor utilization engine 352 may calculate the total network traffic routed by the network device during the time window (808), and may also calculate the average per-process processor utilization consistent with each time the processor utilization of the network device within the time window is higher than the specified threshold (810).
[0130] Therefore, the processor usage engine 352 can input the features determined above into the anomaly detection model 356. For example, the processor usage engine 352 can input the values of counter 1, avg_cpu, avg_process_cpu, the total network traffic routed by the network device during the time window, and the average per-process processor usage consistent with each time the processor utilization of the network device exceeds a specified threshold within the time window into the anomaly detection model 356 (812). The anomaly detection model 356 can output an anomaly score (which can be a score between 0 and 1) and the most relevant features associated with the determined anomaly score based on the input information (814).
[0131] Processor usage engine 352 can determine whether the anomaly score is greater than an anomaly score threshold (e.g., 0.6) (816). If processor usage engine 352 determines that the anomaly score is not greater than the anomaly score threshold ("No" at 816), then processor usage engine 352 can determine that the high processor usage of the network device is not caused by anomalous behavior (817). If processor usage engine 352 determines that the anomaly score is greater than the anomaly score threshold ("Yes" at 816), then processor usage engine 352 can determine that the high processor usage of the network device is caused by anomalous behavior and can save the calculated values of counter 1, avg_cpu, and avg_process_cpu. Processor usage engine 352 can also generate measures for user intervention or to automatically remedy the anomaly by terminating the root cause process (818).
[0132] The processor usage engine 352 can repeat the process for additional batches of processor usage statistics (e.g., processor usage statistics for other time windows) and can correlate the results of such processes with characteristics of previous batches of processor usage statistics, such as via exponential averaging (820).
[0133] Figure 9 This is a flowchart illustrating exemplary operations performed by an exemplary network management system. Figure 9 It is about Figure 3 The network management system is described using the 300 standard. For example... Figure 9 As shown, one or more processors 306 of the network management system 300 can obtain processor usage statistics (902) for one or more network devices (e.g., switch 146). The processor usage statistics for each network device include overall processor usage statistics and per-process processor usage statistics.
[0134] Processor 306 may determine aggregated processor usage statistics across a time window for a given network device among one or more network devices based on processor usage statistics (904). To aggregate processor usage statistics across a time window for a given network device among one or more network devices based on processor usage statistics, processor(s) 306 may determine a count for the given network device of the total processor usage exceeding a specified high processor usage threshold within the time window. For a given network device, processor 306 may also determine the average total processor usage of the given network device within the time window. To aggregate processor usage statistics across a time window for each of the one or more network devices, processor(s) 306 may determine the average processor usage of each process executing at the given network device within the time window.
[0135] In some examples, processor 306 may determine that the count of times the overall processor usage of a given network device exceeds a specified high processor usage threshold within a time window is greater than a high processor usage frequency threshold, and may, in response to determining that the count of times the overall processor usage of the given network device exceeds the specified high processor usage threshold within the time window is greater than the high processor usage frequency threshold, determine that the aggregated overall processor usage of the given network device exceeds the baseline threshold.
[0136] In some embodiments, in order to analyze the aggregated per-process processor usage of a given network device, one or more processors 306 may determine the total network traffic routed through the given network device during a time window, and based on the total network traffic routed through the given network device during the time window and the average processor usage of each process executing at the given network device within the time window, identify one or more processes as the root cause of the anomalous behavior of the given network device.
[0137] Processor 306 can analyze the aggregated per-process processor usage of a given network device based on the aggregated overall processor usage exceeding a baseline threshold to identify one or more processes as the root cause of the anomalous behavior of the given network device (906). To identify the one or more processes as the root cause of the anomalous behavior of the given network device, processor(s) 306 can input the total network traffic routed via the given network device during the time window and the average processor usage of each process executing at the given network device within the time window into anomaly detection model 356 to identify the one or more processes as the root cause of the anomalous behavior of the given network device.
[0138] In some embodiments, an anomaly detection model 356 is trained using machine learning to perform heuristic detection of anomalous behavior that is the root cause of high processor usage by a network device. In some embodiments, the anomaly detection model 356 outputs an anomaly score. The processor 306 may determine that the anomaly score output by the anomaly detection model 356 is greater than an anomaly score threshold, and in response to determining that the anomaly score is greater than the anomaly score threshold, may determine that the high processor usage of a given network device is caused by anomalous behavior of the given network device.
[0139] Processor 306 may generate remedial measures to remedy the root cause (908). To generate remedial measures, processor(s) 306 may automatically terminate one or more processes identified as the root cause of the anomalous behavior of a given network device. The one or more processes include user-space processes. The one or more processes may also include system-space processes that have been whitelisted for automatic termination.
[0140] The techniques described in this invention can be implemented at least in part in hardware, software, firmware, or any combination thereof. For example, various aspects of the described techniques can be implemented within one or more processors, including one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuits, and any combination of these components. The terms "processor" or "processing circuitry" can generally refer to any of the aforementioned logic circuits, alone or in combination with other logic circuits, or any other equivalent circuitry. A control unit, including hardware, can also perform one or more of the techniques of this invention.
[0141] Such hardware, software, and firmware may be implemented within the same device or in separate devices to support the different operations and functions described in this invention. Furthermore, any of the described units, modules, or components may be implemented together or individually as discrete but interoperable logical devices. Describing different features as modules or units is intended to highlight different functional aspects and does not necessarily imply that these modules or units must be implemented by separate hardware or software components. Rather, the functionality associated with one or more modules or units may be performed by separate hardware or software components, or integrated within common or separate hardware or software components.
[0142] The techniques described in this invention can also be implemented or encoded in a computer-readable medium containing instructions (e.g., a computer-readable storage medium). Instructions embedded or encoded in a computer-readable storage medium can cause a programmable processor or other processor to perform the methods, for example, when the instructions are executed. The computer-readable storage medium may include random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, hard disk, CD-ROM, floppy disk, magnetic tape cassette, magnetic media, optical media, or other computer-readable media.
[0143] Various embodiments have been described. These and other embodiments are within the scope of the following claims.
Claims
1. A network management system, comprising: Memory; and One or more processors are coupled to the memory and configured to: Obtain processor usage statistics for one or more network devices; For a given network device among the one or more network devices, aggregated processor usage statistics across time windows are determined based on the processor usage statistics; Based on the fact that the aggregated overall processor usage of the given network device exceeds a baseline threshold, analyze the per-process processor usage of the aggregated network device to identify one or more processes as the root cause of the anomalous behavior of the given network device; and Generate remedies to address the root cause.
2. The network management system according to claim 1, wherein, In order to determine the aggregated processor usage statistics for the given network device, the one or more processors are configured to determine, for the given network device, a count of the number of times the overall processor usage of the given network device exceeds a specified high processor usage threshold within the time window.
3. The network management system according to claim 2, wherein, The one or more processors are further configured to: The count of the number of times the overall processor usage of the given network device exceeds the specified high processor usage threshold within the time window is greater than the high processor usage frequency threshold; and Based on the fact that the count of the number of times the overall processor usage of the given network device exceeds the specified high processor usage threshold within the time window is greater than the high processor usage frequency threshold, it is determined that the aggregated overall processor usage of the given network device exceeds the baseline threshold.
4. The network management system according to any one of claims 1 to 3, wherein, In order to determine the aggregated processor usage statistics for the given network device, the one or more processors are configured to determine, for the given network device, at least one of the average overall processor usage of the given network device within the time window and the average processor usage of each process executing at the given network device within the time window.
5. The network management system according to claim 4, wherein, To analyze the per-process processor usage of the aggregation of the given network device, the one or more processors are further configured to: Determine the total network traffic routed via the given network device during the time window; and Based on the total network traffic routed via the given network device during the time window and the average processor usage of each process executed at the given network device within the time window, the one or more processes are determined as the root cause of the anomalous behavior of the given network device.
6. The network management system according to claim 5, wherein, To determine that the one or more processes are the root cause of the anomalous behavior of the given network device, the one or more processors are further configured to input the total network traffic routed via the given network device during the time window and the average processor usage of each process executed at the given network device within the time window into an anomaly detection model to determine that the one or more processes are the root cause of the anomalous behavior of the given network device.
7. The network management system according to claim 6, wherein, The anomaly detection model is trained via machine learning to perform heuristic detection of anomalous behavior that is the root cause of high processor usage in network devices.
8. The network management system according to any one of claims 6 and 7, wherein, The anomaly detection model outputs anomaly scores, and the one or more processors are further configured to: It is determined that the anomaly score output by the anomaly detection model is greater than the anomaly score threshold; and Based on the determination that the anomaly score is greater than the anomaly score threshold, it is determined that the high processor usage of the given network device is caused by the anomalous behavior of the given network device.
9. The network management system according to any one of claims 1 to 8, wherein, In order to generate the remedy, the one or more processors are configured to automatically terminate the one or more processes that are determined to be the root cause of the anomalous behavior of the given network device.
10. The network management system according to claim 9, wherein, The one or more processes include system space processes that have been whitelisted for automatic termination.
11. A method comprising: The network management system obtains processor usage statistics for one or more network devices from one or more processors. The one or more processors determine aggregated processor usage statistics across a time window based on the processor usage statistics for a given network device among the one or more network devices; Based on the fact that the aggregated overall processor usage of the given network device exceeds a baseline threshold, the aggregated per-process processor usage of the given network device is analyzed by one or more processors to determine one or more processes as the root cause of the anomalous behavior of the given network device; and The one or more processors generate remedies to remedy the root cause.
12. The method according to claim 11, wherein, Determining the aggregated processor usage statistics for the given network device further includes: determining, by the one or more processors, a count of the number of times the overall processor usage of the given network device exceeds a specified high processor usage threshold within the time window.
13. The method of claim 12, further comprising: The count of the number of times the overall processor usage of the given network device exceeds the specified high processor usage threshold within the time window, as determined by the one or more processors, is greater than the high processor usage frequency threshold; and The one or more processors determine that the aggregated overall processor usage of the given network device exceeds the baseline threshold based on a count greater than the specified high processor usage threshold for the number of times the overall processor usage of the given network device exceeds the specified high processor usage threshold within the time window.
14. The method according to any one of claims 11 to 13, wherein, Determining the aggregated processor usage statistics across the time window based on the processor usage statistics also includes: The one or more processors determine, for the given network device, at least one of the average overall processor usage of the given network device within the time window and the average processor usage of each process executing on the given network device within the time window.
15. The method according to claim 14, wherein, Analyzing the per-process processor usage of the aggregation of the given network device further includes: The total network traffic routed via the given network device during the time window is determined by the one or more processors; and The one or more processors determine the one or more processes as the root cause of the anomalous behavior of the given network device based on the total network traffic routed via the given network device during the time window and the average processor usage of each process executed at the given network device within the time window.
16. The method according to claim 15, wherein, Determining the one or more processes as the root cause of the anomalous behavior of the given network device further includes having the one or more processors input the total network traffic routed via the given network device during the time window and the average processor usage of each process executed at the given network device within the time window into an anomaly detection model to determine the one or more processes as the root cause of the anomalous behavior of the given network device.
17. The method according to claim 16, wherein, The anomaly detection model is trained via machine learning to perform heuristic detection of anomalous behavior that is the root cause of high processor usage in network devices.
18. The method according to any one of claims 16 and 17, wherein, The anomaly detection model outputs an anomaly score, and the method further includes: The one or more processors determine that the anomaly score output by the anomaly detection model is greater than an anomaly score threshold; and The one or more processors determine, based on the determination that the anomaly score is greater than the anomaly score threshold, that the high processor usage of the given network device is caused by the anomalous behavior of the given network device.
19. The method according to any one of claims 11 to 18, wherein, Generating the remedy also includes having the one or more processors automatically terminate the one or more processes that are determined to be the root cause of the anomalous behavior of the given network device.
20. A non-transitory computer-readable storage medium, comprising instructions that, when executed by one or more processors of a network management system, cause the one or more processors to: Obtain processor usage statistics for one or more network devices; For a given network device among the one or more network devices, aggregated processor usage statistics across time windows are determined based on the processor usage statistics; Based on the fact that the aggregated overall processor usage of the given network device exceeds a baseline threshold, analyze the per-process processor usage of the aggregated network device to identify one or more processes as the root cause of the anomalous behavior of the given network device; and Generate remedies to address the root cause.
Citation Information
Patent Citations
Link status monitoring based on packet loss detection
US10200264B2
Stateful load balancing in a stateless network
US10277506B2
Network packet flow controller with extended session management
US10432522B2
Intent-based analytics
US10756983B2
Method for conveying AP error codes over BLE advertisements
US10862742B2