Network anomaly detection

By constructing a multi-factor anomaly detection model, identifying seasonal patterns using time-series network data, and selecting an appropriate anomaly detection model, the efficiency and resource optimization issues of anomaly detection in complex wireless network systems are solved, achieving efficient and accurate anomaly detection and root cause analysis.

CN121907665APending Publication Date: 2026-04-21JUNIPER NETWORKS INC
View PDF 17 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JUNIPER NETWORKS INC
Filing Date
2025-09-17
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing network management systems struggle to efficiently detect and identify anomalies in complex wireless network systems, especially across multiple sites. In particular, when the root cause of the anomaly involves organizational-level issues, existing technologies often fail to effectively optimize resources and provide rapid responses.

Method used

By constructing a multi-factor anomaly detection model, seasonal patterns are identified using time-series network data. Appropriate anomaly detection models (such as threshold, baseline, or deep learning models) are selected to predict feature values. Anomalies are detected based on the difference between the predicted and actual feature values. Combined with root cause analysis, the utilization of computing resources is optimized.

Benefits of technology

It enables efficient detection and rapid response to network anomalies, reduces the consumption of computing resources, and improves the accuracy of anomaly detection and the ability to identify organizational-level problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907665A_ABST
    Figure CN121907665A_ABST
Patent Text Reader

Abstract

The invention relates to network anomaly detection. Techniques are described that identify, by a network management system (NMS), a seasonal pattern of the number of devices collected at a site over time; predicting, by the NMS, a number of devices for the time window based on the seasonal pattern and the number of devices determined for one or more previous time windows; detecting, by the NMS, an anomaly during the time window based on a difference between the actual number of devices determined for the time window and the predicted number of devices for the time window; and determining, by the NMS, a root cause of the anomaly at the site.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Utility Application No. 19 / 253,455, filed June 27, 2025, which claims the benefit of U.S. Provisional Application No. 63 / 709,023, filed October 18, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to computer networks, and more specifically to the monitoring and troubleshooting of computer networks. Background Technology

[0004] In commercial settings or sites such as offices, hospitals, airports, stadiums, or retail stores, complex wireless network systems (including networks of wireless access points (APs)) are typically installed throughout the premises to provide wireless network services to one or more wireless client devices (or simply "clients"). An AP is a physical electronic device that enables other devices to wirelessly connect to a wired network using various wireless network protocols and technologies, such as one or more wireless LAN protocols conforming to the IEEE 802.11 standard (i.e., "WiFi"), Bluetooth / Bluetooth Low Energy (BLE), mesh networking protocols such as ZigBee, or other wireless network technologies. Many different types of wireless client devices (such as laptops, smartphones, tablets, wearables, appliances, and Internet of Things (IoT) devices) combine wireless communication technologies and can be configured to connect to a compatible wireless access point to access a wired network when these devices are within range of the access point. When a client device is running a cloud-based application (such as a VoIP application, streaming video application, gaming application, or video conferencing application), data is exchanged from the client device to the cloud-based application server during the application session via one or more access points and one or more wired network devices (e.g., switches, routers, and / or gateway devices). Summary of the Invention

[0005] Typically, this disclosure describes one or more techniques for detecting network anomalies at a site using a multi-factor anomaly detection model. A network management system can detect anomalies at a site based on time-series network data that indicates at least two feature values ​​associated with corresponding characteristics of a network site. For example, the network management system can collect time-series network data with feature values ​​indicating a first feature: the number of access point (AP) devices identified as active AP devices reporting statistics at the site throughout various time windows. The network management system can also collect time-series network data with feature values ​​indicating a second feature: the number of client devices connected to AP devices at the site throughout various time windows.

[0006] A network management system can identify seasonal patterns associated with at least two features corresponding to collected time-series network data. The system can identify seasonal patterns as functions of at least two features (e.g., the y-axis) of the network data with respect to time (e.g., the x-axis), indicating the statistical behavior of the at least two features within a given time window. For example, the system can graphically represent time-series network data collected within a time window (e.g., a day, a week, etc.) to identify seasonal patterns indicating: regular patterns of feature indicators, where the feature values ​​of at least two features exhibit consistent and stable statistical behavior within the time window; complex patterns of feature indicators, where the feature values ​​of at least two features exhibit consistent and stable statistical behavior within the time window; or random patterns of feature indicators, where the feature values ​​of at least two features exhibit inconsistent statistical behavior within the time window.

[0007] In some examples, the network management system can select different styles of multi-feature anomaly detection models (e.g., threshold anomaly detection models, baseline anomaly detection models, general anomaly detection models applicable to more than two sites, and / or fine-tuned anomaly detection models applicable to specific sites) based on identified seasonal patterns to optimize the computational resources used during anomaly detection. For example, the network management system can maintain a first-style anomaly detection model as a threshold model that can apply heuristics or rules to predict feature values ​​based on expected value ranges. The network management system can maintain a second-style anomaly detection model as a baseline model that can apply statistical averaging and logistic regression to predict feature values. Compared to deep learning models, baseline models can consume fewer computational resources (e.g., processing cycles, memory utilization, power utilization, etc.). The network management system can maintain a third-style anomaly detection model as a fine-tuned machine learning model. In this example, and in instances where the identified seasonal patterns indicate stable and relatively regular patterns, the network management system can select a second-style anomaly detection model to apply during anomaly detection to avoid consuming the additional computational resources associated with implementing a third-style anomaly detection model. In instances where the identified seasonal patterns indicate stable and relatively complex patterns, the network management system can select a third-style anomaly detection model to effectively detect potential anomalies.

[0008] A network management system can apply a selected anomaly detection model to predict the feature values ​​of features associated with time-series network data. For example, the network management system can receive and output the anomaly detection model results as predicted data or pseudo-data indicating the predicted feature values ​​of at least two features within a time window of collected feature values. The network management system can implement the anomaly detection model to generate predictions for at least two features within a time window, indicating the expected feature values ​​of at least two features within the time window.

[0009] A network management system can detect anomalies at sites associated with network data comprising at least two features, based on predictions of feature values ​​output by a selected anomaly detection model. For example, the network management system can detect anomalies at a site by comparing actual feature values ​​within network data collected during a time window with predicted feature values ​​output by the selected anomaly detection model. The network management system can trigger the determination of the root cause of the detected anomalies. The network management system can determine the root cause of the detected anomalies by performing root cause analysis based on at least two features of the collected network data. For example, the network management system can determine whether the root cause of an anomalous client disconnection at a site is due to a problem associated with an individual site (e.g., the AP device at the site) or a higher-level problem (e.g., a switch, edge device, or other organizational issue).

[0010] In some examples, based on the network management system's identification of multiple sites within an organization as associated with a specific anomaly, the network management system can determine whether the root cause of the detected anomaly is an organization-level root cause that may require adjustments to network devices at the affected sites. The network management system can then generate recommendations for adjustments to mitigate potential organization-level root causes that may have contributed to the anomaly detected at multiple sites within the organization.

[0011] In one example, this disclosure relates to a network management system (NMS) including a memory and processing circuitry communicating with the memory, and the processing circuitry being configured to: identify seasonal patterns in the number of devices collected at a site over time; predict the number of devices for a time window based on the seasonal patterns and the number of devices determined for one or more previous time windows; detect anomalies during the time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices for the time window; and determine the root cause of the anomalies at the site.

[0012] In another example, this disclosure relates to a method comprising: identifying by an NMS a seasonal pattern in the number of devices collected at a site over time; predicting by the NMS the number of devices in a time window based on the seasonal pattern and the number of devices determined for one or more previous time windows; detecting by the NMS anomalies during the time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices in the time window; and determining by the NMS the root cause of the anomalies at the site.

[0013] In another example, this disclosure relates to a computer-readable storage medium including instructions that, when executed by one or more programmable processors, cause the one or more programmable processors to: identify a seasonal pattern in the number of devices collected at a site over time; predict the number of devices for a time window based on the seasonal pattern and the number of devices determined for one or more previous time windows; detect anomalies during the time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices for the time window; and determine the root cause of the anomalies at the site.

[0014] In one example, this disclosure relates to an NMS including a memory and processing circuitry communicating with the memory, and the processing circuitry being configured to: monitor real-time seasonal patterns of multiple features of network data collected at one of multiple sites associated with an organization during training; assign the site to one of two or more pattern types based on the real-time seasonal patterns at the site, wherein the two or more pattern types include a random pattern type, a regular pattern type, and a complex pattern type; assign one of two or more anomaly detection models to the site based on the pattern type of the site, wherein the anomaly detection model is associated with the pattern type of the site; and use the anomaly detection model assigned to the site to detect anomalies in multiple features of the network data collected at the site.

[0015] In another example, this disclosure relates to a method comprising: monitoring, during training, real-time seasonal patterns of multiple features of network data collected at one of multiple sites associated with an organization by an NMS; assigning the site to one of two or more pattern types by the NMS based on the real-time seasonal patterns at the site, wherein the two or more pattern types include a random pattern type, a regular pattern type, and a complex pattern type; assigning one of two or more anomaly detection models to the site by the NMS based on the pattern type of the site, wherein the anomaly detection model is associated with the pattern type of the site; and using the anomaly detection model assigned to the site to detect anomalies in multiple features of the network data collected at the site by the NMS.

[0016] In another example, this disclosure relates to a computer-readable storage medium, including instructions that, when executed by one or more programmable processors, cause the one or more programmable processors to: monitor real-time seasonal patterns of multiple features of network data collected at one of a plurality of sites associated with an organization during training; assign the site to one of two or more pattern types based on the real-time seasonal patterns at the site, wherein the two or more pattern types include a random pattern type, a regular pattern type, and a complex pattern type; assign one of two or more anomaly detection models to the site based on the pattern type of the site, wherein the anomaly detection model is associated with the pattern type of the site; and use the anomaly detection model assigned to the site to detect anomalies in multiple features of the network data collected at the site.

[0017] Details of one or more examples of the technology disclosed herein are illustrated in the following figures and description. Other features, objects, and advantages of these technologies will be apparent from the description and figures, as well as from the claims. Attached Figure Description

[0018] Figure 1A This is a block diagram of an example network system including a network management system, based on one or more technologies according to this disclosure.

[0019] Figure 1B It is shown Figure 1A A block diagram providing further examples and details of the network system.

[0020] Figure 2 This is a block diagram of an example access point device based on one or more technologies according to this disclosure.

[0021] Figure 3 This is a block diagram of an example network management system based on one or more technologies disclosed herein.

[0022] Figure 4 This is a block diagram of an example user device equipment based on one or more technologies of this disclosure.

[0023] Figure 5 This is a block diagram of an example network node (such as a router or switch) according to one or more technologies of this disclosure.

[0024] Figure 6 Example features of network data collected within an exemplary time window according to one or more techniques of this disclosure are shown.

[0025] Figure 7 An example comparison of the actual feature value of a feature used for anomaly detection according to one or more techniques of this disclosure with the predicted feature value of that feature is shown.

[0026] Figure 8 This is a flowchart illustrating example operations for detecting anomalies according to one or more techniques of this disclosure.

[0027] Figure 9 This is a flowchart illustrating example operations for selecting an anomaly detection model to detect anomalies according to one or more techniques of this disclosure. Detailed Implementation

[0028] Figure 1A This is a block diagram of an example network system 100 including a network management system (NMS) 130 according to one or more technologies disclosed herein. The example network system 100 includes multiple sites 102A-102N, at which a network service provider manages one or more wireless networks 106A-106N, respectively. Although in Figure 1A In this disclosure, each site 102A-102N is shown as including a single wireless network 106A-106N, but in some examples, each site 102A-102N may include multiple wireless networks, and this disclosure is not limited in this respect.

[0029] Each site 102A-102N includes multiple network access server (NAS) devices, such as access points (APs) 142, switches 146, or routers (not shown). For example, site 102A includes multiple APs 142A-1 to 142A-M. Similarly, site 102N includes multiple APs 142N-1 to 142N-M. Each AP 142 can be any type of wireless access point, including but not limited to commercial or enterprise APs, routers, or any other device connected to a wired network and capable of providing wireless network access to client devices within the site.

[0030] Each site 102A-102N also includes multiple client devices (also known as User Equipment (UE)) representing various wireless-enabled devices within each site, typically referred to as UE or client device 148. For example, multiple UEs 148A-1 to 148A-N are currently located at site 102A. Similarly, multiple UEs 148N-1 to 148N-K are currently located at site 102N. Each UE 148 can be any type of wireless client device, including but not limited to mobile devices such as smartphones, tablets or laptops, personal digital assistants (PDAs), wireless terminals, smartwatches, smart rings, or other wearable devices. UE 148 may also include wired client-side devices, such as IoT devices, such as printers, security devices, environmental sensors, or any other device connected to a wired network and configured to communicate over one or more wireless networks 106.

[0031] To provide wireless network services to UE 148 and / or communicate via wireless network 106, AP 142 and other wired client-side devices at site 102 are directly or indirectly connected to one or more network devices (e.g., switches, routers, etc.) via physical cables (e.g., Ethernet cables). Figure 1A In the example, site 102A includes switch 146A, and each of APs 142A-1 to 142A-M at site 102A is connected to this switch. Similarly, site 102N includes switch 146N, and each of APs 142N-1 to 142N-M at site 102N is connected to this switch. Although in Figure 1AThe diagram appears to show each site 102 comprising a single switch 146 and all APs 142 of a given site 102 connected to that single switch 146. However, in other examples, each site 102 may include more or fewer switches and / or routers. Furthermore, APs and other wired client-side devices at a given site may connect to more than two switches and / or routers. Additionally, more than two switches at a site may be interconnected with each other and / or connected to more than two routers, for example, via a mesh or partial mesh topology in a center-radial architecture. In some examples, the interconnected switches and routers comprise a wired local area network (LAN) hosting a wireless network 106 at site 102.

[0032] Example network system 100 also includes various network components for providing network services within a wired network. As an example, these network components include an authentication, authorization, and accounting (AAA) server 110 for authenticating users and / or UE 148; a dynamic host configuration protocol (DHCP) server 116 for dynamically assigning network addresses (e.g., IP addresses) to UE 148 after authentication; a domain name system (DNS) server 122 for resolving domain names to network addresses; multiple servers 128A-128X (collectively referred to as "Server 128") (e.g., web servers, database servers, file servers, etc.); and a network management system (NMS) 130. Figure 1A As shown, various devices and systems of network 100 are coupled together via one or more networks 134 (e.g., the Internet and / or corporate intranets).

[0033] exist Figure 1A In the example, NMS 130 is a cloud-based computing platform that manages wireless networks 106A-106N at one or more sites 102A-102N. As further described herein, NMS 130 provides management tools and implements an integrated suite of various technologies disclosed herein. Typically, NMS 130 can provide a cloud-based platform for wireless network data acquisition, monitoring, activity logging, reporting, predictive analytics, network anomaly identification, and alarm generation. In some examples, NMS 130 outputs notifications (such as alarms, warnings, graphical indicators on dashboards, log messages, text / SMS messages, email messages, etc.) and / or suggestions regarding wireless network issues to a site or network administrator (“admin”) who interacts with and / or operates admin device 111. Furthermore, in some examples, NMS 130 operates in response to configuration input received from an administrator who interacts with and / or operates admin device 111.

[0034] The administrator and administrator device 111 may include IT personnel and administrator computing devices associated with one or more sites 102. The administrator device 111 may be implemented as any suitable device for presenting output and / or accepting user input. For example, the administrator device 111 may include a display. The administrator device 111 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or an administrator. According to one or more aspects of this disclosure, the administrator device 111 may, for example, represent a workstation, laptop or notebook computer, desktop computer, tablet computer, or any other computing device that can be operated by a user and / or present a user interface. The administrator device 111 may be physically separate from and / or located at a different location from the NMS 130, such that the administrator device 111 can communicate with the NMS 130 via network 134 or other communication methods.

[0035] In some examples, one or more NAS devices (e.g., AP 142, switch 146, or router) may be connected to edge devices 150A-150N via physical cables (e.g., Ethernet cables). Edge device 150 includes a cloud-managed wireless local area network (LAN) controller. Each edge device 150 may include a locally deployed device at site 102 that communicates with NMS 130 to extend certain microservices from NMS 130 to the locally deployed NAS device, while using NMS 130 and its distributed software architecture for scalable and resilient operation, management, troubleshooting, and analysis.

[0036] Each network device in network system 100 (e.g., servers 110, 116, 122 and / or 128, AP 142, UE 148, switch 146, and any other server or device attached to or forming part of network system 100) may include a system log or error log module, wherein each of these network devices records the status of the network device, including normal operation status and error status. Throughout this disclosure, one or more network devices in network system 100 (e.g., servers 110, 116, 122 and / or 128, AP 142, UE 148, and switch 146) may be considered “third-party” network devices when owned by and / or associated with an entity different from NMS 130, such that NMS 130 does not receive, collect, or otherwise access the status and other data recorded by the third-party network devices. In some examples, edge device 150 can provide a proxy through which the status and other data recorded by third-party network devices can be reported to NMS 130.

[0037] In some examples, the NMS 130 monitors network data 137 (e.g., one or more Service Level Expectation (SLE) metrics) received from wireless networks 106A-106N at each site 102A-102N, and manages network resources (such as APs 142 at each site) to deliver a high-quality wireless experience to end users, IoT devices, and clients at the sites. For example, the NMS 130 may include a Virtual Network Assistant (VNA) 133 that implements an event processing platform to provide real-time insights for IT operations and streamline troubleshooting, and to automatically take corrective actions or provide recommendations to proactively resolve wireless network issues. For example, the VNA 133 may include an event processing platform configured to process concurrent streams of network data 137 from sensors and / or agents associated with nodes within AP 142 and / or network 134. For example, the VNA 133 of the NMS 130 may include underlying analytics and network error detection engines and alerting systems according to various examples described herein. The VNA 133's underlying analytics engine can apply historical data and models to inbound event streams to calculate assertions, such as the predicted occurrence of identified anomalies or events constituting network error conditions. Furthermore, the VNA 133 can provide real-time alerts and reports to notify site or network administrators of any predicted events, anomalies, or trends via administrator device 111, and can perform root cause analysis and automatic or assisted error remediation. In some examples, the VNA 133 of the NMS 130 can apply machine learning techniques to identify the root causes of error conditions detected or predicted from the network data stream 137. If the root cause can be resolved automatically, the VNA 133 can invoke one or more corrective actions to correct the root cause of the error condition, thereby automatically improving underlying SLE metrics and also automatically improving the user experience.

[0038] Further examples of the operation implemented by the VNA 133 of the NMS 130 are described in the following patents: U.S. Patent No. 9,832,082, published November 28, 2017, entitled "Monitoring Wireless AccessPoint Events"; U.S. Publication No. US 2021 / 0306201, published September 30, 2021, entitled "Network System Fault Resolution Using a Machine Learning Model"; U.S. Patent No. 10,985,969, published April 20, 2021, entitled "Systems and Methods for a Virtual Network Assistant"; and U.S. Patent No. 10,985,969, published March 23, 2021, entitled "Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault". U.S. Patent No. 10,958,585 entitled “Detection”, published on March 23, 2021, entitled “Method for Spatio-Temporal Modeling”, and U.S. Patent No. 10,958,537 entitled “Method for Conveying AP Error Codes Over BLE Advertisements”, published on December 8, 2020, the entire contents of all these patents are incorporated herein by reference.

[0039] In operation, NMS 130 observes, collects, and / or receives network data 137, which may take the form of time-series data, for example, extracted from messages, counters, and statistics. Depending on one implementation, a computing device is part of NMS 130. Depending on other implementations, NMS 130 may include one or more computing devices, dedicated servers, virtual machines, containers, services, or other forms of environments for performing the techniques described herein. Similarly, computing resources and components implementing VNA 133 may be part of NMS 130, may run on other servers or execution environments, or may be distributed across nodes within network 134 (e.g., routers, switches, controllers, gateways, etc.).

[0040] According to one or more techniques of this disclosure, NMS 130 is configured to detect network anomalies at a site. NMS 130 can use an anomaly detection model to detect anomalies at the site, and the anomaly detection model can be intelligently selected for that site to save computational resources associated with executing various versions of anomaly detection models stored in a repository of anomaly detection models (also referred to herein as an "anomaly detection model repository" or "AD repository"). NMS 130 can provide the selected anomaly detection model with at least two features stored in network data 137 collected within a time window (e.g., 8 hours, 1 day, 1 week, 1 month, etc.). NMS 130 can implement the selected anomaly detection model to output predicted feature values ​​of the at least two features within the time window of the collected network data. For example, NMS 130 can implement the anomaly detection model to predict feature values ​​for the next hour immediately following 3 hours of network data collected. NMS 130 can detect anomalies at site 102 based on the difference between the actual collected feature values ​​of the network data 137 collected within the time window and the predicted feature values ​​of the time window output by the anomaly detection model. NMS 130 can determine the root cause of the anomaly based on the characteristics of network data 137 collected at the site in site 102 where the anomaly was detected. In some examples, NMS 130 can determine that the root cause of the anomaly is an organizational issue based on the determination that multiple sites owned by the organization have similar anomalies that have been detected.

[0041] exist Figure 1AIn the example, NMS 130 may include an anomaly detection and root cause analysis module 135 (also referred to herein as “AD and RCA module 135”) and an anomaly detection model manager 136 (also referred to herein as “AD model manager 136”). In operation, AD and RCA module 135 may process multiple features of network data 137 collected within a time window to identify seasonal patterns associated with the features measured within the time window. Features stored in network data 137 may include time-series data indicating metrics, statistics, or other messages reported by network devices at site 102 that indicate specific behavior of devices at site 102 during a particular time window. For example, network data 137 may store a first feature of statistics reported by AP device 142A indicating the number of active client devices 148A connected to AP device 142A throughout the various time windows. Network data 137 may store a second feature of multiple AP devices 142A reporting statistics indicating the number of active or reporting AP devices 142A throughout the various time windows.

[0042] The AD and RCA module 135 can process at least two features of network data 137 collected within a time window to identify seasonal patterns of the at least two features. For example, the AD and RCA module 135 can process a first set of feature values ​​of a first feature and a second set of feature values ​​of a second feature to identify seasonal patterns of the first and second features within the time window, where the first feature indicates the number of active AP devices 142A at site 102A reporting statistics within the time window (e.g., three hours), and the second feature indicates the number of active client devices 148A at site 102A connected to AP devices 142A within the time window. The AD and RCA module 135 can process the first and second features by stacking the first and second features. The AD and RCA module 135 can also stack the first and second features by aggregating the time-series feature values ​​of the first and second features collected within the time window. For example, AD and RCA module 135 can aggregate (e.g., average, mean, etc.) the first set of feature values ​​and the second set of feature values ​​into buckets of feature values ​​of the first and second features corresponding to consecutive time intervals within the time window, in order to identify seasonal patterns of the first and second features within the time window.

[0043] The AD and RCA module 135 can identify seasonal patterns of a first and a second feature within a time window by generating functions that indicate the behavior of aggregated feature values ​​of the first and second features. For example, the AD and RCA module 135 can create a function based on aggregated feature values ​​collected at a site up to the current time, indicating a distribution of regular seasonal patterns (e.g., normal distribution, t-distribution, constant distribution, linear distribution, etc.) that follow a simple, consistent statistical pattern. In another example, the AD and RCA module 135 can create a function based on aggregated feature values ​​collected at a site up to the current time, indicating a distribution of complex seasonal patterns (e.g., fractal patterns, Markov chains, clustered spatial distributions, etc.) that follow a complex, consistent statistical pattern. In yet another example, the AD and RCA module 135 can create a function based on aggregated feature values ​​collected at a site up to the current time, indicating a distribution of random seasonal patterns (e.g., random or pseudo-random behavior of the features) that appear to follow an inconsistent statistical pattern. The AD and RCA modules 135 can send an indication of a seasonal pattern (e.g., a function) of at least two characteristics of the network data 137 within a time window to the AD model manager 136.

[0044] The AD model manager 136 can select an anomaly detection model from a repository of anomaly detection models based on indications of seasonal patterns received from the AD and RCA modules 135. The AD model manager 136 can select anomaly detection models from the repository to optimize the utilization of computational resources associated with performing anomaly detection. The AD model manager 136 can maintain a repository of anomaly detection models that includes different versions or styles of anomaly detection models with varying complexities (e.g., different numbers of parameters for anomaly detection algorithms, different numbers of neural network layers, different weights or biases in machine learning models, etc.). The AD model manager 136 can also maintain a repository of anomaly detection models that includes a mapping from anomaly detection model versions to clusters corresponding to pattern types of seasonal patterns identified in at least two features of the network data 137. In this way, the AD model manager 136 can develop and deploy anomaly detection models that can be adapted to various pattern types of network behavior features specific to site 102 in the network data 137.

[0045] AD Model Manager 136 can create different versions of anomaly detection models in an anomaly detection model repository based on historical network data stored in network data 137. In some examples, AD Model Manager 136 can first create a threshold anomaly detection model, which can be configured to predict feature values ​​using heuristics or rule-based methods. AD Model Manager 136 can generate a baseline anomaly detection model based on the collected network data 137, utilizing a first set of historical network data from network data 137. For example, AD Model Manager 136 can generate a baseline anomaly detection model as a Long Short-Term Memory (LSTM) network model trained on the first set of historical network data collected during the initialization of the techniques described herein, to predict feature values ​​for site 102 with identified features exhibiting seasonal patterns associated with regular pattern types. For example, if site 102A is assigned to a seasonal pattern cluster associated with a regular pattern type, AD Model Manager 136 selects the baseline anomaly detection model instead of the threshold anomaly detection model to more accurately predict the feature values ​​for anomaly detection at site 102A. If the real-time seasonal pattern of the features observed at site 102A is identified as a random pattern type, the AD model manager 136 can continue to select a threshold anomaly detection model.

[0046] As AD and RCA modules 135 collect subsequent historical network data, they can identify more complex seasonal patterns in the features of site 102. AD and RCA modules 135 can send historical network data and corresponding indications of complex pattern types to AD model manager 136. AD model manager 136 can create a general anomaly detection model as a machine learning model (e.g., a deep learning neural network) based on the indications of complex pattern types and the corresponding historical network data of network data 137. This machine learning model is configured to predict feature values ​​for two or more sites in site 102 that are observed to have features following seasonal patterns associated with complex pattern types. For example, AD model manager 136 can create the general anomaly detection model as a neural network trained to predict feature values ​​based on historical network data identified as being associated with complex pattern types. AD model manager 136 can select a general anomaly detection model for two or more sites in site 102 that have features identified as having seasonal patterns associated with complex pattern types. AD model manager 136 can train the general anomaly detection model based on historical network data of network data 137 that is observed to have features with specific complex pattern types. The AD model manager 136 can select a general anomaly detection model from instances of seasonal patterns corresponding to complex pattern types received from the AD and RCA modules 135. In instances where the AD model manager 136 receives indications of seasonal patterns associated with regular pattern types, the AD model manager 136 can select a baseline anomaly detection model to save computational resources (e.g., processing cycles, memory usage, power consumption, etc.) associated with executing more complex general anomaly detection models.

[0047] In some examples, the AD model manager 136 can fine-tune a general anomaly detection model to create different versions of the anomaly detection model specific to sites in site 102. When the AD model manager 136 receives subsequent historical network data from sites in site 102, it can create retrained and / or fine-tuned versions of the general model (e.g., as a transformer model) adapted to predict feature values ​​associated with complex seasonal patterns observed at sites in site 102. For example, the AD model manager 136 can receive indications from the AD and RCA modules 135 of complex seasonal patterns of features already observed at site 102A. The AD model manager 136 can fine-tune the general anomaly detection model to predict feature values ​​for site 102A based on historical network data 137 associated with features having complex seasonal patterns specific to site 102A. The AD model manager 136 can select a fine-tuned anomaly detection model in response to receiving an indication that the most recent set of network data for site 102A has features associated with complex seasonal patterns specific to site 102A. The AD model manager 136 can select a general anomaly detection model in the following instance: The AD model manager 136 receives an indication of a seasonal pattern of features observed at site 102A, which is associated with a shared pattern type that has been observed at more than one site in 102. In this way, the AD model manager 136 can save computational resources associated with executing a fine-tuned anomaly detection model when regular or shared seasonal feature patterns are observed at site 102A.

[0048] The AD model manager 136 can generate seasonal pattern clusters of seasonal patterns, each seasonal pattern cluster indicating various pattern types of features of site 102 collected for a specific time window. The AD model manager 136 can map seasonal pattern clusters of one or more pattern types that identify at least two features of site 102 to corresponding versions of anomaly detection models in an anomaly detection model repository. For example, the AD model manager 136 can map seasonal pattern clusters of seasonal patterns associated with complex pattern types to a version of the anomaly detection model that has been retrained and / or fine-tuned to predict feature values ​​given complex pattern types within a feature display time window. The AD model manager 136 can maintain the mapping from seasonal pattern clusters to anomaly detection models in the AD model repository.

[0049] During the inference phase of multi-feature anomaly detection, the AD model manager 136 can receive from the AD and RCA modules 135 an indication of the real-time seasonality patterns of the real-time feature data 137 collected at the site up to the current time (e.g., the most recent time-series network data 137 indicating at least two features within a time window). The AD model manager 136 can determine the seasonality pattern clusters of the AD model repository based on the real-time seasonality pattern indications received from the AD and RCA modules 135. For example, the AD model manager 136 can assign a site to a seasonality pattern cluster based on real-time network data sent by the site that exhibits a real-time seasonality pattern associated with the pattern type of the seasonality pattern cluster in the AD model repository. The AD model manager 136 can determine, for example, by comparing the function of the real-time seasonality pattern with the function associated with the pattern type (e.g., comparing phase shifts or transformations, periodicity and frequency, and / or graphical shape or behavior, performing differential analysis, comparing Fourier transforms or representative complex functions, etc.) the pattern type that the site exhibits associated with the seasonality pattern cluster. The AD model manager 136 can assign sites associated with real-time seasonal patterns to seasonal pattern clusters based on a function that determines the real-time seasonal pattern and a function that determines the pattern type of the seasonal pattern cluster (e.g., by a threshold amount). The AD model manager 136 can select the version of the anomaly detection model mapped to the seasonal pattern cluster assigned to the site. The AD and RCA modules 135 can implement the selected version of the anomaly detection model to detect anomalies at the site until the AD model manager 136 refreshes the seasonal pattern cluster assigned to that site. In other words, after a training period, the AD model manager 136 can statically maintain the assignment of pattern types and associated anomaly detection models to sites. The AD manager 136 can re-evaluate and / or reassign sites to seasonal pattern clusters at periodic intervals (e.g., weekly, monthly, etc.) rather than iteratively. In other words, during a training period, the AD model manager can change the pattern type and associated anomaly detection model assigned to a site over time based on changes in the real-time seasonal pattern monitored at the site. In this way, the AD model manager 136 can select an anomaly detection model that is specific to the pattern of observed real-time network data behavior collected at a specific site in site 102, taking into account the computing resources used during anomaly detection.

[0050] The AD model manager 136 can send an instance of the selected anomaly detection model to the AD and RCA modules 135. The AD and RCA modules 135 can execute the instance of the anomaly detection model to predict feature values ​​for at least two features within a time window corresponding to the real-time collected network data of network data 137. In other words, the AD and RCA modules 135 can execute an instance of the anomaly detection model to generate predictions of feature values ​​indicating expected feature values ​​within a time window. For example, if the features include the number of devices collected at site 102A over time (e.g., the number of active client devices 148A and the number of active AP devices 142A reported to NMS 130), the AD and RCA modules 135 can execute an instance of the selected anomaly detection model to predict the number of devices within a time window (e.g., the last 4 hours) based on seasonal patterns of features identified at site 102A.

[0051] The AD and RCA module 135 can detect anomalies based on the actual feature values ​​of network data 137 collected within a time window of the observed feature and the prediction of the feature values ​​within that time window. For example, the AD and RCA module 135 can compare the actual feature values ​​of the network data 137 collected within the time window with the predicted or expected feature values ​​output by the selected anomaly detection model. Based on the comparison, the AD and RCA module 135 can determine the difference between the actual feature values ​​and the predicted feature values ​​as an anomaly at the site regarding the observed feature. For example, the AD and RCA module 135 can determine an anomaly at site 102A as the difference (e.g., two standard deviations) between the actual feature value of the number of devices (e.g., active client devices 148A at site 102A connected to AP device 142A and active AP devices 142A reporting statistics to NMS 130) and the predicted feature value of the number of devices output by the selected anomaly detection model.

[0052] The AD and RCA module 135 can determine the root cause of anomalies detected at a site. The AD and RCA module 135 can determine the root cause of the detected anomaly based on the actual feature values ​​of at least two features stored in the network data 137. For example, the AD and RCA module 135 can determine the root cause of an anomaly indicating the number of devices associated with a second feature of the active client device 148A and AP device 142A at site 102A, based on the difference between the actual feature value and the predicted feature value of the number of devices. The AD and RCA module 135 can generate recommendations to mitigate or resolve the determined root cause. For example, the AD and RCA module 135 can output recommendations to the administrator device 111.

[0053] In some examples, AD and RCA module 135 can determine the scope of a specific anomaly. For example, AD and RCA module 135 can identify an anomaly at site 102A and the same anomaly at site 102N. Based on the anomalies detected by AD and RCA module 135 at sites 102A and 102N, AD and RCA module 135 can determine that the scope of the root cause of the anomaly may be associated with organizational problems of the organization owning sites 102A and 102N. AD and RCA module 135 can generate recommendations to mitigate organizational problems. For example, AD and RCA module 135 can generate recommendations instructing for modification or reconfiguration of AP 142, switch 146, etc., at sites 102A and 102N to mitigate client disconnection anomalies. AD and RCA module 135 can output recommendations to administrator device 111. For example, AD and RCA module 135 can output recommendations as notifications output by applications running at administrator device 111.

[0054] The technology disclosed herein provides one or more technical advantages and practical applications. For example, the technology enables the detection of site-specific anomalies in real-time or near real-time. NMS 130 can process multivariate, time-series network data to detect anomalies associated with features indicated in the network data (e.g., the number of client devices, the number of AP devices, combinations of the number of client devices and the number of AP devices, etc.). NMS 130 can develop and maintain various versions of anomaly detection models that can be trained to effectively predict feature values ​​identified as having various seasonal patterns. NMS 130 can develop versions of the anomaly detection model that can be specifically trained and designed to predict feature values ​​of features observed at a particular site. In this way, NMS 130 can use the feature value predictions generated by the anomaly detection model to detect anomalies in the network data features of a site, which has been specifically trained to predict the feature values ​​of the site. NMS 130 can intelligently select which version of the anomaly detection model to execute for anomaly detection. By executing a baseline anomaly detection model for features identified as having stable, regular seasonal patterns, the NMS 130 can save computational resources (e.g., processing cycles, memory usage, power consumption) associated with executing complex anomaly detection models for eigenvalue prediction (e.g., fine-tuning or retraining anomaly detection models). In general, the NMS 130 can detect anomalies in real-time or near real-time to determine root causes and generate recommendations for resolving potential network problems associated with the detected anomalies.

[0055] Although the techniques of this disclosure are described in this example as being performed by NMS 130, the techniques described herein can be performed by any other computing device, system, and / or server, and this disclosure is not limited thereto. For example, one or more computing devices configured to perform the functions of the techniques of this disclosure may reside in a dedicated server or be included in any other server besides NMS 130, or may be distributed throughout network 100 and may or may not be part of NMS 130.

[0056] Figure 1B It is shown Figure 1A A block diagram showing further details of another example of a network system. In this example... Figure 1B The NMS 130 is shown, which is configured to operate based on an artificial intelligence / machine learning-based computing platform that provides services from "clients" (e.g., user equipment 148 connected to wireless network 106 and wired LAN 175). Figure 1B (The far left) crosses over to the “cloud” (e.g., cloud-based application services 181 that can be hosted by computing resources within data center 179). Figure 1B The far right) offers full automation, insights, and assurance (WiFi assurance, wired assurance, and WAN assurance).

[0057] As described herein, NMS 130 provides management tools and an integrated suite of technologies disclosed herein. Typically, NMS 130 can provide a cloud-based platform for wireless network data acquisition, monitoring, activity logging, reporting, predictive analytics, network anomaly detection, and alarm generation. For example, the network management system 130 can be configured to proactively monitor and adaptively configure network 100 to provide automated capabilities. Furthermore, VNA 133 includes a natural language processing engine to provide AI-driven support and troubleshooting, anomaly detection, AI-driven location services, and AI-driven radio frequency (RF) optimization with reinforcement learning.

[0058] like Figure 1BAs shown in the example, the AI-driven NMS 130 also provides configuration management, monitoring, and automated supervision of the software-defined wide area network (SD-WAN) 177, which operates as an intermediate network that communicatively couples the wireless network 106 and wired LAN 175 to the data center 179 and application services 181. Typically, the SD-WAN 177 provides seamless, secure, traffic-engineered connectivity between the “radiating” router 187A that hosts the wireless network 106 (such as a branch or campus network) and the “central” router 187B that goes further up the cloud stack toward the cloud-based application services 181. The SD-WAN 177 typically runs on and manages the overlay network on the underlying physical wide area network (WAN), which provides connectivity for geographically separated customer networks. In other words, SD-WAN 177 extends software-defined networking (SDN) capabilities to WAN and allows networks to decouple the underlying physical network infrastructure from virtualized network infrastructure and applications, enabling the network to be configured and managed in a flexible and scalable manner.

[0059] In some examples, the underlying routers of the SD-WAN 177 can implement a stateful, session-based routing scheme in which routers 187A and 187B dynamically modify the contents of the original packet headers originating from client device 148 to direct traffic along a selected path (e.g., path 189) toward application service 181 without using tunnels and / or additional labeling. In this way, routers 187A and 187B can be more efficient and scalable for large networks because the use of tunnelless, session-based routing allows routers 187A and 187B to achieve considerable network resource efficiency by avoiding the need to perform encapsulation and decapsulation at tunnel endpoints. Furthermore, in some examples, each router 187A and 187B can independently perform path selection and traffic engineering to control the packet flow associated with each session without using a centralized SDN controller for path selection and label distribution. In some examples, routers 187A and 187B implement session-based routing as Secure Vector Routing (SVR) provided by Juniper Networks.

[0060] Additional information regarding session-based routing and SVR is described in the following patents: U.S. Patent No. 9,729,439, published August 8, 2017, entitled "Computer Network Packet Flow Controller"; U.S. Patent No. 9,729,682, published August 8, 2017, entitled "Network Apparatus and Method for Processing Sessions Using a Packet Signature"; U.S. Patent No. 9,762,485, published September 12, 2017, entitled "Network Packet Flow Controller with Extended Session Management"; and U.S. Patent No. 9,762,485, published January 16, 2018, entitled "Router with Optimized Statistical Functions". U.S. Patent No. 9,871,748, entitled "Optimized Statistical Functionality"; U.S. Patent No. 9,985,883, published on May 29, 2018, entitled "Name-Based Routing System and Method"; U.S. Patent No. 10,200,264, published on February 5, 2019, entitled "Link Status Monitoring Based on Packet Loss Detection"; U.S. Patent No. 10,277,506, published on April 30, 2019, entitled "Stateful Load Balancing in a Stateless Network"; and U.S. Patent No. 10,277,506, published on October 1, 2019, entitled "Network Packet Flow Controller with Extended Session Management". U.S. Patent No. 10,432,522, entitled “PACKET FLOWCONTROLLER WITH EXTENDED SESSION MANAGEMENT”; and U.S. Patent Application Publication No. 2020 / 0403890, entitled “IN-LINE PERFORMANCE MONITORING”, published on December 24, 2020, the entire contents of each of these patents are incorporated herein by reference.

[0061] In some examples, the AI-driven NMS 130 can enable intent-based configuration and management of network system 100, including the construction, presentation, and execution of intent-driven workflows for configuring and managing devices associated with wireless network 106, wired LAN network 175, and / or SD-WAN 177. For example, declarative requirements express the desired configuration of network components without specifying precise native device configurations and control flows. By utilizing declarative requirements, what should be done is specified, rather than how it should be done. Declarative requirements can contrast with the necessary instructions that describe the exact device configuration syntax and control flows required to achieve the configuration. By utilizing declarative requirements instead of imperative instructions, users and / or user systems alleviate the burden of determining the exact device configurations needed to achieve the user / system's desired results. For example, when utilizing various types of devices from different vendors, specifying and managing exact imperative instructions for configuring each device in the network is often difficult and cumbersome. The types and kinds of devices in the network can change dynamically as new devices are added and devices fail. Managing various types of devices from different vendors with different configuration protocols, syntaxes, and software versions to configure a cohesive device network is often challenging. Therefore, the management and configuration of network devices become more efficient by requiring only the user / system to specify declarative requirements (which specify the expected results applicable across various types of devices). Further examples of details and techniques for intent-based network management systems are described in the following patents: U.S. Patent No. 10,756,983, entitled "Intent-based Analytics," and U.S. Patent No. 10,992,543, entitled "Automatically generating anintent-based network model of an existing computer network," the entire contents of which are incorporated herein by reference.

[0062] According to the techniques described in this disclosure, NMS 130 can detect anomalies at network sites. The AD and RCA modules 135 of NMS 130 can detect anomalies during a time window of observed features by executing an instance of an anomaly detection model received from the AD model manager 136 of NMS 130. The AD model manager 136 can select a version of the anomaly detection model to send to the AD and RCA modules 135 based on seasonal patterns of feature data 137 collected within the time window (e.g., network data collected at the site up to the current time). For example, the AD model manager 136 can select a baseline anomaly detection model based on associating the identified seasonal pattern with a regular pattern type. In another example, the AD model manager 136 can select a fine-tuned anomaly detection model trained to predict feature values ​​for a specific site based on associating the identified seasonal pattern with a complex pattern type. In yet another example, the AD model manager 136 can select a threshold anomaly detection model based on associating the identified seasonal pattern with a random pattern type. The AD and RCA modules 135 can detect anomalies observed at a site based on the difference between the predicted feature values ​​within a time window output by the instance of the anomaly detection model and the actual feature values ​​of the network data 137 collected within that time window.

[0063] In one implementation, AD and RCA module 135 can identify seasonal patterns in the number of devices collected at site 102A over time. For example, AD and RCA module 135 can process time-series data of network data 137 associated with site 102A to determine the number of active client devices 148A within a week and the number of AP devices 142A that report statistics associated with the active client devices 148A within that week. AD and RCA module 135 can identify seasonal patterns in the number of active client devices 148A and active AP devices 142A throughout the week as real-time seasonal patterns. AD and RCA module 135 can send an indication of the seasonal pattern of the number of devices within that week to AD model manager 136.

[0064] AD Model Manager 136 can assign site 102A to one of two or more pattern types, including at least a random pattern type, a regular pattern type, and a complex pattern type. AD Model Manager 136 can maintain seasonal pattern clusters associated with one of the two or more pattern types. For example, AD Model Manager 136 can maintain a first seasonal pattern cluster for the random pattern type, a second seasonal pattern cluster for the regular pattern type, and a third seasonal pattern cluster for the complex pattern type. AD Model Manager 136 can develop anomaly detection models for predicting feature values ​​associated with each feature in the pattern type. For example, AD Model Manager 136 can develop a first anomaly detection model as a threshold anomaly detection model for the seasonal pattern cluster of the random pattern type, a second anomaly detection model as a baseline anomaly detection model for the seasonal pattern cluster of the regular pattern type, and a third anomaly detection model as a fine-tuned anomaly detection model for the seasonal pattern cluster of the complex pattern type. When AD Model Manager 136 collects additional network data for site 102, AD Model Manager 136 can create additional versions of the anomaly detection model for the additional pattern types. For example, AD model manager 136 can create a generic anomaly detection model for seasonal patterns of a specific pattern type shared between at least two sites in site 102. In this way, AD model manager 136 can maintain versions of anomaly detection models of varying complexity that consume different amounts of computational resources during execution. AD model manager 136 can select the version of the anomaly detection model based on the pattern type assigned to the site, according to the identified seasonal patterns received from AD and RCA modules 135.

[0065] The AD and RCA module 135 can execute an instance of a selected anomaly detection model to predict the number of devices in a time window based on seasonal patterns and the number of devices determined for one or more previous time windows. For example, the AD and RCA module 135 can predict the number of AP devices 142A and client devices 148A at site 102A within a time window (e.g., one week) based on seasonal patterns and the number of devices determined for one or more previous time windows (e.g., the number of devices indicated in network data collected within a previous one-week time window). The AD and RCA module 135 can detect anomalies at site 102A during the time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices for that time window. For example, the AD and RCA module can compare the actual number of active client devices 148A and active AP devices 142A at site 102A within a one-week time window, as indicated in network data 137, with the predicted number of devices during the one-week time window output by the selected anomaly detection model.

[0066] The AD and RCA module 135 can determine the root cause of detected anomalies. For example, the AD and RCA module 135 can analyze characteristic data in network data 137 to identify the root cause of the anomaly. The AD and RCA module 135 can determine whether more than one site experienced the anomaly. Based on the AD and RCA module 135 detecting anomalies at more than one site owned by the organization, the AD and RCA module 135 can determine that the root cause of the anomaly is an organizational issue. The AD and RCA module 135 can generate recommendations for network topology adjustments or network configurations that can mitigate and / or resolve the root cause. The AD and RCA module 135 can output recommendations to administrators of the sites associated with the anomaly.

[0067] Figure 2 This is a block diagram of an example access point (AP) device 200 according to one or more technologies of this disclosure. Figure 2 The example access point 200 shown can be used to implement, as described in this article... Figure 1A Any of the AP 142 shown and described. Access point 200 may include, for example, a Wi-Fi, Bluetooth and / or Bluetooth Low Energy (BLE) base station or any other type of wireless access point.

[0068] exist Figure 2 In the example, access point 200 includes a wired interface 230, wireless interfaces 220A-220B, one or more processors 206, memory 212, and input / output 210 coupled together via bus 214. These components can exchange data and information via bus 214. Wired interface 230 represents a physical network interface and includes a receiver 232 and a transmitter 234 for sending and receiving network communications (e.g., packets). Wired interface 230 couples access point 200 directly or indirectly to wired network devices within a wired network, such as Ethernet cables, via cables such as Ethernet cables. Figure 1A One of the switches in the 146.

[0069] The first wireless interface 220A and the second wireless interface 220B represent wireless network interfaces, and respectively include receiver 222A and receiver 222B, each receiver including a receiving antenna, through which access point 200 can receive signals from wireless communication devices (such as wireless communication devices). Figure 1A The UE 148 in the interface receives wireless signals. The first wireless interface 220A and the second wireless interface 220B also include transmitters 224A and 224B, respectively. Each transmitter includes a transmitting antenna, through which the access point 200 can transmit wireless signals to wireless communication devices (such as…). Figure 1A(UE 148 in the example). In some examples, the first wireless interface 220A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second wireless interface 220B may include a Bluetooth interface and / or a Bluetooth Low Energy (BLE) interface.

[0070] Processor 206 is a programmable, hardware-based processor configured to execute software instructions, such as software instructions for defining software or computer programs, which are stored in a computer-readable storage medium (such as memory 212), such as a non-transitory computer-readable medium including storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause one or more processors 206 to perform the techniques described herein.

[0071] Memory 212 includes one or more devices configured to store programming modules and / or data associated with the operation of access point 200. For example, memory 212 may include a computer-readable storage medium (such as a non-transitory computer-readable medium that includes storage devices (e.g., disk drives or optical drives) or memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory), which stores instructions to cause one or more processors 206 to perform the techniques described herein.

[0072] In this example, memory 212 stores executable software, including an application programming interface (API) 240, a communication manager 242, configuration settings 250, a device status log 252, a data storage 254, and a log controller 255. The device status log 252 includes a list of events specific to access point 200. Events can include logs of both normal and error events, such as memory status, reboot or restart events, crash events, cloud disconnection with self-recovery events, low link speed or link speed fluctuation events, Ethernet port status, Ethernet interface packet errors, upgrade failure events, firmware upgrade events, configuration changes, etc., along with the time and date stamp for each event. The log controller 255 determines the device's log level based on instructions from NMS 130. Data storage 254 can store any data used and / or generated by access point 200, including data collected from UE 148, such as data used to calculate one or more SLE metrics, which is sent by access point 200 for cloud-based management of wireless network 106A by NMS 130.

[0073] Input / output (I / O) 210 represents physical hardware components that enable user interaction, such as buttons, displays, etc. Although not shown, memory 212 typically stores executable software for controlling the user interface for input received via I / O 210. Communication manager 242 includes program code that, when executed by processor 206, allows access point 200 to communicate with UE 148 and / or network 134 via interface 230 and / or any of 220A-220B. Configuration settings 250 include any device settings for access point 200, such as the radio settings of each of wireless interfaces 220A-220B. These settings can be configured manually or can be remotely monitored and managed by NMS 130 to optimize wireless network performance periodically (e.g., hourly or daily).

[0074] As described herein, AP device 200 can measure network data from status log 252 and report it to NMS 130. Network data may include the number of devices (e.g., the number of client devices connected to AP device 200), event data, telemetry data, and / or other SLE-related data. Network data may include various parameters indicating the performance and / or status of the wireless network. These parameters may be measured and / or determined by one or more UE devices and / or one or more APs in the wireless network. NMS 130 may determine one or more SLE metrics based on the SLE-related data received from the APs in the wireless network and store the SLE metrics as network data 137 (…). Figure 1A ).

[0075] Figure 3 This is a block diagram of an example network management system (NMS) 300 according to one or more technologies of this disclosure. The NMS 300 can be used to implement, for example... Figures 1A to 1B The NMS 130 is used in this example. In such an example, the NMS 300 is responsible for monitoring and managing one or more wireless networks 106A-106N at sites 102A-102N.

[0076] The NMS 300 includes a communication interface 330, one or more processors 306, a user interface 310, a memory 312, and a database 318. The components are coupled together via a bus 314, through which they can exchange data and information. In some examples, the NMS 300 receives data from client devices 148, APs 142, switches 146, and other network nodes within the network 134 (e.g., ...). Figure 1BThe NMS 300 receives data from one or more of the routers (187) in the network, which can be used to calculate one or more SLE metrics and / or update network data 316 in the database 318. The NMS 300 analyzes this data for cloud-based management of the wireless networks 106A-106N. In some examples, the NMS 300 may be... Figure 1A Part of another server or any other server shown.

[0077] Processor 306 executes software instructions (such as software instructions for defining software or computer programs) stored in a computer-readable storage medium (such as memory 320), which is a non-transitory computer-readable medium such as including storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, the computer-readable medium storing instructions to cause one or more processors 306 to perform the techniques described herein.

[0078] The communication interface 330 may include, for example, an Ethernet interface. The communication interface 330 couples the NMS 300 to a network and / or the Internet, such as... Figure 1A This refers to any network and / or any local area network (LAN) in network 134 shown. Communication interface 330 includes a receiver 332 and a transmitter 334, through which the NMS 300 receives data from client devices 148, APs 142, switches 146, servers 110, 116, 122, 128, and / or forms networks such as... Figure 1A Any other network node, device, or entity within the network system 100 shown may receive or send data and information to it. In some scenarios described herein, where the network system 100 includes “third-party” network devices owned and / or associated with different entities of the NMS 300, the NMS 300 does not receive, collect, or otherwise access network data from these third-party network devices.

[0079] The data and information received by the NMS 300 may include, for example, telemetry data, SLE-related data, or data from client devices 148, APs 142, switches 146, or other network nodes used by the NMS 300 for remotely monitoring the performance of the wireless network 106A-106N and application sessions from client devices to cloud-based application servers (e.g., Figure 1BThe NMS 300 receives event data from one or more of the routers (187) in the network. The NMS 300 can also send data via the communication interface 330 to any network device, such as client device 148, AP 142, switch 146, other network nodes within the network 134, and administrator device 111, to remotely manage portions of the wireless network 106A-106N and the wired network.

[0080] Memory 312 includes one or more devices configured to store programming modules and / or data associated with the operation of NMS 300. For example, memory 312 may include a computer-readable storage medium, such as a non-transitory computer-readable medium that includes storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause one or more processors 306 to perform the techniques described herein.

[0081] In this example, memory 312 includes API 320, SLE module 322, Virtual Network Assistant (VNA) / AI engine 350, and Radio Resource Management (RRM) engine 360. According to the disclosed technology, VNA / AI engine 350 includes an anomaly detection and root cause analysis module 335 (also referred to herein as "AD and RCA module 335") and an anomaly detection (AD) model manager 336. Figure 3 The AD and RCA modules 335 and AD model manager 336 in Figure 1 can be examples or alternative implementations of the AD and RCA modules 135 and AD model manager 136, respectively. Figure 3In the example, the AD model manager 336 may include a model selector 352, a fine-tuning module 354, and a model repository 380. The model selector 352 may include computer-readable instructions for selecting an anomaly detection model to be executed by the AD and RCA modules 335 during anomaly detection, according to the techniques described herein. The fine-tuning module 354 may include computer-readable instructions for retraining and fine-tuning versions of the anomaly detection model executed by the AD and RCA modules 335 during anomaly detection, according to the techniques described herein. The model repository 380 may include a storage device configured to store versions of the anomaly detection model executed by the AD and RCA modules 335 during anomaly detection, according to the techniques described herein. The model repository 380 may store a mapping from anomaly detection model versions to seasonal pattern clustering, which the model selector 352 may utilize to select versions of the anomaly detection model for anomaly detection. The NMS 300 may also include any other programming modules, software engines, and / or interfaces configured for remote monitoring and management of portions of the wireless network 106A-106N and wired network, including AP 142 / 200, switch 146, or other network devices (e.g., Figure 1B Remote monitoring and management of any of the routers in the router (187).

[0082] SLE module 322 enables the establishment and tracking of thresholds for SLE metrics for each network 106A-106N. SLE module 322 also analyzes SLE-related data collected by APs (such as any of AP 142) from UEs in each wireless network 106A-106N. For example, APs 142A-1 to 142A-N collect SLE-related data from UEs 148A-1 to 148A-N currently connected to wireless network 106A. This data is sent to NMS 300, where SLE module 322 performs actions to determine one or more SLE metrics for each UE 148A-1 to 148A-N currently connected to wireless network 106A. In addition to any network data collected by one or more APs 142A-1 to 142A-N in wireless network 106A, this data is also sent to NMS 300 and stored as network data 316, for example, in database 318. Figure 3In the example, network data 316 may include historical network data 317 and network data 319. Historical network data 317 may include time-series network data (e.g., statistics, metrics, or other messages associated with the network connectivity of client device 148A and AP device 142A at site 102A) indicating characteristics of network connectivity at the site, which has been collected through typical operation of the network devices. For example, historical network data 317 may include time-series network data for a first characteristic indicating the number of client devices connected to each AP device at the site, which may be used to determine a second characteristic of active AP devices based on which AP devices are reporting network data. Network data 319 may include real-time, time-series network data of characteristics reported by AP devices within an observation time window, for comparison with predicted characteristic values ​​of characteristics as described herein.

[0083] RRM Engine 360 ​​monitors one or more metrics at each site 102A-102N to learn and optimize the RF environment at each site. For example, RRM Engine 360 ​​can monitor coverage and capacity SLE metrics for wireless network 106 at site 102 to identify potential SLE coverage and / or capacity issues in wireless network 106 and adjust the radio settings of access points at each site to resolve the identified issues. For example, RRM Engine 360 ​​can determine the channel and transmit power distribution among all AP142s in each network 106A-106N. For example, RRM Engine 360 ​​can monitor events, power, channels, bandwidth, and the number of clients connected to each AP. RRM Engine 360 ​​can also automatically change or update the configuration of one or more AP142s at site 102 to improve coverage and capacity SLE metrics and thus provide users with an improved wireless experience.

[0084] The VNA / AI engine 350 analyzes data received from network devices and its own data to identify when an unexpected anomalous state is encountered at a network device. For example, the VNA / AI engine 350 can identify the root cause of any unexpected or anomalous state, such as any poor SLE metric indicating connectivity problems at one or more network devices. Furthermore, the VNA / AI engine 350 can automatically invoke one or more corrective actions designed to resolve the root cause of one or more identified poor SLE metrics. Examples of corrective actions that can be automatically invoked by the VNA / AI engine 350 may include, but are not limited to, invoking the RRM360 to restart one or more APs, adjusting / modifying the transmit power of a specific radio in a specific AP, adding an SSID configuration to a specific AP, changing the channel on an AP or a group of APs, etc. Corrective actions may also include restarting switches and / or routers, invoking the download of new software to APs, switches, or routers, etc. These corrective actions are given for illustrative purposes only, and this disclosure is not limited in this respect. If automatic corrective actions are unavailable or insufficient to address the root cause, the VNA / AI engine 350 can proactively provide notifications of recommended corrective actions to be taken by IT personnel (e.g., site administrators using Administrator Device 111 or network administrators) to resolve network errors.

[0085] According to one or more techniques disclosed herein, the VNA / AI engine 350 can effectively detect anomalies associated with network data 319. The AD and RCA modules 335 of the VNA / AI engine 350 can process features of the network data 319 collected within a time window to identify seasonal patterns in the features of the network data 319. For example, the AD and RCA modules 335 can aggregate portions (e.g., 10-minute buckets) of time-series feature values ​​of the network data 319 collected during a time window indicating a time interval up to the current time. The AD and RCA modules 335 can identify seasonal patterns based on the aggregated time-series feature values ​​of the network data 319. The AD and RCA modules 335 can send an indication of the identified seasonal patterns of the features of the network data 319 collected within the time window to the AD model manager 336.

[0086] The model selector 352 of the AD model manager 336 can select an anomaly detection model version from the model repository 380 based on indications of real-time seasonal patterns of features of network data 319 collected during a time window. The model selector 352 can select the anomaly detection model version by assigning seasonal pattern clusters from the model repository 380 to the real-time seasonal patterns of features of network data 319. The AD model manager 336 can generate seasonal pattern clusters, each corresponding to a pattern type observable for a set of features contained in historical network data 317. The AD model manager 336 can generate a first seasonal pattern cluster corresponding to a regular pattern type, indicating that the set of features follows a regular pattern (e.g., normal, t-distribution, linear distribution, etc.) of consistent distribution of feature values ​​within the time window. The AD model manager 336 can generate a second seasonal pattern cluster corresponding to a first complex pattern type, indicating that the set of features follows a first specific distribution of feature values ​​within the time window. The AD model manager 336 can generate a third seasonal pattern cluster corresponding to a second complex pattern type, whereby the second complex pattern type indicates that the set of features follows a second specific distribution of feature values ​​within a time window. The AD model manager 336 can assign an anomaly detection model configured to predict feature values ​​based on a specific pattern type to the corresponding seasonal pattern cluster associated with that specific pattern type. The AD model manager 336 can store the mapping between seasonal pattern clusters of pattern types and corresponding anomaly detection models in the model repository 380.

[0087] Model selector 352 can select an anomaly detection model version based on the version of the anomaly detection model mapped to a seasonal pattern cluster, which is assigned to the identified real-time seasonal pattern of the network data 319. Model selector 352 can assign the real-time seasonal pattern to a seasonal pattern cluster stored in model repository 380 by either clustering (e.g., K-means clustering) or mapping the real-time seasonal pattern to a seasonal pattern cluster in model repository 380. For example, model selector 352 can assign the real-time seasonal pattern to a seasonal pattern cluster based on determining that the real-time seasonal pattern resembles a specific pattern type corresponding to the seasonal pattern cluster (e.g., the seasonal pattern cluster corresponds to a distribution that behaves similarly to the real-time seasonal pattern). Model selector 352 can select an anomaly detection model version mapped to the assigned seasonal pattern cluster, which is configured to predict feature values ​​for representing features of the network data 319 that exhibit the pattern type associated with the real-time seasonal pattern.

[0088] The model selector 352 can send an instance of the selected anomaly detection model to the AD and RCA modules 335. The AD and RCA modules 335 can execute the instance of the selected anomaly detection model to predict the feature values ​​of the observed network data 319 within a time window. The AD and RCA modules 335 can determine the anomalies of the features of the network data 319 based on a comparison between the predicted feature values ​​of the features within the time window and the actual feature values ​​of the features of the network data 319 collected during the time window.

[0089] Prior to the inference phase where the VNA / AI engine 350 identifies anomalies at a site, during the training cycle, the AD model manager 336 can monitor seasonal patterns of features in network data 316 collected at network sites associated with the organization. Based on the monitored seasonal patterns and historical network data 317, the AD model manager 336 can generate a version of the anomaly detection model to be stored in the model repository 380. The AD model manager 336 can generate a threshold anomaly detection model as an anomaly detection algorithm that implements rules and heuristics to predict feature values ​​based on the pattern types of features assigned to a site. For example, the AD model manager 336 can generate a threshold anomaly detection model during the initialization of the techniques described herein to predict feature values ​​for anomaly detection.

[0090] When the NMS 300 stores additional network data in the historical network data 317 indicating features of multiple sites, the AD model manager 336 can generate additional anomaly detection models to adapt to trends observed in the historical network data 317. For example, the AD model manager 336 can generate a baseline anomaly detection model as a neural network trained to predict feature values ​​associated with features observed to have a regular pattern type indicating consistent behavior (e.g., normal distribution function, t-distribution function, linear function, etc.). The AD model manager can additionally or alternatively generate a general anomaly detection model based on the historical network data 317, which can be used to predict feature values ​​associated with features observed at multiple sites having a specific pattern type indicating consistent behavior shared by multiple sites. The model selector 352 can select the baseline anomaly detection model when the features collected by the sites indicate a regular pattern type, saving computational resources associated with executing a slightly more complex general anomaly detection model.

[0091] The fine-tuning module 354 of the AD model manager 336 can retrain and / or fine-tune the baseline anomaly detection model to adapt it to site-specific complex seasonal patterns. For example, based on historical network data 317 with a sufficient amount of feature data (e.g., three weeks of feature data), the fine-tuning module 354 can fine-tune the baseline anomaly detection model to predict feature values ​​associated with complex pattern types that may have been observed at a particular site. The fine-tuning module 354 can train a version of the anomaly detection model as a deep learning model (e.g., a transformer model) based on the historical network data 317 to predict feature values ​​for a set of features in the historical network data 317 that have been identified as having complex pattern types. The fine-tuning module 354 can assign the deep learning model to seasonal pattern clusters associated with complex seasonal patterns. The fine-tuning module 354 can store the deep learning anomaly detection model and the mapping from the deep learning anomaly detection model to the seasonal pattern clusters in the model repository 380. When model selector 352 receives an indication that the features at the observation site follow a seasonal pattern associated with a complex pattern type mapped to the fine-tuned anomaly detection model, model selector 352 can select the fine-tuned anomaly detection model. When the features at the observation site follow a seasonal pattern associated with a regular pattern type, model selector 352 can select the baseline anomaly detection model; thereby saving computational resources associated with executing the fine-tuned anomaly detection model.

[0092] The AD model manager 336 can store in the model repository 380 a mapping from seasonal pattern clusters associated with corresponding pattern types to versions of anomaly detection models. For example, the model repository 380 can maintain a first mapping from seasonal pattern clusters associated with random pattern types to a threshold anomaly detection model that displays features of seasonal patterns associated with that random pattern. The model repository 380 can maintain a second mapping from seasonal pattern clusters associated with regular pattern types to a baseline anomaly detection model that displays features of seasonal patterns associated with that regular pattern. The model repository 380 can maintain a third mapping from seasonal pattern clusters associated with complex pattern types to a fine-tuned anomaly detection model that displays features of seasonal patterns associated with that complex pattern.

[0093] In some examples, the anomaly detection model stored in model repository 380 may include one or more supervised ML models. These supervised ML models are trained using historical network data 317 as training data, including pre-collected, labeled network data received from network devices (e.g., client devices, APs, switches, and / or other network nodes), to identify statistical patterns in the feature metrics of the network data used to generate predictions for anomaly detection. The supervised ML model for the anomaly detection model may include one of the following: LSTM model, neural network, logistic regression, Naive Bayes, support vector machine (SVM), etc. In other examples, the anomaly detection model in model repository 380 may include unsupervised ML models. The anomaly detection model in model repository 380 can be trained in batches or periodically. For example, after new feature data is stored in historical network data 317 (e.g., new historical network data collected within the past week), fine-tuning module 354 may retrain or otherwise fine-tune the anomaly detection model in model repository 380. Although not explicitly stated in [the original text]... Figure 3 As shown, in some examples, database 318 can store training data, and VNA / AI engine 350 or dedicated training module can be configured to train an anomaly detection model based on the training data to determine appropriate weights for one or more features across the training data.

[0094] Model selector 352 can assign anomaly detection models from model repository 380 to sites based on real-time seasonal patterns at the sites. For example, model selector 352 can monitor real-time seasonal patterns of features of network data 319 collected at the sites during the training period. Model selector 352 can assign sites to seasonal pattern clusters associated with pattern types in model repository 380 based on the real-time seasonal patterns of features of network data 319 collected at the sites. For example, model selector 352 can compare the real-time seasonal patterns of a site with the corresponding pattern types associated with seasonal pattern clusters maintained by model repository 380. Model selector 352 can select anomaly detection models mapped to the assigned pattern types. AD model manager 336 can change the assignment of seasonal pattern clusters to sites based on subsequent real-time network data indicating different seasonal patterns during the training period. AD model manager 336 can maintain the assignment of sites to seasonal pattern clusters as static after the training period. In this way, the AD model manager 336 will not have to continuously select anomaly detection models until a training cycle is triggered (e.g., weekly, monthly, etc.); thus, it saves computational resources associated with iteratively allocating and selecting anomaly detection models.

[0095] The AD and RCA module 335 can detect anomalies based on the version of the anomaly detection model selected by the model selector 352. For example, in an instance where the AD and RCA module 335 executes a baseline anomaly detection model to predict feature values ​​for a set of features of real-time network data 319 collected during a time window (e.g., where the set of features of real-time network data 319 is identified as a seasonal pattern following a regular t-distribution), the AD and RCA module 335 can detect anomalies based on the mean of the predicted feature values ​​of that set of features relative to the mean of the actual feature values ​​of that set of features. The AD and RCA module 335 can also detect anomalies based on determining a threshold amount by which the mean of the predicted feature values ​​differs from the mean of the actual feature values ​​within the time window. In some cases, the AD and RCA module 335 can detect anomalies based on the fact that the actual feature values ​​collected during the time window are one or more standard deviations (e.g., 2.5 standard deviations) larger than the predicted feature values ​​within that time window. In some examples, where the AD and RCA module 335 executes an instance of a finely tuned anomaly detection module to predict feature values ​​for a set of features of network data 319 collected during a time window (e.g., where the set of features of network data 319 is identified as following a seasonal pattern of a complex pattern type), the AD and RCA module 335 may detect anomalies based on the standard deviation or quantile associated with a comparison of the predicted feature values ​​with the actual feature values ​​within the time window.

[0096] In one implementation, AD and RCA module 335 can detect anomalies at sites associated with client device disconnections at a site. For example, AD and RCA module 335 can identify seasonal patterns in the number of devices indicated in network data 319 as seasonal patterns in the number of client devices at a site and seasonal patterns in the number of AP devices reported as statistics at a site. AD and RCA module 335 can send an indication of the identified seasonal patterns in the number of devices to AD model manager 336. Model selector 352 of AD model manager 336 can select an anomaly detection model based on the identified seasonal patterns in the number of devices. Model selector 352 can send an instance of the selected anomaly detection model to AD and RCA module 335. AD and RCA module 335 can execute the instance of the selected anomaly detection model to predict the number of devices in a time window (e.g., 1 day) as the expected number of active client devices and active AP devices at a site throughout the time window. AD and RCA module 335 can detect anomalies of client device disconnections during that time window based on the difference between the actual number of devices indicated in network data 319 and the predicted number of devices output by the selected anomaly detection model.

[0097] The AD and RCA module 335 can determine the root cause of a detected anomaly. The AD and RCA module 335 can analyze the feature data of network data 319 to determine whether one of the features in the feature data is the root cause of the detected anomaly. In some examples, the AD and RCA module 335 can determine the root cause of the anomaly at the organizational level, in cases where the AD and RCA module 135 determines that anomalies were detected at more than one site owned by the organization. The AD and RCA module 335 can generate recommendations on how to resolve and / or mitigate the identified root cause. The AD and RCA module 335 can output these recommendations via user interface 310.

[0098] The technology disclosed herein provides one or more technical advantages and practical applications. For example, NMS 300 can detect anomalies associated with real-time network data 319. NMS 300 can maintain various versions of anomaly detection models, each configured to predict feature values ​​corresponding to features with specific seasonal patterns. NMS 300 can cluster feature data of real-time network data 319 of a network site into seasonal pattern clusters to determine the appropriate version of the anomaly detection model to apply during anomaly detection. By maintaining anomaly detection models trained to predict feature values ​​following specific seasonal patterns, NMS 300 can execute specific versions of the anomaly detection model to optimize computational resources used for anomaly detection. NMS 300 can update the anomaly detection model based on site-specific feature data by fine-tuning or retraining a baseline anomaly detection model or a general anomaly detection model to predict feature values ​​for site-specific features. In this way, NMS 300 can detect site-specific anomalies and generate recommendations for addressing or mitigating the identified root causes of detected anomalies.

[0099] Although the technology of this disclosure is described in this example as being performed by NMS 130, the technology described herein can be performed by any other computing device, system, and / or server, and this disclosure is not limited thereto. For example, one or more computing devices configured to perform the functions of the technology of this disclosure may reside in a dedicated server or be included in any other server besides NMS 130, or may be distributed throughout network 100 and may or may not be part of NMS 130.

[0100] Figure 4 An example user equipment (UE) device 400 according to one or more technologies of this disclosure is shown. Figure 4 The example UE device 400 shown can be used to implement as described in this article. Figure 1AAny UE 148 shown and described. UE device 400 may include any type of wireless client device, and this disclosure is not limited thereto. For example, UE device 400 may include mobile devices such as smartphones, tablets or laptops, personal digital assistants (PDAs), wireless terminals, smartwatches, smart rings, or any other type of mobile or wearable device. In some examples, UE 400 may also include wired client-side devices, such as IoT devices, such as printers, security sensors or devices, environmental sensors, or any other device connected to a wired network and configured to communicate over one or more wireless networks.

[0101] UE device 400 includes a wired interface 430, wireless interfaces 420A-420C, one or more processors 406, memory 412, and user interface 410. The various components are coupled together via bus 414, through which they can exchange data and information. The wired interface 430 represents a physical network interface and includes a receiver 432 and a transmitter 434. If needed, the wired interface 430 can be used via cable (such as...). Figure 1A One of the Ethernet cables 144) directly or indirectly couples the UE400 to wired network devices (such as Ethernet cables 144) within the wired network. Figure 1A (One of the switches 146).

[0102] The first wireless interface 420A, the second wireless interface 420B, and the third wireless interface 420C each include a receiver 422A, a receiver 422B, and a receiver 422C, respectively. Each receiver includes a receiving antenna, through which the UE 400 can receive signals from a wireless communication device (such as a wireless communication device, e.g., Figure 1A AP 142, Figure 2 The AP 200, other UEs 148, or other devices configured for wireless communication receive wireless signals. The first wireless interface 420A, the second wireless interface 420B, and the third wireless interface 420C also include transmitters 424A, 424B, and 424C, respectively. Each transmitter includes a transmitting antenna, through which the UE 400 can transmit wireless signals to wireless communication devices, such as… Figure 1A AP 142, Figure 2 The AP 200, other UEs 148, and / or other devices configured for wireless communication. In some examples, the first wireless interface 420A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second wireless interface 420B may include a Bluetooth interface and / or a Bluetooth Low Energy interface. The third wireless interface 420C may include, for example, a cellular interface through which the UE device 400 can connect to a cellular network.

[0103] Processor 406 executes software instructions, such as software instructions for defining software or computer programs, which are stored in a computer-readable storage medium (such as memory 412). The computer-readable storage medium is a non-transitory computer-readable medium, such as including storage devices (e.g., disk drives or optical drives) or memories (such as flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause one or more processors 406 to perform the techniques described herein.

[0104] Memory 412 includes one or more devices configured to store programming modules and / or data associated with the operation of UE 400. For example, memory 412 may include a computer-readable storage medium, such as a non-transitory computer-readable medium, including storage devices (e.g., disk drives or optical drives) or memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause one or more processors 406 to perform the techniques described herein.

[0105] In this example, memory 412 includes operating system 440, application 442, communication module 444, configuration settings 450, and data storage 454. Communication module 444 includes program code that, when executed by processor 406, enables UE 400 to communicate using any of wired interface 430, wireless interfaces 420A-420B, and / or cellular interface 450C. Configuration settings 450 includes any device settings of UE 400 configured for each of wireless interfaces 420A-420B and / or cellular interface 420C.

[0106] Data storage 454 may include, for example, a status / error log, which includes a list of events specific to UE 400. These events may include both normal and error events at the log level based on instructions from NMS 130. Data storage 454 may store any data used and / or generated by UE 400, such as data used to calculate one or more SLE metrics or identify relevant behavioral data, which is collected by UE 400 and sent directly to NMS 130 or to any AP 142 in wireless network 106 for further transmission to NMS 130.

[0107] As described herein, UE 400 can measure network data from data storage 454 and report it to NMS 130. Network data may include event data, telemetry data, and / or other SLE-related data. Network data may include various parameters indicating the performance and / or status of the wireless network. NMS 130 can determine one or more SLE metrics and store the SLE metrics as network data 137 based on SLE-related data received from the UE or client equipment in the wireless network. Figure 1A ).

[0108] Optionally, the UE device 400 may include an NMS agent 456. The NMS agent 456 is a software agent of the NMS 130 installed on the UE 400. In some examples, the NMS agent 456 may be implemented as a software application running on the UE 400. The NMS agent 456 collects information from the UE 400, including detailed client device attributes, including insights into the UE 400's roaming behavior. This information provides insights into client roaming algorithms, as roaming is a client device decision. In some examples, the NMS agent 456 may display client device attributes on the UE 400. The NMS agent 456 sends client device attributes to the NMS 130 via an AP device connected to the UE 400. The NMS agent 456 may be integrated into a custom application or as part of a native application. The NMS agent 456 may be configured to identify the device connection type (e.g., cellular or Wi-Fi) and the corresponding signal strength. For example, the NMS agent 456 identifies the access point connection and its corresponding signal strength. NMS Proxy 456 can store information specifying the APs identified by UE 400 and their corresponding signal strengths. NMS Proxy 456 or other components of UE 400 also collect information about which APs UE 400 is connected to, and this information also indicates which APs UE 400 is not connected to. UE 400's NMS Proxy 456 sends this information to NMS 130 via the APs it is connected to. In this way, UE 400 sends not only information about UE 400 and the APs it is connected to, but also information about other APs that UE 400 identifies but is not connected to, along with their signal strengths. The APs then forward this information to the NMS, including information about other APs identified by UE 400 besides themselves. This additional level of granularity allows NMS 130 and the end network administrator to better determine the Wi-Fi experience directly from the perspective of the client devices.

[0109] In some examples, the NMS agent 456 also enriches the client device data utilized in the service level. For instance, the NMS agent 456 can go beyond basic fingerprinting to provide supplementary details such as device type, manufacturer, and different versions of the operating system. Within these detailed client attributes, the NMS 130 can display radio hardware and firmware information of the UE 400 received from the NMS client agent 456. The more detail the NMS agent 456 can derive, the stronger the VNA / AI engine's capabilities in advanced device classification. The NMS 130's VNA / AI engine continuously learns and becomes more accurate in its ability to distinguish between device-specific issues and broader device problems, such as specifically identifying which OS version is affecting certain clients.

[0110] In some examples, NMS agent 456 may cause user interface 410 to display a prompt that alerts the end user of UE 400 to enable location permission before NMS agent 456 can report device location, client information, and network connectivity data to NMS. Then, NMS agent 456 will begin reporting connectivity and location data to NMS. In this way, the end user of the client device can control whether NMS agent 456 is enabled to report client device information to NMS.

[0111] Figure 5 This is a block diagram illustrating an example network node 500 according to one or more technologies of this disclosure. In one or more examples, the network node 500 implements attachment to... Figure 1A Network devices or servers 134, such as switch 146, AAA server 110, DHCP server 116, DNS server 122, network server 128, etc., or supporting... Figure 1B Another network device, such as router 187, of one or more of the following: wireless network 106, wired LAN 175, SD-WAN 177, or data center 179.

[0112] In this example, network node 500 includes a wired interface 502 (e.g., an Ethernet interface), a processor 506, input / output 508 (e.g., a display, buttons, keyboard, keypad, touchscreen, mouse, etc.), and memory 512 coupled together via bus 514. The various components can exchange data and information via bus 514. The wired interface 502 couples network node 500 to a network, such as an enterprise network. Although only one interface is shown by way of example, a network node can (and typically does) have multiple communication interfaces and / or multiple communication interface ports. The wired interface 502 includes a receiver 520 and a transmitter 522.

[0113] Memory 512 stores executable software application 532, operating system 540, and data / information 530. Data 530 may include system logs and / or error logs storing event data (including behavioral data) of network node 500. In the example where network node 500 includes a "third-party" network device, the same entity does not own or have access to the AP or wired client-side device and network node 500. Therefore, in the example where network node 500 is a third-party network device, NMS 130 does not receive, collect, or otherwise access network data from network node 500.

[0114] In an example where network node 500 includes a server, network node 500 can receive data and information via receiver 520, such as operation-related information, such as registration requests, AAA services, DHCP requests, Simple Notification Service (SNS) lookups, and web page requests, and send data and information via transmitter 522, such as configuration information, authentication information, web page data, etc.

[0115] In examples where network node 500 includes wired network devices, network node 500 can be connected to one or more access points (APs) or other wired client-side devices (e.g., IoT devices) via wired interface 502. For example, network node 500 may include multiple wired interfaces 502 and / or wired interfaces 502 may include multiple physical ports for connection to multiple APs or other wired client-side devices within the site via appropriate Ethernet cables. In some examples, each AP or other wired client-side device connected to network node 500 can access the wired network via wired interface 502 of network node 500. In some examples, one or more APs or other wired client-side devices connected to network node 500 can each draw power from network node 500 via appropriate Ethernet cables and Power over Ethernet (PoE) ports of wired interface 502.

[0116] In an example where network node 500 includes a session-based router using a stateful, session-based routing scheme, network node 500 can be configured to perform path selection and traffic engineering independently. The use of session-based routing allows network node 500 to avoid using a centralized controller (such as an SDN controller) to perform path selection and traffic engineering, and to avoid using tunnels. In some examples, network node 500 can implement session-based routing as a Security Vector Router (SVR) provided by Juniper Networks. In an example where network node 500 includes a session-based router operating as a network gateway for a site in an enterprise network (e.g., ...), ... Figure 1BIn the case of a router 187A, network node 500 can access the underlying physical WAN (e.g., Figure 1B The SD-WAN 177) and one or more other session-based routers (e.g., those operating as network gateways for other sites in the enterprise network) Figure 1B Router 187B) establishes multiple peer paths (e.g., Figure 1B (Logical path 189). Network node 500, operating as a session-based router, can collect data at the peer path level and report peer path data to NMS 130.

[0117] In an example where network node 500 includes a packet-based router, network node 500 can employ either packet-based or flow-based routing schemes to forward packets according to network paths defined, for example, by a centralized controller that performs path selection and traffic engineering. In an example where network node 500 includes a packet-based router operating as a network gateway for an enterprise network (e.g., ...), ... Figure 1B In the case of a router 187A, network node 500 can access the underlying physical WAN (e.g., Figure 1B The SD-WAN 177) and one or more other packet-based routers (e.g., those operating as network gateways for other sites in the enterprise network) Figure 1B Router 187B) establishes multiple tunnels (e.g., Figure 1B (Logical path 189). Network node 500, operating as a packet-based router, can collect data at the tunnel level, and tunnel data can be retrieved by NMS 130 via API or open configuration protocol, or tunnel data can be reported to NMS 130 by NMS agent 544 or other modules running on network node 500.

[0118] The data collected and reported by network node 500 may include periodically reported data and event-driven data. Network node 500 is configured to collect logical path statistics via bidirectional forwarding detection (BFD) probing and data extracted from messages and / or counters at the logical path level (e.g., peer-to-peer paths or tunnels). In some examples, network node 500 is configured to collect statistics and / or sample other data according to a first periodic interval (e.g., every 3 seconds, every 5 seconds, etc.). Network node 500 may store the collected and sampled data as path data, for example, in a buffer.

[0119] In some examples, network node 500 may optionally include NMS agent 544. NMS agent 544 may periodically create statistics packets at a second periodic interval (e.g., every 3 minutes). The collected and sampled data periodically reported in the statistics packets may be referred to herein as “oc-stats”. In some examples, the statistics packets may also include details about clients connected to network node 500 and associated client sessions. NMS agent 544 may then report the statistics packets to NMS 130 in the cloud. In other examples, NMS 130 may request, retrieve, or otherwise receive statistics packets from network node 500 via API, open configuration protocols, or other communication protocols. Statistics packets created by NMS agent 544 or another module of network node 500 may include a header identifying network node 500 and statistics and data samples from each logical path from network node 500. In still other examples, NMS agent 544 reports event data to NMS 130 in the cloud in response to the occurrence of certain events at network node 500 when an event occurs. Event-driven data can be referred to as "oc-event" in this article.

[0120] Figure 6 Example features 638A, 638B of network data collected within an exemplary time window 639 according to one or more techniques of this disclosure are shown. These may be presented for illustrative purposes only. Figure 1A discuss Figure 6 .

[0121] NMS 130 can collect network data 137 to include feature values ​​of features 638A and 638B within a time window 639. Figure 6 In the example, NMS 130 can collect feature values ​​of feature 638A, which indicates the number of APs 142A. APs 142A report statistics about site 102A to NMS 130 within a time window 639, which defines the time between timestamps 27400 and 28800. NMS 130 can also collect feature values ​​of feature 638B, which indicates the number of client devices 148A connected to active APs 142A at site 102A within a time window 639, which also defines the time between timestamps 27400 and 28800.

[0122] NMS 130, or more specifically, AD and RCA module 135, can identify seasonal patterns in features 638A and 638B collected within time window 639. For example, AD and RCA module 135 can identify the seasonal patterns in features 638A and 638B as complex pattern types following stable, consistent statistical behavior. Figure 6In the example, the seasonal patterns of features 638A and 638B exhibit consistent statistical behavior. The consistent statistical behavior of feature 638A could be a linear function within time window 639. The consistent statistical behavior of feature 638B could be a complex distribution of a normally distributed sequence distributed over a specific portion of time window 639. The AD and RCA module 135 can send an indication to the AD model manager 136 that the seasonal patterns of features 638A and 638B are assigned to a complex pattern type, which can be specific to one or more sites within site 102 within time window 639.

[0123] The AD model manager 136 can select a general or fine-tuned anomaly detection model to predict the expected feature values ​​of features 638A and 638B for time window 639 based on indications of seasonal patterns in features 638A and 638B. For example, the AD model manager 136 can select a general anomaly detection model based on the complex pattern type of seasonal pattern clustering that the seasonal patterns identified for features 638A and 638B are assigned to the general anomaly detection model. In site 102, more than one site has followed... Figure 6 In the case of features of similar pattern types as shown, AD model manager 136 can assign sites associated with features 638A and 638B to a general anomaly detection model. In another example, AD model manager 136 can select a fine-tuned anomaly detection model to predict the expected feature values ​​of features 638A and 638B based on the complex pattern type of seasonal pattern clustering that maps sites associated with features 638A and 638B to a fine-tuned anomaly detection model specifically trained for those sites.

[0124] The AD model manager 136 can send instances of the selected anomaly detection model to the AD and RCA modules 135 to predict the feature values ​​of features 638A and 638B. The AD and RCA modules 135 can execute instances of the selected anomaly detection model to detect anomalies, such as... Figure 7 As described in more detail below. The AD model manager 136 can refresh the assignment of site 102A to seasonal pattern clusters based on subsequent network data indicating different seasonal patterns. The AD model manager 136 can send different anomaly detection models to the AD and RCA modules 135 based on different seasonal patterns. In this way, the NMS 130 can detect anomalies at the site based on specific seasonal patterns of features observed at site 102.

[0125] Figure 7 An example comparison 745 is shown between the actual feature value 749 of a feature used for anomaly detection according to one or more techniques of this disclosure and the predicted feature value 747 of that feature. This can be illustrated for illustrative purposes only. Figure 1ATo describe Figure 7 .

[0126] NMS 130 can be based on features (e.g., Figure 6 Anomalies are detected by comparing the predicted feature value 747 of features 638A and 638B with the actual feature value 749 of that feature. For example, NMS 130 can detect anomalies at a site based on the actual feature value 749 of a feature collected at the site and the predicted feature value 747 of that feature. NMS 130, or more specifically, AD and RCA module 135, can execute an instance of a selected anomaly detection model trained to output the predicted feature value 747 of that feature. For example, AD and RCA module 135 can execute an instance of anomaly detection model trained to generate the predicted feature value 747 based on historical network data indicating historical feature values ​​that follow pattern types similar to the actual feature value 749.

[0127] The AD and RCA module 135 can execute an instance of the anomaly detection model to generate predicted feature values ​​747. The AD and RCA module 135 can generate predicted feature values ​​747 that include the expected feature values ​​of features within a time window 739. The AD and RCA module 135 can generate a comparison 745, comparing the actual feature values ​​749 of the features collected within the time window 739 with the predicted feature values ​​747 of the features within the time window 739. The AD and RCA module 135 can generate the comparison 745 based on a time interval definition 743. Figure 7 In the example, AD and RCA module 135 can apply a time interval definition 743 specifying a scalar input window of 18 data points. AD and RCA module 135 can apply time interval definition 743 to aggregate features of network data 137 to generate a comparison 745. For example, AD and RCA module 135 can apply time interval definition 743 to generate a comparison by aggregating 18 data points of network data 137 to create feature values ​​relative to a time specified in time window 739. Figure 7 The actual characteristic value shown is 749.

[0128] The AD and RCA module 135 can detect anomalies within a time window 739 based on a comparison of the predicted feature value 747 and the actual feature value 749. For example, the AD and RCA module 135 can determine the difference between the predicted feature value 747 and the actual feature value 749 of the time window 739. Figure 7In the example, AD and RCA module 135 can determine that the difference between the predicted feature value 747 within time window 739 and the actual feature value 749 within time window 739 meets an anomaly detection threshold. For example, AD and RCA module 135 can determine that the difference between the predicted feature value 747 between timestamps 400 and 500 within time window 739 and the actual feature value 749 between timestamps 400 and 500 within time window 739 meets an anomaly detection threshold by determining that the actual feature value 749 collected within timestamps 400 and 500 differs from the predicted feature value 747 between timestamps 400 and 500 within time window 739 by at least two standard deviations. AD and RCA module 135 can use network data 137 to perform root cause analysis to determine the root cause of the detected anomaly. AD and RCA module 135 can generate recommendations based on the determined root cause and output these recommendations to the administrator, which may suggest adjustments to mitigate and / or resolve the detected anomaly.

[0129] Figure 8 This is a flowchart illustrating example operations for detecting anomalies according to one or more techniques of this disclosure. It may be used for illustrative purposes only. Figure 1A Let's discuss Figure 8 .

[0130] NMS 130 can identify seasonal patterns in the number of devices collected at a site over time (802). NMS 130 can predict the number of devices in a time window based on the seasonal patterns and the number of devices determined for one or more previous time windows (804). NMS 130 can detect anomalies during a time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices for the time window (806). NMS 130 can determine the root cause of the anomalies at the site (808).

[0131] Figure 9 This is a flowchart illustrating example operations for selecting an anomaly detection model to detect anomalies according to one or more techniques of this disclosure. It may be used for illustrative purposes only. Figure 1A Let's discuss Figure 9 .

[0132] NMS 130 can monitor real-time seasonal patterns (902) of multiple features of network data (e.g., network data 137) collected at one of multiple sites (e.g., site 102A) associated with an organization during a training period. NMS 130 can assign the site to one of two or more pattern types based on the real-time seasonal patterns at that site, where the two or more pattern types include a random pattern type, a regular pattern type, and a complex pattern type (904). NMS 130 can assign one of two or more anomaly detection models to the site based on the site's pattern type, where the anomaly detection model is associated with the site's pattern type (906). NMS 130 can use the anomaly detection model assigned to the site to detect anomalies in multiple features of the network data collected at that site (908).

[0133] The techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. The individual features described as modules, units, or components can be implemented together in an integrated logic device or individually as separate but interoperable logic devices or other hardware devices. In some cases, the individual features of an electronic circuit can be implemented as one or more integrated circuit devices, such as integrated circuit chips or chipsets.

[0134] If implemented in hardware, this disclosure may relate to devices such as processors or integrated circuit devices (such as integrated circuit chips or chipsets). Alternatively or additionally, if implemented in software or firmware, the technology may be implemented at least in part by a computer-readable data storage medium including instructions that, when executed, cause a processor to perform one or more of the methods described above. For example, the computer-readable data storage medium may store such instructions executed by a processor.

[0135] Computer-readable media can form part of a computer program product, which may include packaging material. Computer-readable media may include computer data storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. In some examples, the manufactured product may include one or more computer-readable storage media.

[0136] In some examples, computer-readable storage media may include non-transitory media. The term "non-transitory" can mean that the storage medium is not embodied in a carrier wave or propagating signal. In some examples, non-transitory storage media may store data that can change over time (e.g., in RAM or cache).

[0137] The code or instructions can be software and / or firmware executed by processing circuitry including one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described in this disclosure can be provided within a software module or a hardware module.

Claims

1. A network management system, comprising: Memory; as well as The processing circuit communicates with the memory, and the processing circuit is configured to: Identify seasonal patterns in the number of devices collected at sites over time; The number of devices in a time window is predicted based on the seasonal pattern and the number of devices determined for one or more previous time windows. Anomalies during the time window are detected based on the difference between the actual number of devices determined for that time window and the predicted number of devices for that time window; and Determine the root cause of the anomaly at the site.

2. The network management system according to claim 1, wherein, The number of devices includes the number of active client devices at the site and the number of active access point (AP) devices at the site.

3. The network management system according to claim 1, wherein, Network data is collected from multiple access point (AP) devices operating at the site.

4. The network management system according to any one of claims 1 to 3, wherein, In order to detect the anomaly during the time window, the processing circuit is configured to determine that the difference between the actual number of devices in the time window and the predicted number of devices in the time window is greater than two standard deviations.

5. The network management system according to any one of claims 1 to 3, wherein, In order to predict the number of devices within the time window, the processing circuit is configured to: Based on the seasonal pattern at the site, an anomaly detection model is selected from multiple anomaly detection models; and The anomaly detection model is used to predict the number of devices within the time window.

6. The network management system according to claim 5, wherein, The multiple anomaly detection models include: a threshold model associated with random pattern types, a baseline model associated with regular pattern types, and a fine-tuned machine learning model associated with complex pattern types.

7. The network management system according to claim 6, wherein, The processing circuit is further configured to: Machine learning models are trained based on a first set of historical network data from two or more sites across an organization to create a general machine learning model; and The general machine learning model is fine-tuned based on a second set of historical network data from the site to create the fine-tuned machine learning model for the site.

8. The network management system according to any one of claims 1 to 3, wherein, To determine the root cause, the processing circuit is configured as follows: The anomaly is detected at each of two or more sites associated with the organization, the two or more sites including the aforementioned site; and Determine the root cause of the anomaly across two or more sites associated with the organization.

9. A computer network method, comprising: The Network Management System (NMS) identifies seasonal patterns in the number of devices collected at sites over time. The NMS predicts the number of devices in a time window based on the seasonal pattern and the number of devices determined for one or more previous time windows. The NMS detects anomalies during the time window based on the difference between the actual number of devices determined for the time window and the predicted number of devices for the time window. as well as The NMS determines the root cause of the anomaly at the site.

10. The computer network method according to claim 9, wherein, The number of devices includes the number of active client devices at the site and the number of active access point (AP) devices at the site.

11. The computer network method according to claim 9, wherein, Network data is collected from multiple access point (AP) devices operating at the site.

12. The computer network method according to any one of claims 9 to 11, wherein, Detecting the anomaly during the time window includes determining that the difference between the actual number of devices and the predicted number of devices during the time window is greater than two standard deviations.

13. The computer network method according to any one of claims 9 to 11, wherein, The predicted number of devices within the time window includes: Based on the seasonal pattern at the site, an anomaly detection model is selected from multiple anomaly detection models; and The anomaly detection model is used to predict the number of devices within the time window.

14. The computer network method according to claim 13, wherein, The multiple anomaly detection models include: a threshold model associated with random pattern types, a baseline model associated with regular pattern types, and a fine-tuned machine learning model associated with complex pattern types.

15. The computer network method according to claim 14, further comprising: Train a machine learning model based on the first set of historical network data from two or more sites across an organization to create a general machine learning model; as well as The general machine learning model is fine-tuned based on a second set of historical network data from the site to create the fine-tuned machine learning model for the site.

16. The computer network method according to any one of claims 9 to 11, wherein, The root cause was identified as including: Detect the anomaly at each of two or more sites associated with the organization, the two or more sites including the site; and Determine the root cause of the anomaly across two or more sites associated with the organization.

17. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured to perform a computer network method according to any one of claims 9 to 16 or to be configured as a network management system according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Link status monitoring based on packet loss detection

    US10200264B2

  • Stateful load balancing in a stateless network

    US10277506B2

  • Network packet flow controller with extended session management

    US10432522B2

  • Intent-based analytics

    US10756983B2

  • Method for conveying AP error codes over BLE advertisements

    US10862742B2