Network range for determining root cause of network anomaly
Through the network management system (NMS), the connection event data of the access point (AP) device is summarized, and network exceptions and their root causes are identified in the computer network, which solves the problem of inefficient network fault handling in the prior art, and achieves fast and accurate network exception detection and processing.
Patent Information
- Application Number
- CN202410210003.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-02-26
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to detect network abnormalities and their root causes quickly and accurately in computer networks, resulting in inefficient network failure handling.
The connection event data of the access point (AP) device is obtained through the Network Management System (NMS), and the summary data generates summary data at multiple network-wide levels (service providers, organizations, sites), detects network exceptions and determines its root cause.
It realizes rapid identification of network exceptions and their root causes, improves the efficiency and accuracy of network failure handling, and allows the scope of network exceptions to be determined in a shorter time.
Smart Images

Figure CN120238416A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of U.S. Patent Application No. 18 / 400,613, filed on December 29, 2023, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to computer networks, and more particularly, to monitoring and troubleshooting computer networks. Background Art
[0004] Commercial premises or sites, such as offices, hospitals, airports, stadiums, or retail stores, typically install complex wireless network systems (including wireless access point (AP) networks) throughout the premises to provide wireless network services to one or more wireless client devices (or simply "clients"). An AP is a physical electronic device that uses various wireless network protocols and technologies to enable other devices to wirelessly connect to a wired network. For example, wireless local area network protocols (i.e., "WiFi") compliant with one or more Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards, Bluetooth / Bluetooth Low Energy (BLE), mesh network protocols (e.g., ZigBee), or other wireless network technologies.
[0005] Many different types of wireless client devices (e.g., laptops, smartphones, tablets, wearable devices, appliances, and Internet of Things (IoT) devices) incorporate wireless communication technologies and can be configured to connect to a wireless AP when the device is within range of a compatible wireless AP in order to access a wired network. In the case where a client device runs a cloud-based application, such as a Voice over Internet Protocol (VoIP) application, a streaming video application, a gaming application, or a video conferencing application, data is exchanged from the client device during an application session through one or more APs of the wireless network, one or more wired network devices (e.g., switches and / or routers), and one or more wide area network (WAN) devices (e.g., gateway routers) to reach a cloud-based application server. Summary of the Invention
[0006] In general, the present disclosure describes techniques for detecting network anomalies and determining the scope of the root cause of the detected network anomalies. For example, a network may include multiple access point (AP) devices. A user equipment (UE) may connect to the network through one of the multiple AP devices. A network anomaly may occur when an abnormal or unexpected number of AP devices disconnect in a way that prevents the UE device from connecting to the network to receive services. A network management system (NMS) receives connection event data indicating multiple disconnection events. Each disconnection event may correspond to one of the multiple AP devices that disconnect from the NMS. The NMS can analyze the connection event data to detect anomalies and determine the scope of the root cause of the detected anomalies.
[0007] The scope of a network anomaly may depend on the network entities associated with the root cause of the network anomaly. For example, a network may include multiple network entities. These network entities may be arranged hierarchically such that some network entities encompass larger portions of the network compared to other network entities. The network includes one or more service providers. Each of the one or more service providers provides network services, such as access to the Internet, to sites operated by one or more of a plurality of organizations, while each of the plurality of organizations operates network infrastructure for one or more of the plurality of sites. An organization may include an enterprise or other entity that has computing or network infrastructure at a site and uses a NAS device 108 at the site, a service provider that manages the NAS device 108 at the site, or other entities that participate in using or managing the NAS device 108 at the site.
[0008] Each site may include one or more of the multiple AP devices of the network. Service providers, organizations, and sites are examples of network entities that can be attributed to a network anomaly. Since a service provider can provide services to multiple organizations, each of which includes one or more sites, a network anomaly attributable to or otherwise associated with a service provider has a broader scope than a network anomaly attributable to, experienced at, or otherwise associated with a single site. Similarly, a network anomaly attributable to an organization that manages multiple sites tends to have a broader scope than a network anomaly associated with a single site. It is very beneficial for the NMS to be able to quickly determine the scope of a network anomaly in order to take appropriate remedial measures to address the network anomaly.
[0009] Connection event data can identify the AP devices associated with each of a plurality of disconnection events, the times associated with each of the plurality of disconnection events, and topology data indicating the locations of the AP devices associated with each of the plurality of disconnection events. Accordingly, the NMS can aggregate connection event data at multiple network scope levels based on time and / or location and / or association with different network entities within the network. For example, the NMS can identify disconnection events associated with each of a plurality of network entities. For example, disconnection events associated with multiple sites of a single organization may indicate a network anomaly attributable to that organization, while disconnection events associated with sites across multiple organizations may indicate a network anomaly attributable to the service provider providing network services to those sites. Each of these network entities is associated with one of a plurality of network scope levels (e.g., service provider, organization, site). The NMS can process the connection event data to detect network anomalies and determine the network scope level associated with the root cause of the network anomaly.
[0010] The techniques of the present disclosure can provide one or more improvements to computer-related fields of computer networks, which are integrated into practical applications. For example, the NMS can aggregate connection event data to indicate one or more disconnection events associated with each of a plurality of network entities over a period of time. In this way, the NMS can be allowed to process the connection event data to detect one or more network anomalies while determining the network scope level associated with the root cause of the one or more network anomalies. Compared with systems that sequentially analyze each network scope level from top to bottom or from bottom to top, the NMS can determine the scope of network anomalies in a shorter period of time. That is, aggregating connection event data to indicate disconnection events associated with each network entity allows the NMS to quickly identify the scope of network anomalies.
[0011] In one example, the NMS includes a memory and a processing circuit communicatively coupled to the memory. The processing circuit is configured to obtain connection event data of a plurality of access point (AP) devices. The connection event data indicates a plurality of disconnection events, wherein each of the plurality of disconnection events corresponds to the disconnection of one of the plurality of AP devices. The processing circuit is further configured to: generate summary data from the connection event data according to a plurality of network scope levels; detect one or more network anomalies based on the summary data; determine whether the root cause of the one or more network anomalies is associated with each of the plurality of network scope levels based on the summary data; and output an indication of the determined network scope level associated with the root cause, or perform a remedial action to address the root cause at the determined network scope level.
[0012] In another example, a method includes: a processing circuit of a network management system obtaining connection event data of a plurality of access point (AP) devices, the connection event data indicating a plurality of disconnection events, wherein each of the plurality of disconnection events corresponds to the disconnection of one of the plurality of AP devices, and wherein the processing circuit communicates with a memory of the network management system. Additionally, the method includes: the processing circuit generating summary data from the connection event data according to a plurality of network scope levels; the processing circuit detecting one or more network anomalies based on the summary data. The method further includes: the processing circuit determining whether a root cause of one or more network anomalies is associated with each of the plurality of network scope levels based on the summary data; the processing circuit outputting an indication of the determined network scope level associated with the root cause, or performing a remedial action to address the root cause at the determined network scope level.
[0013] In another example, a computer-readable storage medium includes instructions that, when executed by a processing circuit, cause the processing circuit to: obtain connection event data of a plurality of access point (AP) devices, the connection event data indicating a plurality of disconnection events, wherein each of the plurality of disconnection events corresponds to the disconnection of one of the plurality of AP devices, and wherein the processing circuit communicates with a memory of the network management system. The instructions further cause the processing circuit to: generate summary data from the connection event data according to a plurality of network scope levels; detect one or more network anomalies based on the summary data; determine whether a root cause of one or more network anomalies is associated with each of the plurality of network scope levels based on the summary data; and output an indication of the determined network scope level associated with the root cause, or perform a remedial action to address the root cause at the determined network scope level.
[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1A is a block diagram of an example network system including a network management system (NMS) in accordance with one or more techniques of the present disclosure.
[0016] Figure 1B is a block diagram showing further example details of a Figure 1A network system in accordance with one or more techniques of the present disclosure.
[0017] Figure 2 is a block diagram showing an example access point (AP) in accordance with one or more techniques of the present disclosure.
[0018] Figure 3Block diagram of an example NMS according to one or more techniques of the present disclosure.
[0019] Figure 4 Illustration of an example UE device according to one or more techniques of the present disclosure.
[0020] Figure 5 Flowchart of an example operation for detecting network anomalies and identifying root causes of detected network anomalies according to one or more techniques of the present disclosure. Detailed implementation
[0021] Figure 1A Block diagram of an example network system 100 including a network management system (NMS) 130 according to one or more techniques of the present disclosure. The example network system 100 includes a plurality of sites 102A - 102N (collectively referred to as "sites 102"), where the NMS 130 is used to manage one or more wireless networks 106A - 106N (collectively referred to as "wireless networks 106") respectively. Although in Figure 1A each site in the sites 102 is respectively shown as including a single wireless network in the wireless networks 106, in some examples, any site in the sites 102 may include multiple wireless networks, and the present disclosure is not limited in this regard.
[0022] Each site 102 includes a plurality of network access server (NAS) devices 108A to 108N (collectively referred to as "NAS devices 108"), for example, access points (APs) 142A - 1 to 142N - M (collectively referred to as "AP 142"), switches 146A to 146N (collectively referred to as "switches 146") or routers 147A to 147N (collectively referred to as "routers 147"). The NAS devices 108 may include any network infrastructure device capable of authenticating and authorizing client devices to access the enterprise network. Site 102A includes a plurality of APs 142A - 1 to 142A - M. Similarly, site 102N includes a plurality of APs 142N - 1 to 142N - M. Each AP 142 may be any type of wireless access point, including but not limited to commercial or enterprise access points, routers, or any other device connected to a wired network and capable of providing wireless network access to client devices within the site. An AP may be referred to as an "AP device" in some cases.
[0023] Each site 102A - 102N also includes a plurality of client devices, also known as user equipment devices (UEs), commonly referred to as UEs or client devices 148, representing various wireless-enabled devices within each site. For example, a plurality of UEs 148A-1 to 148A-K are currently located at site 102A. Similarly, a plurality of UEs 148N-1 to 148N-K are currently located at site 102N. Each UE 148 can be any type of wireless client device, including but not limited to mobile devices such as smart phones, tablets, or laptop computers, personal digital assistants (PDAs), wireless terminals, smart watches, smart rings, or other wearable devices. UE 148 can also include wired client devices, for example, Internet of Things (IoT) devices such as printers, security devices, environmental sensors, or any other device connected to a wired network and configured to communicate over one or more wireless networks 106.
[0024] To provide wireless network services to UE 148 and / or communicate over wireless network 106, the APs 142 and other wired client devices at site 102 are directly or indirectly connected to one or more network devices (e.g., switches, routers, or similar devices) via physical cables (e.g., Ethernet cables). In Figure 1A the example of, site 102A includes switch 146A, to which one or more APs 142A-1 to 142A-M at site 102A can be connected, and switch 146A can in turn be connected to router 147A. Similarly, site 102N includes switch 146N, to which one or more APs 142N-1 to 142N-M at site 102N can be connected, and switch 146N can in turn be connected to router 147N. Although in Figure 1A each site 102 includes one switch 146 and one router 147, in other examples, each site 102 can include more or fewer switches and / or routers. Additionally, the APs and other wired client devices at a given site can be connected to two or more switches and / or routers. In some examples, the interconnected switches and routers include a wired local area network (LAN) at site 102 that hosts wireless network 106. Further, two or more switches at a site can be interconnected and / or connected to two or more routers, and two or more routers can be interconnected and / or connected to other routers at other sites, for example, via a mesh or partial mesh topology in a hub-and-spoke architecture to form at least a portion of a wide area network (WAN).
[0025] The example network system 100 also includes various network components for providing network services within a wired network via network 134. For example, it includes an Authentication, Authorization, and Accounting (AAA) server 110 for authenticating users and / or UEs 148, a Dynamic Host Configuration Protocol (DHCP) server 116 for dynamically allocating network addresses (e.g., Internet Protocol (IP) addresses) to UEs 148 after authentication, a Domain Name System (DNS) server 122 for resolving domain names to network addresses, and an NMS 130. As Figure 1A shown, the various devices and systems of network system 100 are connected together via one or more networks 134, which may include the Internet and / or an enterprise intranet. Network 134 is shown as including an Internet Service Provider network (ISP) 129, where each network is associated, deployed, and managed by a different ISP to provide network services to organizations and individuals. Thus, each ISP 129 includes network infrastructure connected to other networks (e.g., the Internet 127) to provide network services, including Internet access. Each site 102 can obtain network services from any one or more ISPs 129.
[0026] Network anomalies, network failures, and other network problems may occur in network system 100, which may have a negative impact on the ability of UE 148 to receive services. One type of network problem that may occur is when one or more APs of AP 142 disconnect from NMS 130 and are thus unable to provide services to UE 148. This is referred to herein as a "network anomaly". A network anomaly does not always occur when one or more APs disconnect. For example, if AP 142A-1 disconnects from NMS 130, one or more other APs at site 102A may still be connected to the network and be able to provide services to the UEs of UE 148 at site 102A. However, if multiple APs at site 102 or all APs at site 102A disconnect, it may result in the UEs at site 102A being unable to receive services. This is an example of a network anomaly. It may be beneficial for NMS 130 to detect network anomalies and output information corresponding to the detected network anomalies in order to take remedial measures to resolve the network anomalies.
[0027] In some examples, the network system 100 can represent multiple network entities. These network entities can be arranged hierarchically such that lower-level network entities are associated with one or more higher-level network entities, and higher-level network entities are associated with one or more lower-level network entities. A service provider represents the network entity at the top of the hierarchy. For example, ISP 129 can provide services to one or more of a plurality of organizations. Each of the plurality of organizations can represent a network entity, including one or more sites of site 102. For example, a service provider can provide services to a corporate headquarters campus (“organization”) that includes three buildings. Each of the three buildings represents a site within site 102. Each site 102 represents a network entity that includes one or more APs of AP 142.
[0028] Site 102 is lower in the hierarchy than the organization, and the organization is lower in the hierarchy than the service provider. This is because a network anomaly or other network issue affecting the service provider can affect each organization to which the service provider provides services and thus can affect each site within site 102 that is part of an organization receiving services from the service provider experiencing the problem. When the root cause of a network anomaly is at the organization level, the network anomaly can affect one or more sites 102 corresponding to that organization without affecting sites corresponding to other organizations receiving services from the same service provider. A network anomaly can also be concentrated on a single site within site 102 without affecting other sites of the same organization or sites of other organizations. The NMS 130 can determine the scope of a network anomaly by determining the level of the network entity at which the root cause of the network anomaly lies. That is, the NMS 130 can determine whether the root cause of the network anomaly is at the network scope level at the service provider, at the network scope level at the organization, or at the network scope level at the site.
[0029] In Figure 1AIn the example of [[ID=]], the NMS 130 is a computing platform that manages the wireless network 106 at one or more sites 102. The NMS 130 can be cloud-based. As further described herein, the NMS 130 provides a set of integrated management tools and implements various techniques of the present disclosure. Generally, the NMS 130 can provide a cloud-based platform for wireless network data acquisition, monitoring, activity logging, reporting, predictive analysis, network anomaly identification, and alert generation. In some examples, the NMS 130 outputs notifications to a site or network administrator who interacts with and / or operates the management device 111, such as alerts, warnings, graphical indicators on a dashboard, log messages, text / Short Message Service (SMS) messages, email messages, etc. and / or suggestions regarding network anomalies and other network issues. Additionally, in some examples, the NMS 130 operates in response to configuration inputs received from an administrator who interacts with and / or operates the administrator device 111.
[0030] The administrator and the administrator device 111 can include IT personnel and administrator computing devices associated with one or more sites 102. The administrator device 111 can be implemented as any suitable device for presenting output and / or accepting user input. For example, the administrator device 111 can include a display. The administrator device 111 can be a computing system, such as a mobile or non-mobile computing device operated by a user and / or an administrator. According to one or more aspects of the present disclosure, the administrator device 111 can, for example, represent a workstation, a laptop or notebook computer, a desktop computer, a tablet computer, or any other computing device that can be operated by a user and / or present a user interface. The administrator device 111 can be physically separated from and / or located at a different location from the NMS 130 such that the administrator device 111 can communicate with the NMS 130 via the network 134 or other communication means.
[0031] In some examples, one or more NAS devices 108 (e.g., AP 142, switch 146, or router 147) can be connected to the edge devices 150A - 150N (collectively referred to as "edge devices 150") via a physical cable (e.g., an Ethernet cable). The edge devices 150 include cloud-managed LAN controllers. Each edge device 150 can include an internal device at one of the sites in the site 102 that communicates with the NMS 130 to extend some microservices from the NMS 130 to the internal NAS devices 108 while using the NMS 130 and its distributed software architecture for scalable and resilient operation, management, troubleshooting, and analysis.
[0032] Each network device of network system 100 (e.g., servers 110, 116, 122, AP 142, UE 148, switch 146, router 147, and any other server or device attached to or forming part of network system 100) may include a system log or error log module, where each of these network devices records the state of the network device, including normal operating states and error conditions. Throughout this disclosure, one or more network devices of network system 100 (e.g., servers 110, 116, 122, AP 142, UE 148, switch 146, and router 147) may be considered "third-party" network devices when owned and / or associated with an entity different from NMS130, such that NMS 130 does not receive, collect, or otherwise access the recorded states and other data of the third-party network devices. In some examples, edge device 150 may provide a proxy through which the recorded states and other data of the third-party network devices can be reported to NMS 130.
[0033] NMS 130 may include processing circuitry 131 and memory 132. Processing circuitry 131 may include fixed-function circuitry and / or programmable processing circuitry. Processing circuitry 131 may include any one or more of a microprocessor, a controller, a digital signal processor (DSP), a graphics processing unit (GPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or equivalent discrete or analog logic circuitry. In some examples, processing circuitry 131 may include multiple components, such as one or more microprocessors, one or more controllers, one or more DSPs, GPUs, TPUs, one or more ASICs, or one or more FPGAs, and any combination of other discrete or integrated logic circuitry, which may be physically located in one or more devices in one or more physical locations.
[0034] Processing circuitry 131 may be capable of processing instructions stored in memory 132. In some examples, memory 132 includes a computer-readable medium that includes instructions that, when executed by processing circuitry 131, cause NMS 130 and processing circuitry 131 to perform the various functions attributed to them herein. Memory 132 may include any volatile, non-volatile, magnetic, optical, or electrical medium, such as random access memory (RAM), read only memory (ROM), non-volatile RAM (NVRAM), electrically erasable programmable ROM (EEPROM), ferroelectric RAM (FRAM), dynamic random access memory (DRAM), flash memory, or any other digital medium.
[0035] Memory 132 can be configured to store a Virtual Network Assistant (VNA) 133, including a network anomaly scope model 135, network data 137, and connection event data 138. In some examples, the NMS 130 monitors network data 137 received from the wireless networks 106A through 106N at each site 102A through 102N, respectively, such as one or more Service Level Expectation (SLE) metrics. The NMS 130 can monitor network data 137 through connection sessions 128A through 128N (collectively referred to as "connection sessions 128") established with multiple NAS devices 108 at site 102, such as Transmission Control Protocol (TCP) sessions. Figure 1A Examples of connection sessions 128A and 128N are described, but each NAS device 108 can establish a separate connection session with the NMS 130. Although not shown in Figure 1A each connection session of connection sessions 128 traverses network 134, typically including one or more ISPs 129 that provide network services to the sites hosting the NAS devices 108 that have connection sessions with the NMS 130. The connection session 128 between the NAS device 108 and the NMS 130 can be established as a management path. The NAS device 108 can establish other connection sessions as data paths to one or more cloud-based applications, application servers, and / or data centers. In some examples, the NAS device can use the same path to the NMS 130 as both a management path and a data path.
[0036] The NMS 130 manages network resources, such as NAS devices 108 at each site, to deliver a high-quality wireless experience to end users, IoT devices, and clients at that site. For example, the NMS 130 may include a VNA 133, which is a virtual network assistant that implements an event handling platform for providing real-time insights and simplified troubleshooting for IT operations, and automatically takes corrective actions or provides recommendations to proactively address wireless network issues. For example, the VNA 133 may include an event handling platform that is configured to process hundreds or thousands of concurrent network data streams 137 from sensors and / or agents associated with APs 142 and / or nodes within the network 134. For example, according to various examples described herein, the VNA 133 of the NMS 130 may include an underlying analysis and network error identification engine and an alert system. The underlying analysis engine of the VNA 133 may apply historical data and models to the inbound event stream to compute assertions, such as identified anomalies or predicted occurrences of events that constitute a network error condition. Additionally, the VNA 133 may provide real-time alerts and reports to notify site or network administrators via an administrator device 111 of any predicted events, anomalies, trends, and may perform root cause analysis and automated or assisted error resolution. In some examples, the VNA 133 of the NMS 130 may apply machine learning techniques to identify the root cause of error conditions detected or predicted from the network data stream 137. If the root cause can be automatically resolved, the VNA 133 may invoke one or more corrective actions to correct the root cause of the error condition, thereby automatically improving the underlying SLE metrics and also automatically improving the user experience.
[0037] A description of further exemplary details of the operations implemented by the VNA 133 of the NMS 130 can be found in U.S. Patent No. 9,832,082, entitled "Monitoring Wireless Access Point Events," issued November 28, 2017, U.S. Publication No. US2021 / 0306201, entitled "Network System Fault Resolution Using a Machine Learning Model," issued September 30, 2021, U.S. Patent No. 10,985,969, entitled "Systems and Methods for a Virtual Network Assistant," issued April 20, 2021, U.S. Patent No. 10,958,585, entitled "Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault Detection," issued March 23, 2021, and U.S. Patent No. 10,958,537, entitled "Method for Spatio-Temporal Modeling," issued March 23, 2021, the entire contents of all of which are incorporated herein by reference.
[0038] In operation, the NMS 130 observes, collects, and / or receives network data 137, which can be in the form of data extracted from messages, counters, statistics, and the like. In Figure 1AIn the example of, NMS 130 also observes, collects, and / or receives connection event data 138 of NAS device 108. For each NAS device 108, the connection event data 138 includes one or more connection events of a connection session between the NAS device and NMS 130, e.g., connection or disconnection events, where the connection session is provided by the service provider. In some examples, the connection event data 138 indicates one or more connection events and / or one or more disconnection events corresponding to each AP of AP 142. In some cases, NMS 130 can receive the connection event data 138 from each AP of AP 142 independently of receiving the network data 137. In some examples, NMS 130 can receive the connection event data 138 for one or more time periods without receiving the network data 137 for one or more time periods. NMS 130 can process the connection event data 138 to identify network anomalies or other network characteristics without processing the network data 137. In some examples, NMS 130 can process the network data 137 and the connection event data 138 simultaneously to identify network anomalies or other network characteristics.
[0039] Regarding AP 142, connection events may occur when one or more APs of AP 142 are connected to NMS 130 through a connection session in connection session 128. For example, a connection event may occur when AP 142A-1 is connected to NMS 130 through connection session 128A. These connection sessions may represent TCP sessions or sessions of other communication protocols. For each of the multiple connection events, the connection event data 138 may include a timestamp indicating the time when the connection event occurred, the AP in AP 142 corresponding to the connection event, and topological information indicating the location of the AP corresponding to the connection event in the network. For example, the topological information may indicate the site 102 where the AP is located, the organization corresponding to the site, and the service provider serving the site. NMS 130 can store the topological information independently of the connection event data 138 and key the identifier of the AP in the connection event data 138 to the stored topological information to determine the site 102 hosting the AP, the organization corresponding to the site, and the service provider serving the site.
[0040] A disconnection event may occur when the AP in the AP 142 disconnects from the NMS 130 to end a session with one or more other devices. For example, when the connection session 128A terminates and disconnects the APs 142A-1 to 142A-M in the AP 142 from the NMS 130, a disconnection event may occur. For each disconnection event among multiple disconnection events, the connection event data 138 may include a timestamp indicating the time when the disconnection event occurred, the AP in the AP 142 corresponding to the disconnection event, and topology information indicating the location of the AP corresponding to the disconnection event in the network. The topology information may indicate the site 102 where the AP is located, the organization corresponding to the site, and the service provider providing services to the organization.
[0041] In some examples, the processing circuit 131 of the NMS 130 is configured to obtain the connection event data 138 indicating multiple disconnection events from the AP 142. Each disconnection event among the multiple disconnection events may correspond to the disconnection of the AP in the AP 142 from the NMS 130. The processing circuit 131 may be configured to apply the network anomaly scope model 135 to process the connection event data 138 to detect one or more network anomalies. In some examples, a network anomaly may correspond to one or more of the multiple disconnection events. A disconnection event is not necessarily equivalent to a network anomaly, but a large number of disconnection events of the APs 142 hosted at one or more sites 102 may be evidence of a network anomaly that interrupts the services provided to the UEs 148.
[0042] According to the techniques of the present disclosure, the processing circuit 131 may apply the network anomaly scope model 135 to determine whether the root cause of one or more network anomalies is associated with each of multiple network scope levels.
[0043] In some examples, the processing circuit 131 of the NMS 130 is configured to apply the network anomaly scope model 135 to determine whether the root cause of one or more detected network anomalies is associated with each of the network scope levels independent of other network scope levels. This means that the network anomaly scope model 135 can simultaneously determine whether the root cause is associated with each of the multiple network scope levels without some determinations depending on other determinations. For example, the network anomaly scope model 135 is configured to independently determine whether the root cause is associated with an organization without first determining whether the root cause is associated with a site.
[0044] To determine whether the root cause of one or more detected anomalies is associated with each of multiple network scope levels, the NMS 130 can aggregate connection event data 138 to organize disconnection events based on time, network entities, and scope. For example, the NMS 130 can aggregate connection event data 138 to indicate one or more disconnection events corresponding to each of one or more service providers, one or more disconnection events corresponding to each of multiple organizations, and one or more disconnection events corresponding to each of the sites in site 102. A particular disconnection event may be associated with multiple network entities. For example, a disconnection event involving disconnection of the connection session 128A to the NMS 130 of AP142A-M is associated with site 102A, at least one ISP in ISP 129 that provides network services for the site, and the organization that manages site 102A. The disconnection event is associated with a timestamp indicating the time when the disconnection event occurred. By aggregating the connection event data 138 based on network entities, network scope levels, and time, the processing circuit 131 can apply the network anomaly scope model 135 to identify the network scope levels associated with the root cause of one or more detected anomalies, which is more efficient compared to a system that does not aggregate the connection event data in this way.
[0045] When the root cause of a network anomaly is associated with a network scope level, it means that the network entities at that network scope level have encountered one or more problems, resulting in anomalies manifested throughout the network entities. The multiple network scope levels can include a service provider network scope level, an organization network scope level, and a site network scope level. The service provider network scope level corresponds to one or more service providers. The organization network scope level corresponds to one or more organizations. The site network scope level corresponds to site 102. The processing circuit 131 of the NMS 130 can apply the network anomaly scope model 135 to determine whether the root cause of one or more network anomalies is associated with the service provider network scope level, whether the root cause of one or more network anomalies is associated with the organization network scope level, and whether the root cause of one or more network anomalies is associated with the site network scope level based on the connection event data 138.
[0046] In some examples, to aggregate connection event data 138 based on network entities, network scope levels, and time, processing circuitry 131 may generate network scope information by identifying each of a plurality of disconnection events in the following connection event data 138: AP 142 in the AP corresponding to the disconnection event, station 102 in the station corresponding to the disconnection event, organization in the plurality of organizations corresponding to the disconnection event; and service provider in one or more service providers corresponding to the disconnection event. NMS 130 may receive information from AP 142 and / or organization, which indicates AP MAC address, station, organization, and AP public IP address. Based on the public IP address of the AP, NMS 130 may obtain information about the service provider that provides network services for the AP. When a disconnection event occurs in one of the connection sessions 128, the now-disconnected connection session is associated with the IP address of the AP that is the endpoint of the connection session. NMS 130 may use this IP address to look up the station, organization, and service provider of the corresponding AP. Network anomaly scope model 135 may detect one or more network anomalies based on network scope information. Network anomaly scope model 135 may determine whether the root cause of one or more network anomalies is associated with each of the plurality of network scope levels based on network scope information.
[0047] In some examples, network anomaly scope model 135 may include a transformer model. Processing circuitry 131 of NMS 300 may generate an input matrix based on connection event data 138, the matrix including a plurality of entries. Each of the plurality of entries in the input matrix includes connection event data corresponding to one of a plurality of network entities, the connection event data indicating the degree of impact of a network failure at the network entity. This data may be used as a scaling factor. The degree of impact of a network failure at the network entity corresponding to each of the plurality of entries in the input matrix is greater for network entities involving more access points. For example, a network failure of one of the ISPs 129 may affect tens of thousands of APs, while a network failure attributable to an organization may only affect thousands of APs. A network failure at one of the stations in station 102 may affect hundreds of APs. The scaling factor for a particular network entity scope may be based on the average or median of the AP disconnections at that network entity scope. In one example, to generate the input matrix, data of a plurality of organizations (e.g., 1000) and a plurality of ISPs 129 (e.g., 500) are superimposed into the entries / number of rows (e.g., 1500 entries) of the input matrix.
[0048] To detect one or more network anomalies, the processing circuit 131 is configured to apply the transformer model of the network anomaly scope model 135 to generate an output matrix based on the input matrix, the output matrix including a plurality of entries corresponding to a plurality of entries of the input matrix. Each entry of the plurality of entries of the output matrix includes a severity score, the severity score indicating the probability of a network anomaly among one or more network anomalies existing at the network entity corresponding to the entry. To detect one or more network anomalies, the processing circuit 131 is configured to determine the severity score of each entry of the plurality of entries of the output matrix based on the number of disconnection events associated with the network entity corresponding to the entry over a period of time. The processing circuit 131 may compare the severity score of each entry of the plurality of entries of the output matrix with one or more anomaly thresholds. The processing circuit 131 may detect one or more network anomalies based on comparing the severity score of each entry of the plurality of entries of the output matrix with one or more anomaly thresholds.
[0049] In some examples, each entry of the plurality of entries of the input matrix of the transformer model of the network anomaly scope model 135 further includes scope information indicating a network scope level among a plurality of network scope levels corresponding to the network entity among the plurality of network entities. The processing circuit 131 is configured to determine whether the root cause of one or more network anomalies is associated with each network scope level among the plurality of network scope levels based on the detected one or more network anomalies and the scope information of each entry of the plurality of entries of the input matrix. For example, by processing the input matrix to generate the output matrix, the network anomaly scope model 135 may be configured to determine whether the root cause of one or more network anomalies is associated with each network scope level among the plurality of network scope levels based on the severity scores associated with the network entities of each network scope level.
[0050] According to one particular implementation, the computing device is part of the NMS 130. According to other implementations, the NMS 130 may include one or more computing devices, dedicated servers, virtual machines, containers, services, or other forms of environments for performing the techniques described herein. Similarly, the computing resources and components implementing the VNA 133 may be part of the NMS 130, may be executed on other servers or execution environments, or may be distributed to nodes within the network.
[0051] Although the techniques of the present disclosure are described in this example as being performed by the NMS 130, the techniques described herein can be performed by any other computing device, system, and / or server, and the present disclosure is not limited to this aspect. For example, one or more computing devices configured to perform the functions of the techniques of the present invention may reside in a dedicated server, or be included in any other server other than the NMS 130, or may be distributed throughout the network 100 and may or may not form part of the NMS 130.
[0052] Figure 1B is a block diagram showing further example details of a Figure 1A network system according to one or more techniques of the present disclosure. In this example, Figure 1B shows the NMS 130 configured to operate according to an artificial intelligence / machine learning-based computing platform that provides comprehensive automation, insights, and assurance (Wi-Fi assurance, wired assurance, and WAN assurance), ranging from "client" (e.g., user device 148 Figure 1B connected to the wireless network 106 and the wired LAN 175 ( Figure 1B at the far left)) to "cloud", e.g., a cloud-based application service 181 hosted by computing resources within the data center 179 (
[0053] at the far right).
[0054] The NMS 130 includes a processing circuit 131 and a memory 132. The memory 132 is configured to store a VNA 133, including a network anomaly range model 135, network data 137, and connection event data 136. In some examples, the NMS 130 is configured to operate according to an artificial intelligence or a machine learning computing platform, and use the processing circuit 131 to train the network anomaly range model 135 based on training data stored in the memory 132 of the NMS 130. In some examples, the training data may include multiple sets of training connection event data. Each set of training connection event data in the multiple sets of training connection event data may indicate whether the set of training connection event data indicates a network anomaly, a network range level associated with the network anomaly, a network entity associated with the network anomaly, or any combination thereof.
[0055] As Figure 1B shown in the example of, the NMS 130 also provides configuration management, monitoring, and automated supervision of a software-defined wide area network (SD-WAN) 177 that operates as an intermediate network communicatively coupling the wireless network 106 and the wired LAN 175 to the data center 179 and cloud-based application services 181. Generally, the SD-WAN 177 provides seamless, secure, traffic engineering connections between a "spoke" router 187A of the wired LAN 175 that hosts the wireless network 106 (e.g., a branch or campus network) and a "hub" router 187B further up in the cloud stack towards the cloud-based application services 181. The SD-WAN 177 typically operates and manages an overlay network over the underlying physical wide area network (WAN) that provides connections to geographically separated customer networks. In other words, the SD-WAN 177 extends software-defined network (SDN) capabilities to the WAN and allows the network to decouple the underlying physical network infrastructure from the virtualized network infrastructure and applications, such that the network can be configured and managed in a flexible and scalable manner.
[0056] In some examples, the underlying routers of the SD-WAN 177 can implement a stateful, session-based routing scheme, where routers 187A, 187B dynamically modify the content of the original packet headers initiated by the client device 148 to direct traffic to the cloud-based application service 181 along a selected path (e.g., path 189) without the need to use tunnels and / or additional labels. In this way, routers 187A, 187B can be more efficient and scalable for large networks because the use of tunnel-less, session-based routing can enable routers 187A, 187B to achieve significant network resources by eliminating the need to perform encapsulation and decapsulation at the tunnel endpoints. Additionally, in some examples, each router 187A, 187B can independently perform path selection and traffic engineering to control the packet flow associated with each session without the need to use a centralized SDN controller for path selection and label distribution. In some examples, routers 187A, 187B implement session-based routing as Secure Vector Routing (SVR) provided by Juniper Networks, Inc.
[0057] In some examples, the NMS 130 may implement intent-based configuration and management of the network system 100, including implementing the construction, presentation, and execution of intent-driven workflows for configuring and managing devices associated with the wireless network 106, the wired LAN 175, and / or the SD-WAN 177. For example, declarative requirements express the desired configuration of network components without specifying the exact local device configuration and control flow. By leveraging declarative requirements, what should be done can be specified rather than how it should be done. Declarative requirements can be contrasted with imperative instructions that describe the exact device configuration syntax and control flow for implementing the configuration. By leveraging declarative requirements rather than imperative instructions, the burden on the user and / or user system of determining the exact device configuration required to achieve the desired result of the user / system is alleviated. For example, when leveraging various different types of devices from different vendors, it is typically difficult and burdensome to specify and manage precise imperative instructions to configure each device of the network. As new devices are added and device failures occur, the types and variety of network devices can change dynamically. Managing a cohesive network of devices from different vendors with different configuration protocols, syntax, and software versions to configure the devices is typically difficult to achieve. Thus, by only requiring the user / system to specify declarative requirements that specify the desired results applicable to various different types of devices, the management and configuration of network devices becomes more efficient. Further exemplary details and techniques of intent-based NMS are described in U.S. Patent No. 10,756,983 entitled "Intent-Based Analysis" and U.S. Patent No. 10,992,543 entitled "Automatically Generating an Intent-Based Network Model of an Existing Computer Network," both of which are incorporated herein by reference.
[0058] Figure 2 is a block diagram of an example AP 200 in accordance with one or more techniques of the present disclosure. Figure 2 The example AP 200 shown in Figure 1A can be used to implement any of the APs 142 shown and described in Figure 2 . The AP 200 may include, for example, a Wi-Fi, Bluetooth, and / or Bluetooth Low Energy (BLE) base station or any other type of wireless AP. In an example of
[0059] the first wireless interface 220A and the second wireless interface 220B represent wireless network interfaces and include a receiver 222A and a receiver 222B, respectively, each receiver including a receiving antenna through which the AP 200 can receive from wireless communication devices (e.g.,Figure 1A The radio signals of the UE 148) therein. The first radio interface 220A and the second radio interface 220B also respectively include transmitters 224A and 224B, and each transmitter includes a transmitting antenna. The AP 200 can transmit radio signals to wireless communication devices, such as Figure 1A the UE 148 therein. In some examples, the first radio interface 220A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second radio interface 220B may include a Bluetooth interface and / or a Low Energy Bluetooth (BLE) interface. Although the AP 200 includes two radio interfaces 220, in some cases, the AP 200 may include more than two or less than two radio interfaces.
[0060] The wired interface 230 represents a physical network interface, including a receiver 232 and a transmitter 234, for sending and receiving network communications, such as data packets. The wired interface 230 directly or indirectly connects the AP 200 to a wired network device (e.g., Figure 1A one of the switches 146) in the wired network through a cable (e.g., an Ethernet cable). Although the AP 200 is shown as including one wired interface 230, the AP 200 may include more than one wired interface or no wired interface.
[0061] The processing circuit 206 may include one or more programmable hardware-based processors configured to execute software instructions stored in a computer-readable storage medium (e.g., the memory 208), such as software instructions for defining software or a computer program, e.g., a non-transitory computer-readable medium including a storage device (e.g., a disk drive or an optical disc drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause the processing circuit 206 to perform the techniques described herein.
[0062] The memory 208 includes one or more devices configured to store programming modules and / or data associated with the operation of the AP 200. For example, the memory 208 may include a computer-readable storage medium, e.g., a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical disc drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause the processing circuit 206 to perform the techniques described herein.
[0063] In this example, the memory 208 stores executable software, including an Application Programming Interface (API) 240, a communication manager 242, configuration / radio settings 250, a device status log 252, data 254, and a log controller 255. The device status log 252 includes a list of events specific to the AP 200. The events can include logs of normal events and error events, such as, memory status, reboot or restart events, crash events, cloud disconnection events with self - recovery, low link speed or link speed swing events, Ethernet port status, Ethernet interface packet errors, upgrade failure events, firmware upgrade events, configuration changes, etc., along with the time and date stamps for each event. The log controller 255 determines the logging level of the device based on instructions from the NMS 130.
[0064] The data 254 can store any data used and / or generated by the AP 200. For example, the connection event data 256 can include information corresponding to one or more connection events and one or more disconnection events corresponding to the AP 200. For example, the connection event data 256 can include time stamps indicating the time of each connection event in the one or more connection events. The connection event data 256 can include time stamps indicating the time of occurrence of each disconnection event in the one or more disconnection events. For each disconnection event and each connection event, the connection event data 256 can indicate the AP 200 associated with the event. This can include indicating the AP 200 and / or indicating the location of the AP 200 in the network topology. The network data 258 can include data collected from the UE 148, for example, data used to calculate one or more SLE metrics, which are transmitted by the AP 200 for cloud - based management of the wireless network 106A by the NMS 130.
[0065] Input / Output (I / O) 210 represents physical hardware components capable of interacting with a user, such as, buttons, displays, etc. Although not shown, the memory 208 generally stores executable software for controlling a user interface regarding inputs received via the I / O 210. The communication manager 242 includes program code that, when executed by the processing circuit 206, allows the AP200 to communicate with the UE 148 and / or the network 134 via any wireless interface 220 and / or wired interface 230. The configuration settings 250 include any device settings of the access point 200, such as, radio settings for each wireless interface 220. These settings can be configured manually or can be remotely monitored and managed by the NMS 130 to optimize the wireless network performance periodically (e.g., hourly or daily).
[0066] As described herein, the AP 200 can measure network data from the device status log 252 and report it to the NMS 130. Additionally or alternatively, the AP 200 can measure and report connection event data 256 and / or network data 258 to the NMS 130. The network data collected by the device status log 252 and the network data 258 can include event data, telemetry data, and / or other SLE-related data. The network data can include various parameters indicating the performance and / or status of the wireless network. These parameters can be measured and / or determined by one or more UE devices and / or one or more APs in the wireless network. The NMS 130 can determine one or more SLE metrics based on the SLE-related data received from the APs in the wireless network and store the SLE metrics as network data 137( Figure 1A ).
[0067] Figure 3 is a block diagram of an example NMS 300 according to one or more techniques of the present disclosure. The NMS 300 can be used to implement, for example, Figures 1A to 1B the NMS 130 in. In such an example, the NMS 300 is responsible for monitoring and managing one or more wireless networks 106 at the site 102 in Figure 1A respectively. In some examples, the NMS 300 is responsible for monitoring and managing the wireless network 106 at the site 102 respectively. Additionally or alternatively, the NMS 300 is responsible for monitoring and managing the network 134, the servers 110, 116, and / or 122, or any combination thereof.
[0068] The NMS 300 includes a processing circuit 306, a memory 308, a user interface 310, a communication interface 312, and a database 318. The memory 308 is used to store the API 322 and the VNA 360 (including the network anomaly range model 362). The communication interface 312 includes a receiver 324 and a transmitter 326. The database 318 is used to store the network data 368 and the connection event data 370. The connection event data 370 includes connection event information 372, disconnection event information 374, a connection event timestamp 376, and a disconnection event timestamp 378. Although Figure 3 shows the database 318 separated from the memory 308, in some examples, the memory 308 is configured to store the database 318.
[0069] Various components are coupled together by a bus 314, and the components can exchange data and information through the bus 314. The processing circuit 306 can be an example of the processing circuit 131 in Figures 1A to 1B . The memory 308 can be an example of the memory 132 in Figures 1A to 1B . The VNA 360 can be an example of the VNA 133 in Figures 1A to 1B . The network anomaly range model 362 can beFigures 1A to 1B An example of the intermediate network anomaly range model 135. The network data 368 can be Figures 1A to 1B An example of the intermediate network data 137. The connection event data 370 can be Figures 1A to 1B An example of the intermediate connection event data 138.
[0070] In some examples, the NMS 300 receives data from one or more of the AP 142, switch 146, router 147, UE 148, router 187, and other network nodes inside the network system 100, and this data can be used to calculate one or more metrics corresponding to the network system 100. The NMS 300 can analyze this data for cloud-based monitoring and / or management of the wireless network 106 at the site 102, monitoring and / or management of the servers 110, 116, and / or 122, and monitoring and / or management of the network 134. In some examples, the NMS 300 can be Figure 1A Part of another server as shown, or part of any other server.
[0071] The processing circuit 306 can include fixed-function circuitry and / or programmable processing circuitry. The processing circuit 306 can include any one or more of a microprocessor, a controller, a DSP, a GPU, a TPU, an ASIC, an FPGA, or equivalent discrete or analog logic circuitry. In some examples, the processing circuit 306 can include multiple components, such as one or more microprocessors, one or more controllers, one or more DSPs, GPUs, TPUs, one or more ASICs, or one or more FPGAs, and any combination of other discrete or integrated logic circuitry, which can be physically located in one or more devices in one or more physical locations. The processing circuit 306 can execute software instructions stored in a computer-readable storage medium (e.g., the memory 308), e.g., software instructions that define software or a computer program, e.g., a non-transitory computer-readable medium including a storage device (e.g., a disk drive or an optical disc drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause the processing circuit 306 to perform the techniques described herein.
[0072] Processing circuit 306 may be capable of processing instructions stored in memory 308. In some examples, memory 308 includes a computer-readable medium that includes instructions that, when executed by processing circuit 306, cause NMS 300 and processing circuit 306 to perform the various functions attributed to them herein. Memory 308 may include any volatile, non-volatile, magnetic, optical, or electrical medium, such as RAM, ROM, NVRAM, EEPROM, FRAM, DRAM, flash memory, or any other digital medium. Memory 308 may include one or more devices configured to store programming modules and / or data associated with the operation of NMS 300. For example, memory 308 may include a computer-readable storage medium, such as a non-transitory computer-readable medium including a storage device (e.g., a disk drive or an optical drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause processing circuit 306 to perform the techniques described herein.
[0073] A user, such as an administrator, may interact with NMS 300 via user interface 310. User interface 310 may include a display, such as a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), or other type of screen, that processing circuit 306 may utilize to present information related to NMS 300, network devices, or other devices of network system 100. Additionally, user interface 310 may include an input organization for receiving input from the user. The input organization may include any one or more of, for example, buttons, a keypad (e.g., an alphanumeric keypad), a peripheral pointing device, a touchscreen, or another input organization that allows the user to navigate and provide input in the user interface presented by processing circuit 306 of NMS 300. In other examples, user interface 310 further includes an audio circuit for providing audible notifications, instructions, or other sounds to the user, receiving voice commands from the user, or both. Memory 308 may include instructions for operating user interface 310.
[0074] Communication interface 312 may include, for example, an Ethernet interface. Communication interface 312 couples NMS 300 to a network and / or the Internet, such as, Figure 1A any of the networks 134 shown, and / or any local area network. Communication interface 312 includes a receiver 324 and a transmitter 326, via which NMS 300 communicates to / from client device servers 110, 116, and / or 122, switch 146, router 147, UE 148, and / or forms as Figures 1A to 1BAny other network node, device, or system that is part of the network system 100 shown receives / transmits data and information. In some scenarios described herein where the network system 100 includes "third-party" network devices owned and / or associated with entities different from the NMS 300, the NMS 300 does not receive, collect, or access network data from the third-party devices.
[0075] The NMS 300 can receive connection event data 370 through the receiver 324 of the communication interface 312. The connection event data 370 can include corresponding information of multiple connection events corresponding to the AP 142 and corresponding information of multiple disconnection events corresponding to the AP 142. For example, the connection event data 370 includes connection event information 372 corresponding to multiple connection events and connection event timestamps 376. The connection event data 370 includes disconnection event information 374 corresponding to multiple disconnection events and disconnection event timestamps 378.
[0076] Each connection event among the multiple connection events corresponds to one of the APs in the AP 142, and each disconnection event among the multiple disconnection events corresponds to one of the APs in the AP 142. A connection event occurs when an AP in the AP 142 connects to the NMS 130 through the connection session 128 of the connection session. The connection session can include a session according to the TCP protocol or a session according to other communication protocols. A disconnection event occurs when an AP in the AP 142 disconnects from the NMS 130, for example, when one of the connection sessions 128 terminates.
[0077] In some examples, one or more disconnection events may be related to network anomalies that interfere with, degrade, or cut off services to one or more UEs 148. For example, when one or more APs of the AP 142 disconnect from the NMS 130, it may cause one or more UEs of the UE 148 to be unable to receive services through the network. The disconnection of the AP from the NMS 130 does not necessarily cause the UE to disconnect from the NMS 130. For example, when the UE 148A-1 is within the range of more than one AP connected to the network and one of these APs disconnects from the NMS 130, the UE 148A-1 can continue to be connected to the network by remaining connected to one or more APs that are still connected to the network. However, when each AP within the range of the UE 148A-1 disconnects, this may indicate the cutting off of one or more conditions for the UE 148A-1 to receive services, resulting in a network anomaly.
[0078] In some examples, the root cause of a network anomaly may be related to network entities such as service providers, organizations, sites, or other network entities. For example, the root cause of one or more of the APs 142A-1 to 148A-M disconnecting from the NMS 130 may be located at site 102A, and the root cause is not at a higher level in the network hierarchy, e.g., within the organization corresponding to site 102A or within the service provider providing services to site 102A. In other examples, the root cause of one or more of the APs 142A-1 to 148A-M disconnecting from the NMS 130 can be located at the organization corresponding to site 102A or at the service provider providing services to site 102A. In these examples, APs at sites other than site 102A may be affected by the network anomaly because an organization can include more than one site and a service provider can provide services to more than one organization and more than one site.
[0079] It may be beneficial for the NMS 300 to apply the network anomaly scope model 362 to detect one or more network anomalies, identify one or more network entities corresponding to each detected anomaly, and determine whether the root cause of each detected network anomaly is associated with each of multiple network scope levels. For example, it may be beneficial to determine whether the root cause of a network anomaly affecting site 102A is associated with site 102A itself, the organization corresponding to site 102A, the service provider providing services to site 102A, or other network entities. This is because, to correct a network anomaly, it may be beneficial to determine where the root cause of the problem lies.
[0080] The connection event information 372 indicates the access points of the APs 142 corresponding to each of the multiple connection events of the connection event data 370. In some examples, the connection event information 372 includes topology information indicating the location of the APs within the network corresponding to each of the multiple connection events. For example, the connection event information 372 can indicate that the connection event is associated with AP 148A-1, which is located at site 102A, and site 102A is part of an organization among multiple organizations, and a service provider among one or more service providers provides services to the organization and site 102A. In some examples, the connection event information 372 can indicate the APs 142 corresponding to each of the multiple connection events, and the database 318 stores separate topology information indicating the location of each AP 142 in the network topology.
[0081] The disconnection event information 374 indicates the AP of the AP 142 corresponding to each of the multiple disconnection events of the connection event data 370. In some examples, the disconnection event information 374 includes topology information indicating the location of the AP in the network corresponding to each of the multiple connection events. For example, the disconnection event information 374 may indicate that the disconnection event is associated with the APs 148A-M, which are located at site 102A, which is part of an organization among multiple organizations, and a service provider among one or more service providers provides services to the organization and site 102A. In some examples, the disconnection event information 374 may indicate the AP in the AP 142 corresponding to each of the multiple disconnection events, and the database 318 stores separate topology information indicating the location of each AP 142 in the network topology.
[0082] The connection event timestamp 376 may include timestamps indicating the occurrence time of each of the multiple connection events. The disconnection event timestamp 378 may include timestamps indicating the occurrence time of each of the multiple disconnection events. In some examples, the connection event timestamp 376 is saved in the database 318 as part of the connection event information 372. In some examples, the disconnection event timestamp 378 is saved in the database 318 as part of the disconnection event information 374. In some examples, the connection event timestamp 376 and the disconnection event timestamp 378 are very important for identifying network anomalies. For example, a large number of disconnection events occurring in a short period of time may indicate a network anomaly, resulting in service interruption of the UE 148.
[0083] In some examples, when a service provider failure or outage causes a network anomaly, the organizations and sites receiving services from the service provider may experience a high frequency of AP disconnection events. During normal operation, the number of AP disconnection events occurring in an organization per day may be small (e.g., 10 AP disconnection events), but a network anomaly corresponding to a service provider failure may cause a large number of APs (e.g., more than 100 APs) to disconnect at that organization. When a service provider fails, this may have a very large negative impact, and hundreds or thousands of AP disconnection events may occur at the organizations and sites receiving the services of the service provider. In an example where a service provider failure or outage causes a network anomaly, the root cause of the network anomaly is related to the network scope level of the service provider.
[0084] The root cause of a network anomaly may not necessarily be related to the network scope level of the service provider. In some cases, an organization or site may be associated with the root cause of the network anomaly. For example, when the root cause of the network anomaly is located at site 102A, the network anomaly may affect site 102A without spreading to other sites. When the root cause of the network anomaly is located at the organization corresponding to site 102A, the network anomaly may affect site 102A and other sites belonging to that organization. This shows that it may be beneficial to determine the network scope level associated with the root cause of the network anomaly so that remedial measures can be taken to resolve the network anomaly. By using the NMS 300 to determine whether the root cause of the network anomaly is associated with each of the multiple network scope levels, the network system 100 can identify the location of the root cause of the network anomaly.
[0085] The incidence of disconnection events can indicate a network anomaly. For example, the NMS 300 can determine that a network anomaly is occurring based on determining that a large number of disconnection events have occurred over a period of time or based on determining that a large number of disconnection events occur per unit time. The NMS 300 can perform a scope analysis to detect the network scope level associated with the root cause of the network anomaly. Example network scope levels include the site network scope level, the organization network scope level, the cluster network scope level, and the service provider network scope level. It may take a significant amount of time to sequentially determine whether the root cause of the network anomaly is associated with each of the multiple network scope levels one by one. For example, first determining whether the root cause is associated with the site network scope level and then determining whether the root cause is associated with the organization network scope level, and so on, may take a significant amount of time. To determine whether the root cause of the network anomaly is associated with each of the multiple network scope levels in a time - and effort - saving manner, the NMS 300 can simultaneously determine whether each of the multiple network scope levels is associated with the root cause.
[0086] The processing circuit 306 of the NMS 300 may obtain connection event data 370 indicating multiple disconnection events from the AP 142. Each disconnection event among the multiple disconnection events corresponds to one of the APs in the AP 142. The processing circuit 306 may apply the network anomaly scope model 362 to detect one or more network anomalies based on the connection event data 370. For example, the network anomaly scope model 362 may process the disconnection event timestamps 378 to identify one or more network anomalies in response to recognizing a high incidence of disconnection events within a unit time. The processing circuit 306 may apply the network anomaly scope model 362 to determine whether the root cause of one or more network anomalies is associated with each of the multiple network scope levels based on the connection event data 370. The network anomaly scope model 362 may process the connection event data 370 to simultaneously determine whether the root cause is associated with each of the multiple network scope levels.
[0087] In some examples, the network anomaly scope model 362 includes a transformer model that is configured to detect network anomalies across multiple network scope levels while attributing the detected network anomalies. For example, the network anomaly scope model 362 may determine the root cause scope of a network anomaly and generate operations and / or suggestions for resolving the network anomaly. Compared with a loop-based system, the transformer model of the network anomaly scope model 362 may allow data at different network scope levels to be stacked together and processed in a shorter time, while a loop-based system would sequentially determine whether the root cause of a network anomaly is associated with each of the multiple network scope levels. The NMS 300 may apply the transformer model of the network anomaly scope model 362 to detect network anomalies and determine the network scope level of the root cause of the detected anomalies. The NMS 300 generates information that can identify the root cause of the detected anomalies and generates suggestions for remedying the detected network anomalies. In some examples, a high incidence of disconnection events at one or more sites of the site 102 may cause the NMS 300 to perform anomaly scope detection using the transformer model of the network anomaly scope model 362.
[0088] In some examples, the processing circuit 306 of the NMS 300 can generate an input matrix including multiple entries based on the connection event data 370. Each entry among the multiple entries of the input matrix includes connection event data corresponding to one of the multiple network entities, indicating the degree of impact of a network failure at the network entity. To detect one or more network anomalies, the processing circuit 306 is configured to apply a transformer model to generate an output matrix based on the input matrix, where the output matrix includes multiple entries corresponding to the multiple entries of the input matrix. Each entry among the multiple entries of the output matrix includes a severity score, which indicates the probability of an anomaly among one or more anomalies existing at the network entity corresponding to the entry.
[0089] For example, the connection event data 370 can be aggregated into the input matrix for input into the transformer model of the network anomaly scope model 362. To aggregate the connection event data 370, the NMS 300 can place a set of connection event data for each of the multiple network entities into one of the multiple entries of the input matrix. Each of the multiple network entities can be associated with a network scope level. For example, each site of site 102 can be associated with a site network scope level. Each of the multiple organizations corresponding to site 102 can be associated with an organization network scope level. Each of one or more service providers can be associated with a service provider network scope level. One or more patterns of the connection event data 370 can indicate network anomalies associated with a specific network scope level. For example, certain patterns indicate network anomalies associated with a site, certain patterns indicate network anomalies associated with an organization, and certain patterns indicate network anomalies associated with a service provider. Compared with a system that does not aggregate input data, aggregating the connection event data 370 into the input matrix to organize data by network entity can improve the ability of the network anomaly scope model 362 to determine the network scope level associated with a network anomaly.
[0090] The transducer model of the network anomaly scope model 362 can identify different categories of network anomalies. For example, the network anomaly scope model 362 can apply a scaling factor based on the degree of fault impact at each network scope to scale data of different entity types (e.g., sites, organizations, service providers). For example, a network anomaly corresponding to a service provider may affect a large number of APs (e.g., tens of thousands of APs), a network anomaly corresponding to an organization may affect a medium number of APs (e.g., thousands of APs), and a network anomaly corresponding to a site may affect a relatively small number of APs (e.g., hundreds of APs). This indicates that the number of network anomaly disconnection events at the site network scope level may be less than the number of network anomaly disconnection events at the organization network scope level, and the number of network anomaly disconnection events at the organization network scope level may be less than the number of network anomaly disconnection events at the service provider network scope level. In some examples, the scaling factor for each network scope level can be based on the average or median of the network anomaly disconnection events at the corresponding network scope level.
[0091] In some examples, to determine the degree of impact corresponding to each of the multiple network entities to generate the input matrix of the network anomaly scope model 362, the NMS 300 is configured to determine the number of disconnection events corresponding to each of the multiple network entities over a period of time. For example, the disconnection event information 374 can indicate one or more network entities associated with each of the multiple disconnection events. The disconnection event timestamp 378 can indicate the time at which each of the multiple disconnection events occurred. This indicates that the NMS 300 is configured to determine the number of disconnection events associated with each of the multiple network entities over a period of time by analyzing the disconnection event information 374 and the disconnection event timestamp 378.
[0092] The processing circuit 306 is configured to apply the network anomaly scope model 362 to convert an input matrix including multiple entities into an output matrix including multiple entities. To detect one or more network anomalies, the processing circuit 306 is configured to apply the network anomaly scope model 362 to determine a severity score for each entry of the multiple entries of the output matrix based on the number of disconnection events associated with the network entity corresponding to the entry over a period of time. As described above, the severity can be scaled based on the network scope level of the network entity. For example, the number of disconnection events corresponding to a network anomaly at a site may be lower than the number of disconnection events corresponding to an organization, and so on. In any case, the severity score of the input matrix entry can indicate the likelihood that the network entity corresponding to the entry is associated with a network anomaly.
[0093] In some examples, the degree of impact of a network failure at a network entity corresponding to each entry in a plurality of entries of an input matrix is indicated by the number of disconnection events associated with the network entity over a period of time. To detect one or more network anomalies, the processing circuit 306 is configured to apply a network anomaly scope model 362 to determine a severity score for each entry in the plurality of entries of the output matrix based on the number of disconnection events associated with the network entity corresponding to the entry over a period of time. The more disconnection events over a period of time, the higher the severity score, and the fewer disconnection events over a period of time, the lower the severity score. To detect one or more network anomalies, the processing circuit 306 is configured to apply the network anomaly scope model 362 to compare the severity score of each entry in the plurality of entries of the output matrix with one or more anomaly thresholds. The network anomaly scope model 362 can detect one or more network anomalies by comparing the severity score of each entry in the plurality of entries of the output matrix with one or more anomaly thresholds. For example, when the severity score of a network entity exceeds a threshold, this may indicate a network anomaly associated with the network entity.
[0094] To generate the input matrix processed by the network anomaly scope model 362, the processing circuit 306 may aggregate a data set of connection event data 370 corresponding to each of the plurality of network entities. For example, when the plurality of network entities includes 2,000 sites, 1,000 organizations, and 500 service providers, the input matrix may include 3,500 entries, each entry corresponding to one of the plurality of network entities. Since a site may be part of an organization and a service provider provides services to the organization and the site, some data that is part of the entry corresponding to the site may also be part of the organization and / or service provider entries. Additionally, data that is part of the entry corresponding to the service provider may also be part of one or more entries corresponding to the organization and one or more entries corresponding to the site.
[0095] To generate the output matrix based on the input matrix, the network anomaly scope model 362 may determine a severity score corresponding to each entry in the plurality of entries of the output matrix. The network anomaly scope model 362 may apply one or more anomaly thresholds to the severity score of each entry in the plurality of entries of the output matrix to identify one or more network entities that are highly likely to have a network anomaly. For example, the anomaly threshold may be set within a range between 60% and 90% of the access points located within each network entity. For example, when 70% of the APs within site 102A are disconnected from the NMS 130 over a period of time, the network anomaly scope model 362 may determine that a network anomaly is likely to have occurred at site 102A.
[0096] In some cases, the network anomaly scope model 362 may apply the same anomaly threshold to each network entity. In other examples, the network anomaly scope model 362 may apply different anomaly thresholds to different types of network entities (e.g., apply different thresholds to each site, organization, and service provider). The network anomaly scope model 362 may attribute the root cause of the detected network anomaly to the network scope level based on one or more network entities with severity scores above the anomaly threshold. The network anomaly scope model 362 may determine one or more recommended actions to remedy the detected network anomaly. For example, applying one or more anomaly thresholds to the output matrix of the transducer may result in 10 entities being identified as including network anomalies, including 6 organizations and 4 service providers. The network anomaly scope model 362 may attribute the detected network anomaly to the root cause at the service provider network scope level.
[0097] The NMS 300 may generate one or more notifications for outputting to a network administrator (e.g., the administrator device 111) a notification that the root cause of the network anomaly is associated with a service provider among one or more service providers. In some examples, the NMS 300 may generate recommendations for resolving service provider failures or otherwise remedying network failures. In some examples, the NMS 300 may use the severity score corresponding to the detected anomaly to determine whether to generate one or more recommendations for output to the network administrator.
[0098] Although the techniques of the present disclosure are described in this example as being performed by the NMS 300, the techniques described herein may be performed by any other computing device, system, and / or server, and the present disclosure is not limited in this regard. For example, one or more computing devices configured to perform the technical functions of the present disclosure may be located in a dedicated server, or included in any other server other than the NMS 300, or in addition to the NMS 300, included in any other server, may also be distributed throughout the network, and may or may not form part of the NMS 300.
[0099] Figure 4 An example UE device 400 in accordance with one or more techniques of the present disclosure is shown. Figure 4 The example UE device 400 shown in may be used to implement as referred to herein Figure 1AAny UE 148 shown and described. The UE device 400 can include any type of wireless client device, and the present disclosure is not limited to this aspect. For example, the UE device 400 can include mobile devices such as smart phones, tablets or laptops, personal digital assistants (PDAs), wireless terminals, smart watches, smart rings or any other type of mobile or wearable device. In some examples, the UE device 400 can also include wired client devices such as IoT devices such as printers, security sensors or devices, environmental sensors or any other device connected to a wired network and configured to communicate via one or more wireless networks.
[0100] The UE device 400 includes a wired interface 430, wireless interfaces 420A to 420C (collectively referred to as "wireless interface 420"), a processing circuit 406, a memory 408, and a user interface 410. The various elements are coupled together via a bus 414, and the various elements can exchange data and information through the bus. The wired interface 430 represents a physical network interface and includes a receiver 432 and a transmitter 434. If needed, the wired interface 430 can be used to directly or indirectly couple the UE device 400 to a wired network device within a wired network via a cable (e.g., one or more Ethernet cables), such as Figure 1A a switch 146.
[0101] The first, second, and third wireless interfaces 420A, 420B, and 420C respectively include receivers 422A, 422B, and 422C, and each receiver includes a receiving antenna. The UE device 400 can receive wireless signals from a wireless communication device via the receiving antenna, such as Figure 1A AP 142 of Figure 2 AP 200 of, other UEs 148 or other devices configured for wireless communication. The first, second, and third wireless interfaces 420A, 420B, and 420C also respectively include transmitters 424A, 424B, and 424C, and each transmitter includes a transmitting antenna. The UE device 400 can transmit wireless signals to a wireless communication device via these transmitting antennas, such as Figure 1A AP 142 of Figure 2 AP 200 of, other UEs 148 and / or other devices configured for wireless communication. In some examples, the first wireless interface 420A can include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second wireless interface 420B can include a Bluetooth interface and / or a Bluetooth Low Energy interface. The third wireless interface 420C can include, for example, a cellular interface through which the UE device 400 can connect to a cellular network.
[0102] The processing circuitry 406 may include fixed function circuitry and / or programmable processing circuitry. The processing circuitry 406 may include any one or more of a microprocessor, a controller, a DSP, a GPU, a TPU, an ASIC, an FPGA, or equivalent discrete or analog logic circuitry. In some examples, the processing circuitry 406 may include multiple components, such as one or more microprocessors, one or more controllers, one or more DSPs, a GPU, a TPU, one or more ASICs, or one or more FPGAs, as well as any combination of other discrete or integrated logic circuitry, which may be physically located in one or more devices in one or more physical locations. The processing circuitry 406 executes software instructions stored in a computer-readable storage medium (e.g., memory 408), such as software instructions defining a software or computer program, e.g., a non-transitory computer-readable medium including a storage device (e.g., a disk drive or an optical disk drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause the processing circuitry 406 to perform the techniques described herein.
[0103] The processing circuitry 406 may be capable of processing instructions stored in the memory 408. In some examples, the memory 408 includes a computer-readable medium that includes instructions that, when executed by the processing circuitry 406, cause the UE device 400 and the processing circuitry 406 to perform the various functions ascribed to them herein. The memory 408 may include any volatile, non-volatile, magnetic, optical, or electrical medium, such as RAM, ROM, NVRAM, EEPROM, FRAM, DRAM, flash memory, or any other digital medium. The memory 408 includes one or more devices configured to store programming modules and / or data associated with the operation of the UE device 400. For example, the memory 408 may include a computer-readable storage medium, e.g., a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical disk drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions to cause the processing circuitry 406 to perform the techniques described herein.
[0104] A user may interact with the UE device 400 via the user interface 410. The user interface 310 may include a display such as, for example, an LCD, an LED display, an OLED display, or other types of screens, and the processing circuitry 406 may utilize the display to present information related to the UE device 400 and / or one or more services received by the UE device 400. Additionally, the user interface 410 may include an input mechanism for receiving input from the user. The input mechanism may include any one or more of, for example, buttons, a keypad (e.g., an alphanumeric keypad), a peripheral pointing device, a touch screen, or another input mechanism that allows the user to navigate and provide input in the user interface presented by the processing circuitry 406 of the UE device 400. In other examples, the user interface 410 further includes an audio circuit for providing audible notifications, instructions, or other sounds to the user, receiving voice commands from the user, or both. The memory 408 may include instructions for operating the user interface 410.
[0105] In this example, the memory 408 is configured to store an operating system 440, application programs 442, a communication module 444, configuration settings 450, and data 454. The communication module 444 includes program code that, when executed by the processing circuitry 406, causes the UE device 400 to be capable of communicating using either the wired interface 430 and / or the wireless interface 420. The configuration settings 450 include any device settings set for the UE device 400 for each of the wireless interfaces 420.
[0106] The data 454 may include, for example, a status / error log that includes a list of events specific to the UE device 400. Depending on the log level based on instructions from the NMS 130, the events may include logs of normal events and error events. The data 454 may store any data used and / or generated by the UE device 400, such as data for calculating one or more metrics or identifying relevant behavior data, which is collected by the UE device 400 and directly transmitted to the NMS 130 or transmitted to any AP 142 in the wireless network of the wireless network 106 for further transmission to the NMS 130. In some examples, the data 454 indicates one or more events in which the UE device 400 fails to connect to the network via the AP 142 or experiences a degradation in service quality via the AP 142. These one or more events may each correspond to a network anomaly. A network anomaly represents an event in which one or more APs 142 are disconnected from the NMS 130, which means that the UE 400 cannot connect to the network via the disconnected AP.
[0107] As described herein, the UE device 400 may measure and report network data from the data 454 to the NMS 130. The network data may include event data, telemetry data, and / or other data. In some examples, the network data may include data corresponding to one or more sessions between the UE device 400 and the NMS 130. For example, the UE device 400 may form one or more sessions with a service provider device such as a server of a video streaming service. The data 454 may include information corresponding to one or more sessions, the information including information indicating the quality of one or more sessions and / or data corresponding to one or more failed sessions. The network data may include various parameters indicating the performance and / or state of the wireless network.
[0108] Figure 5 is a flowchart showing example operations for detecting network anomalies and identifying root causes of the detected network anomalies according to one or more techniques of the present disclosure. This example operation is described for Figures 1A to 1B the network system 100 and its components. However, Figure 5 the techniques in
[0109] The NMS 130 may obtain connection event data 138 of multiple APs 142, and the connection event data 138 indicates multiple disconnection events (502). In some examples, each disconnection event among the multiple disconnection events corresponds to one of the APs in the AP 142. That is, each disconnection event among the multiple disconnection events may indicate that the AP in the AP 142 is disconnected from the NMS 130. For example, each connection session of the connection session 128 may connect one or more APs of the AP 142 to the NMS 130. Each site of the site 102 may receive services from the ISP in the ISP 129. Each site 102 may be part of an organization among multiple organizations. This shows that network anomalies occurring at the site, at the organization, or at the ISP may cause one or more connection sessions 128 to go offline, resulting in one or more disconnection events.
[0110] In some examples, the connection event data 138 may indicate, for each disconnection event among the multiple disconnection events, the AP corresponding to the disconnection event and the time when the disconnection event occurred. In this way, this allows the NMS 130 to summarize the connection event data 138 to indicate one or more disconnection events corresponding to each of the multiple network entities over a period of time. For example, the NMS 130 is able to determine the site associated with the AP, the organization associated with the AP, and the service provider associated with the AP based on the common IP address of the AP.
[0111] The NMS 130 can generate summary data (504) from connection event data 138 according to multiple network scope levels. In some examples, the NMS 130 can generate summary data based on network scope information. For example, the NMS 130 can generate network scope information by identifying each disconnection event among multiple disconnection events of the following connection event data 138: APs in the APS 142, sites in the site 102, organizations in multiple organizations, and ISPs in the ISP 129. For example, each disconnection event among the multiple disconnection events can be associated with the AP public IP address of the AP corresponding to the disconnection event.
[0112] The NMS 130 can detect one or more network anomalies (506) based on the summary data. In some examples, the NMS 130 applies one or more thresholds to detect one or more network anomalies. The NMS 130 can determine whether the root cause of one or more network anomalies is associated with each network scope level among the multiple network scope levels (508). The NMS 130 can output an indication of the determined network scope level associated with the root cause, or perform a remedial measure to address the root cause at the determined network scope level (510).
[0113] The techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. Various features described as modules, units, or components can be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices or other hardware devices. In some cases, various features of the electronic circuit can be implemented as one or more integrated circuit devices, for example, an integrated circuit chip or a chipset.
[0114] If implemented in hardware, the present disclosure can relate to a device, for example, a processor or an integrated circuit device, for example, an integrated circuit chip or a chipset. Alternatively or additionally, if implemented in software or firmware, the techniques can be at least partially implemented by a computer-readable data storage medium including instructions that, when executed, cause a processor to perform one or more of the above methods. For example, the computer-readable data storage medium can store such instructions executed by the processor.
[0115] The computer-readable medium can form part of a computer program product, which can include packaging materials. The computer-readable medium can include a computer data storage medium, for example, random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. In some examples, the article of manufacture can include one or more computer-readable storage media.
[0116] In some examples, a computer-readable storage medium may include a non-transitory medium. The term "non-transitory" can indicate that the storage medium is not embodied in a carrier wave or a propagated signal. In some examples, the non-transitory storage medium may store data that can change over time (e.g., in RAM or a cache).
[0117] The code or instructions can be software and / or firmware that is executed by a processing circuit, which includes one or more processors, e.g., one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described in this disclosure may be provided within software modules or hardware modules.
Claims
1. A network management system, comprising: Memory; as well as a processing circuit in communication with the memory and configured to: Acquire connection event data of a plurality of access point devices, the connection event data indicating a plurality of disconnection events, wherein each of the plurality of disconnection events corresponds to disconnection of one of the plurality of access point devices; generating aggregate data from the connection event data according to a plurality of network scope levels; detecting one or more network anomalies based on the aggregated data; determining, based on the aggregated data, whether a root cause of the one or more network anomalies is associated with each of the plurality of network-wide levels; and An indication of a determined network-wide level associated with the root cause is output, or a remedial action is performed to address the root cause at the determined network-wide level.
2. The network management system according to claim 1, wherein: The memory is configured to store a model, and wherein, to determine whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels, the processing circuit is configured to apply the model to process the aggregated data to simultaneously determine whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels.
3. The network management system according to claim 1, wherein: The multiple network range levels include: a service provider network-wide level, the service provider network-wide level comprising one or more service providers, wherein each of the one or more service providers provides network services to one or more organizations of the plurality of organizations; an organization network scope level, the organization network scope level including the plurality of organizations, wherein each organization of the plurality of organizations includes one or more sites of the plurality of sites; and A site network-wide level includes the plurality of sites, wherein each site of the plurality of sites includes one or more access point devices of the plurality of access point devices.
4. The network management system according to claim 3, wherein: The processing circuit is configured to simultaneously: determining, based on the aggregated data, whether the root cause of the one or more network anomalies is associated with the service provider network-wide level; determining, based on the aggregated data, whether the root cause of the one or more network anomalies is associated with the organizational network-wide level; as well as Based on the aggregated data, it is determined whether the root cause of the one or more network anomalies is associated with the site network-wide level.
5. The network management system according to claim 3, wherein: The processing circuit is configured to: Network-wide information is generated by identifying each of the following multiple disconnection events: an access point device among the plurality of access point devices corresponding to the disconnection event; a site among the plurality of sites corresponding to the disconnection event; an organization among the plurality of organizations corresponding to the disconnection event; as well as a service provider among the one or more service providers corresponding to the disconnection event; The summary data is generated based on the network-wide information.
6. The network management system according to claim 1, wherein: The processing circuit is further configured to generate the indication to indicate whether the root cause of the one or more network anomalies is associated with each of the plurality of network-wide levels.
7. The network management system according to any one of claims 1 to 6, wherein: The memory is configured to store a transformer model, wherein, in order to generate the summary data, the processing circuit is configured to generate an input matrix comprising a plurality of entries based on the connection event data, wherein each entry of the plurality of entries of the input matrix comprises connection event data corresponding to a network entity among a plurality of network entities, Wherein, to detect the one or more network anomalies, the processing circuit is configured to apply the transformer model to generate an output matrix based on the input matrix, the output matrix comprising a plurality of entries corresponding to the plurality of entries of the input matrix, and wherein each of the plurality of entries of the output matrix comprises a severity score indicating a probability that a network anomaly among the one or more network anomalies exists at the network entity corresponding to the entry.
8. The network management system according to claim 7, in, Each entry of the plurality of entries of the input matrix comprises connection event data corresponding to one of a plurality of network entities, the connection event data being scaled according to a scaling factor indicative of a network-wide level fault impact level for the network entity, and Wherein, in order to detect the one or more network anomalies, the processing circuit is configured to determine the severity score of each of the multiple entries of the output matrix based on a number of disconnection events associated with the network entity corresponding to the entry over a period of time.
9. The network management system according to claim 7, wherein: In order to detect the one or more network anomalies, the processing circuit is further configured to: comparing the severity score of each entry of the plurality of entries of the output matrix to one or more anomaly thresholds; as well as The one or more network anomalies are detected based on comparing the severity score of each of the plurality of entries of the output matrix to the one or more anomaly thresholds.
10. The network management system according to claim 7, in, Each entry of the plurality of entries of the input matrix further includes scope information indicating a network scope level of the plurality of network scope levels corresponding to the network entity of the plurality of network entities, and Wherein, the processing circuit is configured to determine whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels based on the detected one or more network anomalies and the scope information of each of the multiple entries of the input matrix.
11. A network management method, comprising: Acquiring, by processing circuitry of a network management system, connection event data for a plurality of access point devices, the connection event data indicating a plurality of disconnection events, wherein each of the plurality of disconnection events corresponds to a disconnection of an access point device among the plurality of access point devices, wherein the processing circuitry is in communication with a memory of the network management system; The processing circuit generates summary data from the connection event data according to a plurality of network scope levels; The processing circuitry detects one or more network anomalies based on the aggregated data; The processing circuitry determines, based on the aggregated data, whether a root cause of the one or more network anomalies is associated with each of the plurality of network-wide levels; and The processing circuitry outputs an indication of a determined network-wide level associated with the root cause or performs a remedial action to address the root cause at the determined network-wide level.
12. The method according to claim 11, wherein: The memory is configured to store a model, and wherein determining whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels includes the processing circuit applying the model to process the aggregated data to simultaneously determine whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels.
13. The method according to claim 11, wherein: The multiple network range levels include: a service provider network-wide level, the service provider network-wide level comprising one or more service providers, wherein each of the one or more service providers provides network services to one or more organizations of the plurality of organizations; an organization network scope level, the organization network scope level including the plurality of organizations, wherein each organization of the plurality of organizations includes one or more sites of the plurality of sites; and A site network-wide level includes the plurality of sites, wherein each site of the plurality of sites includes one or more access point devices of the plurality of access point devices.
14. The method according to claim 13, further comprising simultaneously: the processing circuitry determining, based on the aggregated data, whether the root cause of the one or more network anomalies is associated with the service provider network-wide level; The processing circuitry determines, based on the aggregated data, whether the root cause of the one or more network anomalies is associated with the organizational network-wide level; and The processing circuitry determines, based on the aggregated data, whether the root cause of the one or more network anomalies is associated with the site network-wide level.
15. The method according to claim 13, further comprising: The processing circuit generates network range information by identifying each disconnection event of the following plurality of disconnection events: an access point device among the plurality of access point devices corresponding to the disconnection event; a site among the plurality of sites corresponding to the disconnection event; an organization among the plurality of organizations corresponding to the disconnection event; as well as a service provider among the one or more service providers corresponding to the disconnection event; The processing circuit generates the summary data based on the network-wide information.
16. The method according to any one of claims 11 to 15, wherein: The memory is configured to store a transformer model, The generating of the summary data comprises: the processing circuit generating an input matrix comprising a plurality of entries based on the connection event data, wherein each entry of the plurality of entries of the input matrix comprises connection event data corresponding to a network entity among a plurality of network entities, Wherein, detecting the one or more network anomalies comprises: the processing circuit applies the transformer model to generate an output matrix based on the input matrix, the output matrix comprising a plurality of entries corresponding to the plurality of entries of the input matrix, and wherein each of the plurality of entries of the output matrix comprises a severity score indicating a probability of a network anomaly among the one or more network anomalies existing at the network entity corresponding to the entry.
17. The method according to claim 16, in, Each entry of the plurality of entries of the input matrix comprises connection event data corresponding to one of a plurality of network entities, the connection event data being scaled according to a scaling factor indicative of a network-wide level fault impact level for the network entity, and Wherein, detecting the one or more network anomalies comprises: the processing circuit determining the severity score for each of the plurality of entries of the output matrix based on a number of disconnection events associated with the network entity corresponding to the entry over a period of time.
18. The method according to claim 16, wherein: Detecting the one or more network anomalies includes: the processing circuitry comparing the severity score of each of the plurality of entries of the output matrix to one or more anomaly thresholds; and The processing circuitry detects the one or more network anomalies based on comparing the severity score of each of the plurality of entries of the output matrix to the one or more anomaly thresholds.
19. The method according to claim 16, in, Each entry of the plurality of entries of the input matrix further includes scope information indicating a network scope level of the plurality of network scope levels corresponding to the network entity of the plurality of network entities, and The method further includes the processing circuit determining, based on the detected one or more network anomalies and the scope information of each of the multiple entries of the input matrix, whether the root cause of the one or more network anomalies is associated with each of the multiple network scope levels.
20. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured as a network management system according to any one of claims 1 to 10 or to be configured to execute a network management method according to any one of claims 11 to 19.
Citation Information
Patent Citations
Intent-based analytics
US10756983B2
Method for spatio-temporal monitoring
US10958537B2
Methods and apparatus for facilitating fault detection and / or predictive fault detection
US10958585B2
Systems and methods for a virtual network assistant
US10985969B2
Automatically generating an intent-based network model of an existing computer network
US10992543B1