Wide area network problem prediction based on detection of service provider connection exchange
By introducing a continuous switching engine in the network management system (NMS), detecting and analyzing the connection switching situation of NAS devices between different service providers, the problem that the prior art is difficult to distinguish between WAN problems and service provider problems is solved, and effective prediction and troubleshooting of WAN problems is achieved.
Patent Information
- Application Number
- CN202411619558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-11
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to distinguish wide area network (WAN) problems from service provider problems, especially on sites that only support WiFi, where network management systems (NMSs) cannot effectively monitor and troubleshoot WAN problems.
By introducing a continuous switching engine in the Network Management System (NMS), the connection switching situation of NAS devices between different service providers is detected and the root cause of connection switching is predicted based on thresholds is the WAN problem.
It realizes effective prediction and distinction of WAN problems, improves the network management system's capabilities in troubleshooting and network optimization, and reduces dependence on specialized equipment.
Smart Images

Figure CN119996229A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims U.S. Patent Application No. 18 / 943,614, filed on November 11, 2024, and U.S. Provisional Patent Application No. 63 / 598,472, filed on November 13, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present invention relates generally to computer networks and, more particularly, to monitoring and troubleshooting computer networks. Background Art
[0004] Commercial venues or sites, such as offices, hospitals, airports, stadiums, or retail stores, typically install complex wireless network systems throughout the venue, including a network of wireless access points (APs) to provide wireless network services to one or more wireless client devices (or simply "clients"). An AP is a physical electronic device that enables other devices to wirelessly connect to a wired network using various wireless network protocols and technologies, such as a wireless local area network protocol that complies with one or more of the IEEE 802.11 standard (i.e., "WiFi"), Bluetooth / Bluetooth Low Energy (BLE), a mesh network protocol such as ZigBee, or other wireless network technologies.
[0005] Many different types of wireless client devices, such as laptops, smartphones, tablets, wearable devices, appliances, and Internet of Things (IoT) devices, incorporate wireless communication technology and can be configured to connect to a wireless access point to access a wired network when the device is within range of a compatible wireless AP. In the case of a client device running a cloud-based application, such as a Voice over Internet Protocol (VOIP) application, a streaming video application, a gaming application, or a video conferencing application, data travels from the client device through one or more APs of the wireless network, one or more wired network devices (e.g., switches and / or routers), and one or more Wide Area Network (WAN) devices (e.g., gateway routers) during an application session to reach a cloud-based application server. Summary of the invention
[0006] In general, the present disclosure describes one or more techniques for predicting wide area network (WAN) problems based on detection of network access server (NAS) devices (e.g., access point (AP) devices) that continuously switch between connections provided by different service providers. In some examples, the disclosed concepts utilize existing AP devices and WiFi-only data to infer or predict upper-level WAN problems.
[0007] According to the disclosed technology, a network management system (NMS) (i.e., a cloud-based computing platform that manages a wireless network) obtains connection event data for one or more NAS devices at a site, wherein each event in the connection event data includes a connection or disconnection event of a connection session provided by a service provider between the NAS device and the NMS. The NMS is configured to detect the number of connection exchanges within a time window. The connection exchange includes changing the connection session from a first service provider to a second service provider. Based on the number of connection exchanges detected that meet a threshold, the NMS predicts that the root cause of the connection exchange is a WAN problem. The NMS then generates a notification of the predicted root cause of the connection exchange, for example, for presentation to an administrator of the site.
[0008] The disclosed techniques provide one or more technical advantages and practical applications. A NAS device can experience disconnected connection sessions and connection events back and forth between two service providers due to problems with the WAN or the service provider itself. For sites that only support WiFi, the NMS can only see the wireless network based on WiFi data collected from the AP devices, and cannot see the wired network or WAN of these sites. As described herein, these techniques enable the NMS to distinguish between WAN problems and service provider problems, which is the root cause of NAS devices constantly switching between service providers.
[0009] In one example, the present invention is directed to an NMS, which includes a memory and one or more processors in communication with the memory. The NMS is configured to obtain connection event data of one or more NAS devices at a site, wherein each event included in the connection event data includes a connection or disconnection event of a connection session provided by a service provider between a NAS device in the one or more NAS devices and the NMS. The NMS is also configured to detect a number of connection exchanges in the connection event data within a time window, wherein the connection exchange includes a change from a first connection session provided by a first service provider to a second connection session provided by a second service provider; based on the number of connection exchanges detected that meet a threshold, predict a root cause of the connection exchange as a WAN problem; and generate a notification of the predicted root cause of the connection exchange.
[0010] In another example, the present disclosure is directed to a method including obtaining, by an NMS, connection event data of one or more NAS devices at a site, wherein each event included in the connection event data includes a connection or disconnection event of a connection session provided by a service provider between a NAS device in the one or more NAS devices and the NMS. The method also includes detecting, by the NMS, a number of connection switches in the connection event data within a time window, wherein the connection switch includes a change from a first connection session provided by a first service provider to a second connection session provided by a second service provider; predicting, by the NMS, a root cause of the connection switch as a WAN problem based on the detected number of connection switches that meet a threshold; and generating, by the NMS, a notification of the predicted root cause of the connection switch.
[0011] In another example, the present disclosure is directed to a non-transitory computer-readable storage medium including instructions that, when executed, cause one or more processors to obtain connection event data for one or more NAS devices at a site, wherein each event included in the connection event data includes a connection or disconnection event of a connection session provided by a service provider between a NAS device and an NMS in the one or more NAS devices; detect a number of connection exchanges in the connection event data within a time window, wherein the connection exchange includes a change from a first connection session provided by a first service provider to a second connection session provided by a second service provider; predict a root cause of the connection exchange as a WAN problem based on the detected number of connection exchanges that meet a threshold; and generate a notification of the predicted root cause of the connection exchange.
[0012] The details of one or more examples of the technology of the present disclosure are set forth in the accompanying drawings and the following detailed description. Other features, objects, and advantages of the technology will be apparent from the detailed description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1A is a block diagram of an example network system including a network management system, in accordance with one or more techniques of this disclosure.
[0014] Figure 1B It is shown Figure 1A A block diagram of a network system with further example details.
[0015] Figure 2 is a block diagram of an example access point device in accordance with one or more techniques of this disclosure.
[0016] Figure 3 is a block diagram of an example network management system in accordance with one or more techniques of this disclosure.
[0017] Figure 4is a block diagram of an example user equipment device in accordance with one or more techniques of this disclosure.
[0018] Figure 5 is a block diagram of an example network node, such as a router or a switch, in accordance with one or more techniques of this disclosure.
[0019] Figure 6 is a flow diagram illustrating example operation of a network management system in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION
[0020] Figure 1A 1 is a block diagram of an example network system 100 including a network management system (NMS) 130 according to one or more techniques of the present disclosure. The example network system 100 includes a plurality of sites 102A-102N, where a network service provider manages one or more wireless networks 106A-106N, respectively. Figure 1A , each site 102A-102N is illustrated as including a single wireless network 106A-106N, respectively, but in some examples each site 102A-102N may include multiple wireless networks, and the present disclosure is not limited in this respect.
[0021] Each site 102A-102N includes multiple client devices, also known as user equipment devices (UE), which are generally referred to as UE or client devices 148, representing various wireless-enabled devices within each site. For example, multiple UEs 148A-1 to 148A-K are currently located at site 102A. Similarly, multiple UEs 148N-1 to 148N-K are currently located at site 102N. Each UE 148 can be any type of wireless client device, including but not limited to mobile devices such as smartphones, tablets or laptops, personal digital assistants (PDAs), wireless terminals, smart watches, smart rings, or other wearable devices. UE 148 may also include wired client devices, for example, IoT devices such as printers, security devices, environmental sensors, or any other devices connected to a wired network and configured to communicate via one or more wireless networks 106.
[0022] Each site 102A-102N includes a plurality of network access server (NAS) devices 108A-108N, such as access points (APs) 142, switches 146, or routers 147. NAS devices 108 may include any network infrastructure device capable of authenticating and authorizing client devices to access an enterprise network. For example, site 102A includes a plurality of APs 142A-1 to 142A-M. Similarly, site 102N includes a plurality of APs 142N-1 to 142N-M. Each AP 142 may be any type of wireless access point, including but not limited to a commercial or enterprise AP, a router, or any other device connected to a wired network and capable of providing wireless network access to client devices within the site.
[0023] Each site 102A-102N also includes at least one of the gateway devices 187. Each gateway device 187 is located at the boundary of its respective site 102 and is configured to connect one or more networks at its respective site 102 to one or more networks 134 (e.g., the Internet and / or an intranet) via a service provider network ("SP") 160. For example, site 102A includes a gateway device 187A that connects site 102A to SPs 160A, 160B to access network(s) 134. Similarly, site 102N includes a gateway device 187N that connects site 102A to SPs 160A, 160B to access network(s) 134. In some examples, gateway device 187 may include router functionality to route traffic to and from NAS device 108 through network(s) 134. Each of SPs 160 may include a different service provider network that includes means for accessing the Internet and telecommunication lines. As one example, SP 160A may include a broadband service provider that uses a wired connection (such as fiber optic, cable, or telephone line) to connect gateway device 187 to network(s) 134. As another example, SP 160B may include a long term evolution (LTE) service provider or a 5G service provider that uses radio waves to connect gateway device 187 to network(s) 134. In other examples, SP 160 may include other types of service providers that use wired, wireless, or cellular connections.
[0024] To provide wireless network services to UE 148 and / or to communicate over wireless network 106, AP 142 and other wired client devices at site 102 are connected directly or indirectly to one or more network devices (e.g., switches, routers, etc.) via physical cables (e.g., Ethernet cables). Figure 1AIn the example of , site 102A includes switch 146A, to which one or more of APs 142A-1 through 142A-M at site 102A can be connected, and switch 146A can in turn be connected to router 147A. Switch 146A and / or router 147A can be connected to gateway device 187A. Similarly, site 102N includes switch 146N, to which one or more of APs 142N-1 through 142N-M at site 102N can be connected, and switch 146N can in turn be connected to router 147N. Switch 146N and / or router 147N can be connected to gateway device 187N. Although in Figure 1A 146 and a single router 147, but in other examples, each site 102 may include more or fewer switches and / or routers. In addition, the APs and other wired client devices of a given site may be connected to two or more switches and / or routers.
[0025] In some examples, the interconnected switches and routers comprise a wired local area network (LAN) at site 102 that hosts wireless network 106. Gateway devices 187 at site 102 can connect LANs to each other via one or more networks 134 (e.g., the Internet and / or a corporate intranet). In addition, two or more switches at a site can be connected to each other and / or to two or more routers, and two or more routers can be connected to each other and / or to a gateway that is connected to other gateways at other sites (e.g., via a mesh or partially meshed topology in a hub and spoke architecture), forming at least a portion of a wide area network (WAN).
[0026] The example network system 100 also includes various network components for providing network services within the wired network, including, for example, an authentication, authorization, and accounting (AAA) server 110 for authenticating users and / or UEs 148, a dynamic host configuration protocol (DHCP) server 116 for dynamically assigning a network address (e.g., an IP address) to the UEs 148 after authentication, a domain name system (DNS) server 122 for resolving domain names into network addresses, a plurality of servers 128A-128X (collectively referred to as “servers 128”) (e.g., network servers, database servers, file servers, etc.), and a network management system (NMS) 130. Figure 1A As shown, the various devices and systems of network 100 are coupled together via one or more networks 134 (eg, the Internet and / or a corporate intranet).
[0027] exist Figure 1AIn some examples, NMS 130 is a cloud-based computing platform for managing wireless networks 106A-106N at one or more of sites 102A-102N. As further described herein, NMS 130 provides a set of integrated management tools and implements various technologies of the present disclosure. In general, NMS 130 can provide a cloud-based platform for wireless network data collection, monitoring, activity logging, reporting, predictive analysis, network anomaly identification, and alarm generation. In some examples, NMS 130 outputs notifications, such as alarms, alerts, graphical indicators on dashboards, log messages, text / SMS messages, email messages, etc., and / or suggestions about wireless network problems to site or network administrators ("administrators") interacting with and / or operating management devices 111. In addition, in some examples, NMS 130 operates in response to configuration inputs received from administrators interacting with and / or operating management devices 111.
[0028] Administrators and management devices 111 may include IT personnel and administrator computing devices associated with one or more sites 102. Management device 111 may be implemented as any suitable device for presenting output and / or accepting user input. For example, management device 111 may include a display. Management device 111 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or administrator. Management device 111 may, for example, represent a workstation, a laptop or notebook computer, a desktop computer, a tablet computer, or any other computing device that can be operated by a user and / or present a user interface according to one or more aspects of the present disclosure. Management device 111 may be physically separated from NMS 130 and / or located at a different location from NMS 130, so that management device 111 can communicate with NMS 130 via network 134 or other communication means.
[0029] In some examples, one or more of the NAS devices 108, such as the AP 142, the switch 146, or the router 147, can be connected to the edge devices 150A-150N via a physical cable (e.g., an Ethernet cable). The edge device 150 includes a cloud-managed wireless LAN controller. Each edge device 150 may include a local device at the site 102 that communicates with the NMS 130 to extend certain microservices from the NMS 130 to the local NAS device 108, while using the NMS 130 and its distributed software architecture for scalable and resilient operations, management, troubleshooting, and analysis.
[0030] Each network device of the network system 100, such as servers 110, 116, 122 and / or 128, AP 142, UE 148, switch 146, router 147, and any other server or device connected to or forming part of the network system 100, may include a system log or error log module, wherein each of these network devices records the state of the network device, including normal operating states and error conditions. In the present disclosure, one or more network devices of the network system 100, such as servers 110, 116, 122 and / or 128, AP 142, UE 148, switch 146, and router 147, when owned and / or associated with a different entity than NMS 130, may be considered a "third party" network device, such that NMS 130 cannot receive, collect, or otherwise access the recorded state and other data of the third party network device. In some examples, edge device 150 may provide an agent through which the recorded state and other data of the third party network device may be reported to NMS 130.
[0031] In some examples, the NMS 130 obtains network data 137 of the NAS devices 108A-108N at each site 102A-102N, respectively. The network data 137 may include event data, telemetry data, and / or other service level expectation (SLE) related data. The network data 137 may include various parameters indicating the performance and / or status of the wireless network 106A-106N. The NMS 130 may obtain the network data 137 via a connection session (e.g., a transmission control protocol (TCP) session) established with multiple NAS devices 108 at the site 102. The connection session between the NAS device 108 and the NMS 130 may be established as a management path. The NAS device 108 may establish other connection sessions as data paths to one or more cloud-based applications, application servers, and / or data centers. In some examples, the NAS device may use the same path to the NMS 130 as a management path and a data path.
[0032] The connection session between NAS device 108 and NMS 130 may be provided by one or more service providers (e.g., SP 160). The connection session may be established through physical devices and cables, such as switch 146, router 147, and gateway device 187, which enable AP device 142 to access network(s) 134 and, therefore, NMS 130. Figure 1AIn the example shown, the AP 142A-1 at the site 102A has a first connection session 162A with the NMS 130 provided by the first SP 160A (e.g., a broadband service provider), and a second connection session 162B with the NM 130 provided by the second SP 160B (e.g., an LTE service provider). Only one of the connection sessions 162A, 162B can be used or active at a time.
[0033] The NMS 130 manages network resources, such as NAS devices 108 at each site, to provide a high-quality wireless experience to end users, IoT devices, and clients at the site. For example, the NMS 130 may include a virtual network assistant (VNA) 133 that implements an event processing platform for providing real-time insights and simplified troubleshooting for IT operations, and automatically takes corrective actions or provides recommendations to proactively resolve wireless network issues. For example, the VNA 133 may include an event processing platform that is configured to process hundreds or thousands of concurrent network data streams 137 from sensors and / or agents associated with APs 142 and / or nodes within the network 134. For example, according to various examples described herein, the VNA 133 of the NMS 130 may include an underlying analysis and network error identification engine and an alarm system. The underlying analysis engine of the VNA 133 may apply historical data and models to inbound event streams to compute assertions, such as identified anomalies or predicted occurrences of events that constitute network error conditions. In addition, the VNA 133 can provide real-time alerts and reports to notify site or network administrators of any predicted events, anomalies, trends through the management device 111, and can perform root cause analysis and automatic or assisted error remediation. In some examples, the VNA 133 of the NMS 130 can apply machine learning techniques to identify the root cause of the error condition detected or predicted from the network data stream 137. If the root cause can be automatically resolved, the VNA 133 can invoke one or more corrective actions to correct the root cause of the error condition, thereby automatically improving the underlying SLE metric and automatically improving the user experience.
[0034] Further example details of the operations implemented by the VNA 133 of the NMS 130 are disclosed in U.S. Patent No. 9,832,082, entitled “Monitoring Wireless Access Point Events,” issued on November 28, 2017, U.S. Patent No. 2021 / 0306201, entitled “Network System Fault Resolution Using a Machine Learning Model,” issued on September 30, 2021, U.S. Patent No. 10,985,969, entitled “Systems and Methods for aVirtual Network Assistant,” issued on April 20, 2021, U.S. Patent No. 10,958,585, entitled “Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault Detection,” issued on March 23, 2021, and U.S. Patent No. 10,958,585, entitled “Method for Spatio-Temporal Fault Detection and / or Predictive Fault Detection,” issued on February 23, 2021. Modeling” and U.S. Patent 10,958,537, entitled “Method for Conveying AP Error Codes Over BLE Advertisements” issued on December 8, 2020, all of which are incorporated herein by reference in their entirety.
[0035] In operation, NMS 130 observes, collects and / or receives network data 137, which may take the form of data extracted from messages, counters, and statistics, for example. Figure 1A In the example of , NMS 130 also observes, collects and / or receives connection event data 136 of NAS devices 108. For each NAS device 108, the connection event data 136 includes one or more connection events of a connection session between the NAS device and NMS 130, such as a connection or disconnection event, where the connection session is provided by a service provider. According to one specific implementation, the computing device is part of NMS 130. According to other implementations, NMS 130 may include one or more computing devices, dedicated servers, virtual machines, containers, services, or other forms of environments for executing the techniques described herein. Similarly, the computing resources and components that implement VNA 133 may be part of NMS 130, may be executed on other servers or execution environments, or may be distributed to nodes (e.g., routers, switches, controllers, gateways, etc.) within (multiple) networks 134.
[0036] According to one or more techniques of the present disclosure, NMS 130 includes a continuous switching engine 135 configured to predict WAN problems based on detecting that NAS device 108 continuously switches between connections provided by different service providers. In some examples, the disclosed concepts utilize existing AP devices 142 and WiFi-only data to infer or predict upper-level WAN problems.
[0037] According to the disclosed technology, the NMS 130 obtains connection event data 136 of a NAS device at a site (e.g., the AP device 142A at the site 102A). Each event in the connection event data 136 includes a connection or disconnection event of a connection session provided by a service provider between the NAS device 108 and the NMS 130, such as one connection session 162A, 162B provided by the SPs 160A, 160B, respectively, between the AP 142A-1 and the NMS 130. For example, the AP 142A-1 may experience a connection exchange during which a currently active connection session (e.g., a first connection session 162A provided by the first SP 160A) is disconnected, and another connection session (e.g., a second connection session 162B provided by the second SP 160B) is connected to the NMS 130 as an active connection session. For each of the disconnection of the first connection session 162A and the connection of the second connection event 162B, the AP 142A-1 collects event data and reports it to the NMS 130.
[0038] The continuous switching engine 135 of the NMS 130 is configured to detect the number of connection switches that the NAS device experiences at the site within a time window. For example, in some cases, the AP 142A-1 may experience repeated connection switches during which the connection with the NMS 130 changes back and forth between a first connection session 162A provided by the SP 160A and a second connection session 162B provided by the SP 160B.
[0039] Based on the number of detected connection exchanges that meet the threshold, the continuous exchange engine 135 of the NMS 130 predicts that the root cause of the connection exchange is a WAN problem. For example, during the time window, the number of connection exchanges from the first connection session 162A to the second connection session 162B is relatively small, for example, less than three times in one hour, which may be due to a problem with the first SP 160A or a problem with the underlying physical devices and cables associated with the first connection session 62A. However, during the time window, repeated connection exchanges between the first connection session 162A and the second connection session 162B, for example, three or more exchanges in one hour, may be caused by a problem with the underlying physical devices and cables associated with the two connection sessions 162A, 162B, such as the gateway device 187A or other devices in the WAN.
[0040] The continuous exchange engine 135 monitors the connection exchange of connection sessions from multiple APs 142A and / or NAS devices 108A at the site 102A, which all use the same gateway device 187A to access (multiple) networks 134, such as the Internet and the NMS 130. For example, to predict a WAN problem, the continuous exchange engine 352 can determine that all or most of the multiple AP devices 142A at the site 102A are experiencing a continuous exchange of connection sessions between the SPs 160. Based on the predicted WAN problem, the continuous exchange engine 135 of the NMS 130 generates a notification of the predicted root cause of the connection exchange. The NMS 130 can send the notification for presentation to an administrator of the site, such as the administrator device 111.
[0041] The techniques of the present disclosure provide one or more technical advantages and practical applications. A NAS device 108 may experience disconnected connection sessions and connection events between two service providers (e.g., SP 160A, 160B) due to issues with the WAN or the service provider itself. For sites that only support WiFi, the NMS can only view the wireless network based on WiFi data collected from AP devices, and cannot view the wired network or WAN of these sites. As described herein, these techniques enable the NMS 130 to infer or predict WAN issues, rather than service provider issues, as the root cause of connection exchanges between service providers based on WiFi-only data.
[0042] Although the techniques of the present disclosure are described in this example as being performed by NMS 130, the techniques described herein may be performed by any other computing device(s), system(s), and / or server(s), and the present disclosure is not limited in this regard. For example, one or more computing devices configured to perform the functions of the techniques of the present disclosure may reside in a dedicated server, or may be included in any other server outside of NMS 130, may be distributed throughout network 100, and may or may not form part of NMS 130.
[0043] Figure 1B It is shown Figure 1A A block diagram of a network system with further example details is provided. In this example, Figure 1B NMS 130 is shown, which is configured to operate according to an artificial intelligence / machine learning based computing platform that provides services from "clients" (e.g., user devices 148 (connected to wireless network 106 and wired LAN 175) Figure 1B ) to the “cloud” (e.g., cloud-based application services 181), which can be hosted by computing resources within a data center 179 ( Figure 1B on the far right of the ).
[0044] As described herein, NMS 130 provides an integrated set of management tools and implements various technologies of the present disclosure. In general, NMS 130 can provide a cloud-based platform for wireless network data collection, monitoring, activity logging, reporting, predictive analysis, network anomaly identification, and alarm generation. For example, the network management system 130 can be configured to actively monitor and adaptively configure the network 100 to provide autonomous driving capabilities. In addition, VNA 133 includes a natural language processing engine to provide AI-driven support and troubleshooting, anomaly detection, AI-driven location services, and AI-driven radio frequency (RF) optimization with reinforcement learning.
[0045] like Figure 1BAs shown in the example of , the AI-driven NMS 130 also provides configuration management, monitoring, and automated supervision of a software-defined wide area network (SD-WAN) 177, which operates as an intermediate network that communicatively couples the wireless network 106 and the wired LAN 175 to the data center 179 and application services 181. Generally speaking, the SD-WAN 177 provides seamless, secure, traffic-engineered connectivity between the "spoke" routers or gateways 187A of the wired network 175 that hosts the wireless network 106 (such as a branch or campus network) and the "hub" routers or portals 187B further up the cloud stack toward the cloud-based application services 181. The SD-WAN 177 typically operates and manages an overlay network on an underlying physical wide area network (WAN) that provides connectivity to geographically independent customer networks. In other words, the SD-WAN 177 extends software-defined networking (SDN) capabilities to the WAN and allows (multiple) networks to decouple the underlying physical network infrastructure from the virtualized network infrastructure and applications, so that the network can be configured and managed in a flexible and scalable manner.
[0046] In some examples, the underlying routers of the SD-WAN 177 can implement a stateful, session-based routing scheme in which the router or gateway 187A, 187B dynamically modifies the contents of the original packet header from the client device 148 to direct the traffic along a selected path (e.g., path 189) to the application service 181 without the use of tunnels and / or additional labels. In this way, the router or gateway 187A, 187B may be more efficient and scalable for large networks because the use of tunnel-free, session-based routing can enable the router or gateway 1873, 187B to realize considerable network resources by eliminating the need to perform encapsulation and decapsulation at the tunnel endpoints. In addition, in some examples, each router or gateway 187A, 187B can independently perform path selection and traffic engineering to control the packet flow associated with each session without the need to use a centralized SDN controller for path selection and label distribution. In some examples, the router or gateway 187A, 187B implements session-based routing as Security Vector Routing (SVR) provided by Juniper Networks, Inc.
[0047] Additional information about session-based routing and SVR is available in U.S. Patents 9,729,439, entitled “COMPUTER NETWORK PACKET FLOW CONTROLLER,” issued on August 8, 2017; 9,729,682, entitled “NETWORK DEVICE AND METHOD FOR PROCESSING A SESSION USING A PACKET SIGNATURE,” issued on August 8, 2017; 9,762,485, entitled “NETWORK PACKET FLOW CONTROLLERWITH EXTENDED SESSION MANAGEMENT,” issued on September 12, 2017; 9,871,748, entitled “ROUTER WITH OPTIMIZED STATISTICAL FUNCTIONALITY,” issued on January 16, 2018; and 9,871,748, entitled “NAME-BASED ROUTING SYSTEM AND METHOD FOR PROCESSING A SESSION USING A PACKET SIGNATURE,” issued on May 29, 2018. No. 9,985,883, entitled “LINK STATUS MONITORING BASED ON PACKET LOSS DETECTION”, issued on February 5, 2019; No. 10,200,264, entitled “LINK STATUS MONITORING BASED ON PACKET LOSS DETECTION”, issued on April 30, 2019; No. 10,277,506, entitled “STATEFUL LOADBALANCING IN A STATELESS NETWORK”, issued on October 1, 2019; and No. 11,075,824, entitled “IN-LINE PERFORMANCE MONITORING”, issued on July 27, 2021, the entire contents of which are incorporated herein by reference.
[0048] In some examples, the AI-driven NMS 130 can implement intent-based configuration and management of the network system 100, including implementing the construction, presentation, and execution of intent-driven workflows for configuring and managing devices associated with the wireless network 106, the wired LAN network 175, and / or the SD-WAN 177. For example, declarative requirements express the desired configuration of network components without specifying the exact local device configuration and control flow. By using declarative requirements, it is possible to specify what should be done, rather than how it should be done. Declarative requirements may be contrasted with imperative instructions that describe the exact device configuration syntax and control flow to implement the configuration. By utilizing declarative requirements rather than imperative instructions, the user and / or user system is relieved of the burden of determining the exact device configuration required to achieve the user / system desired result. For example, when using a variety of different types of devices from different vendors, it is often difficult and cumbersome to specify and manage the exact command instructions for configuring each device of the network. The types and kinds of network devices may change dynamically as new devices are added and device failures occur. It is often difficult to manage a variety of different types of devices from different vendors with different configuration protocols, syntax, and software versions to configure a cohesive device network. Thus, by requiring only the user / system to specify declarative requirements that specify desired outcomes applicable to a variety of different types of devices, management and configuration of network devices becomes more efficient. Further example details and techniques of intent-based network management systems are described in U.S. Patent No. 10,756,983, entitled “Intent-based Analytics,” and U.S. Utility Patent No. 10,992,543, entitled “Automatically generating an intent-based network model of an existing computer network,” both of which are incorporated herein by reference.
[0049] According to the disclosed technology, the NMS 130 obtains connection event data 136 of a NAS device at a site associated with a wireless network (e.g., the wireless network 106), wherein each event in the connection event data 36 includes a connection or disconnection event of a connection session between the NAS device and the NMS 130 provided by a service provider. The continuous exchange engine 135 of the NMS 130 is configured to detect the number of connection exchanges within a time window, wherein the connection exchange includes a change of the connection session from a first service provider to a second service provider. Based on the number of connection exchanges detected that meet a threshold, the continuous exchange engine 135 of the NMS 130 predicts that the root cause of the connection exchange is a WAN problem, and generates a notification of the predicted root cause of the connection exchange. In some examples, the NMS 130 may only see the wireless network 106. For example, the NMS 130 may only obtain network data 137 of AP devices and / or client devices 148 of the wireless network 106, and may not obtain network data from devices (e.g., switches, routers, and / or gateways) of the wired network 175 or the WAN 177. As described herein, these techniques enable the NMS 130 to infer or predict problems occurring within the WAN 177 rather than service provider issues, which is the root cause of the continuous exchange of NAS devices between service providers based on WiFi-only data.
[0050] Figure 2 is a block diagram of an example access point (AP) device 200 , in accordance with one or more techniques of this disclosure. Figure 2 The example access point 200 shown in FIG. 1 may be used to implement the Figure 1A Any of the APs 142 shown and described. Access point 200 may include, for example, a Wi-Fi, Bluetooth, and / or Bluetooth Low Energy (BLE) base station or any other type of wireless access point.
[0051] exist Figure 2 In the example of FIG. 1 , access point 200 includes a wired interface 230, wireless interfaces 220A-220B, one or more processors 206, memory 212, and input / output 210, which are coupled together via a bus 214, and the various components can exchange data and information on the bus 214. The wired interface 230 represents a physical network interface and includes a receiver 232 and a transmitter 234 for sending and receiving network communications (e.g., packets). The wired interface 230 directly or indirectly couples the access point 200 to a wired network device within a wired network, such as a wireless network device, via a cable (e.g., an Ethernet cable). Figure 1A One of the switches 146.
[0052] The first and second wireless interfaces 220A and 220B represent wireless network interfaces and include receivers 222A and 222B, respectively. Each receiver includes a receiving antenna through which the access point 200 can receive signals from wireless communication devices (such as Figure 1A The first and second wireless interfaces 220A and 220B further include transmitters 224A and 224B, respectively, each of which includes a transmitting antenna, through which the access point 200 can send wireless signals to wireless communication devices (such as UE 148). Figure 1A In some examples, the first wireless interface 220A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second wireless interface 220B may include a Bluetooth interface and / or a Bluetooth low energy (BLE) interface.
[0053] The processor(s) 206 are programmable hardware-based processors that are configured to execute software instructions stored on a computer-readable storage medium (such as the memory 212), for example, instructions defining a software or computer program, wherein the computer-readable storage medium includes a storage device (such as a disk drive or an optical drive) or a memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause the processor(s) 206 to perform the techniques described herein.
[0054] Memory 212 includes one or more devices configured to store programming modules and / or data associated with the operation of access point 200. For example, memory 212 may include a computer-readable storage medium, such as a non-transitory computer-readable medium, including a storage device (e.g., a magnetic disk drive or an optical disk drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause one or more processors 206 to perform the techniques described herein.
[0055] In this example, the memory 212 stores executable software that includes an application programming interface (API) 240, a communication manager 242, configuration settings 250, a device status log 252, a data storage 254, and a log controller 255. The device status log 252 includes a list of events specific to the access point 200. These events may include a log of normal events and error events, such as memory status, restart or reboot events, crash events, cloud disconnection and self-recovery events, low link speed or link speed swing events, Ethernet port status, Ethernet interface packet errors, upgrade failure events, firmware upgrade events, configuration changes, etc., as well as time and date stamps for each event. The log controller 255 determines the logging level of the device based on instructions from the NMS 130. Data 254 can store any data used and / or generated by the access point 200, including data collected from the UE 148, such as data used to calculate one or more SLE metrics, which are sent by the access point 200 for the NMS 130 to perform cloud-based management of the wireless network 106A.
[0056] Input / output (I / O) 210 represents physical hardware components capable of interacting with a user, such as buttons, displays, etc. Although not shown, memory 212 typically stores executable software for controlling a user interface with respect to input received via I / O 210. Communications manager 242 includes program code that, when executed by processor(s) 206, allows access point 200 to communicate with UE 148 and / or network(s) 134 via any of interface(s) 230 and / or 220A-220C. Configuration settings 250 include any device settings of access point 200, such as radio settings for each of wireless interface(s) 220A-220C. These settings can be manually configured or remotely monitored and managed by NMS 130 to optimize wireless network performance on a regular basis (e.g., hourly or daily).
[0057] As described herein, the AP device 200 may measure and report network data and / or data 254 from the device status log 252 to the NMS 130. The network data may include event data, telemetry data, and / or other SLE-related data. The network data may include various parameters indicating the performance and / or status of the wireless network. These parameters may be measured and / or determined by one or more UE devices and / or one or more APs in the wireless network. The AP device 200 may periodically create network data packets according to periodic intervals. The collected and sampled data reported periodically in the statistical data packet may be referred to herein as "oc-stats". In other examples, the NMS 130 may request, retrieve, or otherwise receive the statistical data packet from the AP device 200 via the API 240, the open configuration protocol, or other communication protocols. In other examples, when certain events occur at the AP device 200, such as connection and disconnection events of a connection session with the NMS 130, the AP device 200 reports the event data to the NMS 130 in the cloud, and / or the NMS 130 may observe and record events at the AP device 200 when the event occurs. Event-driven data may be referred to herein as "oc events"
[0058] According to the disclosed technology, the AP device 200 may have a connection session, such as a TCP connection session, provided by one or more service providers with the NMS 130. For example, the AP device 200 may have an active first connection session with the NMS 130 provided by a first service provider (e.g., a broadband service provider), and an inactive second connection session provided to the NMS 130 by a second service provider (e.g., an LTE service provider).
[0059] The AP device 200 may collect and / or report connection event data to the NMS 130 as part of a network data packet or as event-driven data. The connection event data includes connection and / or disconnection events of a connection session between the AP device 200 and the NMS 130. For example, when a first connection session is established between the AP device 200 and the NMS 130, the AP device 200 may collect and / or report a first connection event of the first connection session. When the first connection session becomes unstable, the AP device 200 may experience a connection switch from the first connection session to the second connection session. For example, the AP device 200 may collect and / or report a first disconnection event of the first connection session when the first connection session is disconnected, and collect and / or report a second connection event of the second connection session when the second connection session is established between the AP device 200 and the NMS 130.
[0060] Figure 3 is a block diagram of an example network management system (NMS) 300 according to one or more techniques of the present disclosure. NMS 300 may be used to implement, for example, Figure 1A-1B In these examples, NMS 300 is responsible for monitoring and managing one or more wireless networks 106A-106N at sites 102A-102N, respectively.
[0061] The NMS 300 includes a communication interface 330, one or more processors 306, a user interface 310, a memory 312, and a database 318. The various elements are coupled together via a bus 314, through which the various elements can exchange data and information. In some examples, the NMS 300 receives data from one or more of the client devices 148, APs 142, switches 146, routers 147, gateway devices 187, and other network nodes within the network(s) 134, which data can be used to calculate one or more SLE metrics and / or update network data 316 in the database 318. Figure 3 In the example shown, NMS 300 receives and / or records connection event data 317 for NAS devices 108. For each NAS device, connection event data 317 includes one or more connection events, such as connection or disconnection events, for a connection session between the NAS device and NMS 300, where the connection session is provided by a service provider. NMS 300 analyzes the data for cloud-based management of wireless networks 106A-106N. In some examples, NMS 300 may be Figure 1A It may be part of another server shown in , and may also be part of any other server.
[0062] The processor(s) 306 execute software instructions, e.g., instructions defining a software or computer program, which are stored in a computer-readable storage medium such as memory 312, e.g., a non-transitory computer-readable medium including a storage device such as a disk drive or an optical drive or a memory such as flash memory or RAM or any other type of volatile or non-volatile memory that stores instructions to cause the processor(s) 306 to perform the techniques described herein.
[0063] The communication interface 330 may include, for example, an Ethernet interface. The communication interface 330 couples the NMS 300 to a network and / or the Internet, for example, Figure 1A The communication interface 330 includes a receiver 332 and a transmitter 334 through which the NMS 300 receives and transmits data and information to and from any client device 148, AP 142, switch 146, server 110, 116, 122, 128, and / or any other network node, device, or system forming part of the network system 100, such as Figure 1AIn some scenarios described herein, where the network system 100 includes "third-party" network devices that are owned and / or associated with entities different from the NMS 300, the NMS 300 does not receive, collect, or otherwise access network data from the third-party network devices.
[0064] The data and information received by the NMS 300 may include, for example, telemetry data, SLE-related data, or event data received from one or more of the client devices AP 148, AP 142, switch 146, router 147, gateway device 187, or other network nodes used by the NMS 300 to remotely monitor the performance of the wireless networks 106A-106N and application sessions from client devices to cloud-based application servers. The NMS 300 may also transmit data to any network device, such as the client device 148, AP 142, switch 146, router 147, gateway device 187, other network nodes within the network(s) 134, management device 111, via the communication interface 330 to remotely manage the wireless networks 106A-106N and portions of the wired network and WAN.
[0065] The memory 312 includes one or more devices configured to store programming modules and / or data associated with the operation of the NMS 300. For example, the memory 312 may include a computer-readable storage medium, such as a non-transitory computer-readable medium including a storage device (e.g., a magnetic disk drive or an optical disk drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause the one or more processors 306 to perform the techniques described herein.
[0066] In this example, the memory 312 includes an API 320, an SLE module 322, a virtual network assistant (VNA) / AI engine 350, and a radio resource management (RRM) engine 360. According to the disclosed technology, the VNA / AI engine 350 includes a continuous switching engine 352, which is configured to predict WAN problems based on detecting that the NAS device is continuously switching between connection sessions provided by different service providers. The NMS 300 may also include any other programming modules, software engines, and / or interfaces configured for remote monitoring and management of the wireless networks 106A-106N and wired network portions, including AP 142 / 200, switch 146, router 147 or other network devices (e.g., Figure 1B Remote monitoring and control of any one of the gateway routers 187).
[0067] The SLE module 322 is capable of setting and tracking thresholds for SLE metrics for each network 106A-106N. The SLE module 322 also analyzes SLE-related data collected by an AP (e.g., any one of the APs 142 from a UE in each wireless network 106A-106N). For example, APs 142A-1 to 142A-N collect SLE-related data from UEs 148A-1 to 148A-N currently connected to the wireless network 106A. The data is transmitted to the NMS 300, which is executed by the SLE module 322 to determine one or more SLE metrics for each UE 148A-1 to 148A-N currently connected to the wireless network 106A. In addition to any network data collected by one or more APs 142A-1 to 142A-N in the wireless network 106A, the data is also sent to the NMS 300 and stored in the database 318 as, for example, network data 316.
[0068] The RRM engine 360 monitors one or more metrics of each site 102A-102N in order to learn and optimize the RF environment of each site. For example, the RRM engine 360 can monitor the coverage and capacity SLE metrics of the wireless network 106 at the site 102 to identify potential problems with SLE coverage and / or capacity in the wireless network 106, and adjust the radio settings of the access points at each site to solve the identified problems. For example, the RRM engine can determine the channel and transmit power distribution of all APs 142 in each network 106A-106N. For example, the RRM engine 360 can monitor the events, power, channels, bandwidth, and number of clients connected to each AP. The RRM engine 360 can also automatically change or update the configuration of one or more APs 142 at the site 102 in order to improve the coverage and capacity SLE metrics, thereby providing users with a better wireless experience.
[0069] The VNA / AI engine 350 analyzes the data received from the network device and its own data to identify when an unexpected abnormal state is encountered at one of the network devices. For example, the VNA / AI engine 350 can identify the root cause of any unexpected or abnormal state, for example, any bad SLE (multiple) metrics indicating one or more network device connection problems. In addition, the VNA / AI engine 350 can automatically call one or more corrective measures, aiming to solve the identified (multiple) root causes of one or more bad SLE metrics. Examples of corrective actions that the VNA / AI engine 350 can automatically call may include, but are not limited to, calling RRM 360 to restart one or more APs, adjusting / modifying the transmit power of a specific radio in a specific AP, adding SSID configuration to a specific AP, changing the channel on an AP or a group of APs, etc. Corrective measures may also include restarting switches and / or routers, calling to download new software to APs, switches or routers, etc. These corrective measures are for example purposes only, and the present disclosure is not limited in this regard. If automated corrective actions are not available or do not adequately address the root cause, the VNA / AI engine 350 can proactively provide notifications, including recommended corrective actions for IT personnel (e.g., a site or network administrator using the management device 111) to take to resolve the network error.
[0070] According to one or more techniques of the present disclosure, the continuous switching engine 352 is configured to predict WAN problems based on detecting that the NAS device continuously switches between connection sessions provided by different service providers. In some examples, the disclosed concepts utilize existing NAS devices (e.g., AP devices) and WiFi-only data to infer or predict upper-level WAN problems.
[0071] According to the disclosed technology, the NMS 300 obtains connection event data 317 of a NAS device at a site, wherein each event in the connection event data 217 includes a connection or disconnection event of a connection session between the NAS device provided by a service provider and the NMS 300. The continuous switching engine 352 is configured to detect the number of connection exchanges within a time window, for example, one hour, thirty minutes, ten minutes, etc. The connection exchange includes changing the connection session from a first service provider to a second service provider. Based on the number of connection exchanges detected that meet a threshold, for example, three or more exchanges within a one-hour time window, the continuous switching engine 352 predicts that the root cause of the connection exchange is a WAN problem. The continuous switching engine 352 then generates a notification of the predicted root cause of the connection exchange. The notification can be sent to an administrator computing device, for example, Figure 1A The administrator device 111 in the site is presented to the administrator of the site.
[0072] The disclosed concepts are described herein with respect to AP devices, but are also applicable to other types of NAS devices. The AP devices for the wireless network at the site have a connection session, e.g., a TCP connection, with the NMS 300 via one or more service providers. For a given site, there may be two SPs providing Internet access, e.g., one broadband and one LTE. Figure 1A In the example of , AP 142A-1 of site 102A has an active first connection session 162A provided by SP 160A and an inactive second connection session 162B provided by SP 160B. NMS 300 receives and / or records connection / disconnection events for the connection sessions as connection event data 317, which events include an indication of the service provider with which the connection session was established with NMS 300. For example, each event included in the connection event data 317 may include an address, e.g., an Internet Protocol (IP) address, of the service provider providing the connection session that experienced the reported event.
[0073] AP devices may experience disconnection and connection events of connection sessions back and forth between two SPs due to problems with the WAN or the SPs themselves. For customer sites that only support WiFi, the NMS 300 only has visibility into the wireless network based on WiFi data collected from the AP devices, and has no visibility into the wired network or WAN at these sites. In this case, the NMS 300 traditionally cannot distinguish between WAN problems or SP problems, which is the root cause of the constant switching between connection sessions provided by different SPs.
[0074] Some solutions to this problem include using sensors in dedicated devices deployed to monitor the wide area network. However, the disclosed concept is able to predict WAN problems based on WiFi-only data collected from AP devices without requiring dedicated devices to monitor the WAN or access to WAN-specific or SP-specific data sources. Instead, the continuous exchange engine 352 of the NMS 300 uses WiFi-only data collected from AP devices, namely the connection event data 317, to infer upper-layer WAN problems based on detecting continuous exchanges between connection sessions provided by different SPs.
[0075] The disclosed concepts are intended to leverage existing AP devices and WiFi-only data to infer WAN issues based on continuous connection session exchanges by AP devices. When connection sessions are unstable, the AP device may experience continuous exchanges between SPs, which may indicate a problem with the WAN (e.g., gateway device) rather than the SPs themselves. For example, if the gateway device is not configured correctly, the gateway device itself may continuously exchange between SPs in an attempt to access the Internet. This continuous exchange of the gateway device is reflected in the connection session behavior experienced at the lower-level AP devices.
[0076] In some examples, the AP device may send a report to the NMS 300 indicating connection / disconnection events for connection sessions with the SP and, for each event, identifying the SP providing the connection session. In other examples, the NMS 300 itself observes and records the connection / disconnection events for the connection sessions between the AP device and the NMS 300. In either example, the NMS 300 may store the connection / disconnection events as connection event data 317.
[0077] The continuous exchange engine 352 can detect the connection exchange on the specific AP device based on the corresponding events in the connection event data 317 of the specific AP device. For example, to detect the connection exchange, the continuous exchange engine 352 can detect that the disconnection event of the second connection session may then lead to a subsequent connection event of the first connection session provided by the first SP, which restarts the connection exchange cycle. The continuous exchange engine 352 can detect such connection exchanges that occur continuously within a time window (e.g., one hour, thirty minutes, ten minutes, etc.), and the continuous exchange engine 354 can identify the pattern of the continuous exchange based on the connection event data 317 of the specific AP device.
[0078] In some examples, the AP device may use the same connection session to communicate with the NMS 300 and pass real traffic from a cloud-based application, application server, and / or data center to the client device. In this example, the disclosed solution may be used to identify the root cause of the connection problem for the management path and the data path. In other examples, the AP device may use two different connections, one for the management path to connect to the NMS 300 and the other for the data path to pass traffic. In this example, the disclosed solution may identify the root cause of the connection problem with the NMS 300, which may prevent the configuration file from being transferred from the NMS 300 to the AP device.
[0079] To determine the occurrence of continuous exchanges at the AP device, the continuous exchange engine 352 of the NMS 300 can view the WiFi-only data within a time window (e.g., every hour) to identify whether the connection session between the AP device and the NMS is exchanged multiple times between different SPs within the time window. The continuous exchange engine 352 can increment a counter each time the connection session of the AP device changes within the time window. Based on the counter meeting a threshold during the time window, for example, reaching or exceeding the threshold, the continuous exchange engine 352 detects the continuous exchange of the connection session of the AP device. In some examples, the threshold can be equal to one of three exchanges, five exchanges, ten exchanges, etc., and can vary based on the associated time window.
[0080] The raw data reported from the AP device to the NMS 300 and / or recorded by the NMS 300 indicates connection / disconnection events for connection sessions and, for each event, indicates the IP address of the SP providing the connection session. The IP address is provided by the SP. Based on the data from the AP device in the connection event data 317, the continuous switching engine 352 performs a lookup (reverse lookup) of the IP address to identify the SP name and location. In some examples, the continuous switching engine 352 can use a third-party service to perform a reverse lookup of the IP address to identify the SP. The continuous switching engine 352 then compares the IP address and / or SP name of each connection / disconnection event within the time window to detect a switching event and increment a counter. In some examples, the location of the SP can be used to filter out situations where the AP site is geographically far away from the NMS 300, so that the switching may be a management path problem rather than a WAN problem.
[0081] The continuous exchange engine 352 can be configured to continuously review the WiFi-only data of the site's AP devices included in the connection event data 317 to identify continuous exchange instances at the site's APs according to a time window. In response to the identification of continuous exchange by the site AP, the continuous exchange engine 352 can generate a notification indicating an inferred or predicted WAN problem. The NMS 300 can send the notification to a network administrator of the site or organization via the communication interface 330, for example, Figure 1A The management device 111 in.
[0082] Although generally described in this disclosure as monitoring connection exchanges of a single AP device, the continuous exchange engine 352 monitors connection exchanges of connection sessions from multiple AP and / or NAS devices at a given site. For example, to predict WAN problems, the continuous exchange engine 352 may determine that all or most of the multiple AP devices at the site are experiencing continuous exchanges of connection sessions between service providers.
[0083] In a further example, the continuous exchange engine 352 can consider the scope of data observed, collected and / or recorded by the NMS 300, including the type of AP that is experiencing the continuous exchange, the location of the AP or station that is experiencing the continuous exchange, the duration of the continuous exchange (e.g., a time window or multiple consecutive time windows). In some cases, the continuous exchange engine 352 can determine the severity associated with the detected connection exchange and modify the counter threshold used to identify the continuous exchange within the time window based on the severity of the exchange, or another threshold can be applied that requires the continuous exchange to be detected within multiple consecutive time windows before generating a notification. The notification can include checking the gateway device (e.g., Figure 1A , 1B recommendations on the health of gateway devices 187) and other WAN health metrics.
[0084] In some examples, ML model 380 may include a supervised ML model that is trained using training data that includes pre-collected labeled network data received from network devices (e.g., client devices, APs, switches, and / or other network nodes) to detect AP continuous switching between SPs. The supervised ML model may include one of logistic regression, naive Bayes, support vector machine (SVM), etc. In other examples, ML model 380 may include an unsupervised ML model. Although Figure 3 Not shown, but in some examples, database 318 may store training data, and VNA / AI engine 350 or a dedicated training module may be configured to train ML model 380 based on the training data to determine appropriate weights for one or more features of the training data.
[0085] Although the techniques of the present disclosure are described in this example as being performed by NMS 130, the techniques described herein may be performed by any other computing device(s), system(s), and / or server(s), and the present disclosure is not limited in this regard. For example, one or more computing devices configured to perform the functions of the techniques of the present disclosure may reside in a dedicated server, or be included in any other server outside of NMS 130, or may be distributed throughout network 100, and may or may not form part of NMS 130.
[0086] Figure 4 An example user equipment (UE) device 400 is shown, in accordance with one or more techniques of this disclosure. Figure 4 The example UE device 400 shown in FIG. 4 may be used to implement the Figure 1A Any of the UE 148 shown and described. UE device 400 may include any type of wireless client device, and the present disclosure is not limited in this regard. For example, UE device 400 may include a mobile device such as a smart phone, a tablet or a laptop computer, a personal digital assistant (PDA), a wireless terminal, a smart watch, a smart ring, or any other type of mobile or wearable device. In some examples, UE 400 may also include a wired client device, for example, an Internet of Things device such as a printer, a security sensor or device, an environmental sensor, or any other device connected to a wired network and configured to communicate via one or more wireless networks.
[0087] UE device 400 includes a wired interface 430, wireless interfaces 420A-420C, one or more processors 406, memory 412, and a user interface 410. The various components are coupled together via a bus 414, and the various components can exchange data and information through the bus 414. The wired interface 430 represents a physical network interface and includes a receiver 432 and a transmitter 434. If necessary, the wired interface 430 can be used to communicate via a cable (such as Figure 1A The UE 400 is coupled directly or indirectly to a wired network device within the wired network, such as one of the Ethernet cables 144. Figure 1A One of the switches 146 in.
[0088] The first, second and third wireless interfaces 420A, 420B and 420C include receivers 422A, 422B and 422C, respectively, each of which includes a receiving antenna, through which the UE 400 can receive wireless signals from a wireless communication device, e.g. Figure 1A AP 142 Figure 2 The first, second and third wireless interfaces 420A, 420B and 420C further include transmitters 424A, 424B and 424C, respectively, each of which includes a transmitting antenna, through which the UE 400 can send wireless communication signals to wireless communication devices (such as Figure 1A AP 142 Figure 2 The first wireless interface 420A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz), and the second wireless interface 420B may include a Bluetooth interface and / or a Bluetooth low energy interface. The third wireless interface 420C may include, for example, a cellular interface, through which the UE device 400 can be connected to a cellular network.
[0089] The processor(s) 406 execute software instructions, e.g., instructions defining a software or computer program, which are stored in a computer-readable storage medium such as the memory 412, e.g., a non-transitory computer-readable medium including a storage device such as a disk drive or an optical drive or a memory such as flash memory or RAM or any other type of volatile or non-volatile memory that stores instructions to cause the processor(s) 406 to perform the techniques described herein.
[0090] Memory 412 includes one or more devices configured to store programming modules and / or data associated with the operation of UE 400. For example, memory 412 may include a computer-readable storage medium, such as a non-transitory computer-readable medium, including a storage device (e.g., a magnetic disk drive or an optical disk drive) or a memory (e.g., flash memory or RAM) or any other type of volatile or non-volatile memory that stores instructions to cause one or more processors 406 to perform the techniques described herein.
[0091] In this example, memory 412 includes operating system 440, applications 442, communication module 444, configuration settings 450, and data storage 454. Communication module 444 includes program code that, when executed by processor(s) 406, enables UE 400 to communicate using any of wired interface(s) 430, wireless interfaces 420A-420B, and / or cellular interface 450C. Configuration settings 450 include any device settings of UE 400, including settings for each of wireless interface(s) 420A-420B and / or cellular interface 420C.
[0092] The data store 454 may include, for example, a status / error log that includes a list of events specific to the UE 400. The events may include a log of normal events and error events, depending on the logging level based on instructions from the NMS 130. The data store 454 may store any data used and / or generated by the UE 400, such as data used to calculate one or more SLE metrics or identify relevant behavioral data, which is collected by the UE 400 and transmitted directly to the NMS 130 or to any AP 142 in the wireless network 106 for further transmission to the NMS 130.
[0093] As described herein, UE 400 can measure network data and report it from data storage 454 to NMS 130. Network data can include event data, telemetry data, and / or other SLE-related data. Network data can include various parameters indicating the performance and / or status of a wireless network. NMS 130 can determine one or more SLE metrics based on SLE-related data received from UEs or client devices in the wireless network, and store the SLE metrics as network data 137 ( Figure 1A ).
[0094] Alternatively, the UE device 400 may include an NMS agent 456. The NMS agent 456 is a software agent of the NMS 130 installed on the UE 400. In some examples, the NMS agent 456 may be implemented as a software application running on the UE 400. The NMS agent 456 collects information including detailed client device attributes from the UE 400, including insights into the roaming behavior of the UE 400. This information provides insights into the client roaming algorithm, because roaming is a decision of the client device. In some examples, the NMS agent 456 may display the client device attributes on the UE 400. The NMS agent 456 sends the client device attributes to the NMS 130 via the AP device to which the UE 400 is connected. The NMS agent 456 may be integrated into a custom application or as part of a location application. The NMS agent 456 may be configured to identify the device connection type (e.g., cellular or Wi-Fi) and the corresponding signal strength. For example, the NMS agent 456 identifies the access point connection and its corresponding signal strength. The NMS agent 456 may store information about the APs identified by a given UE 400 and their corresponding signal strengths. The NMS agent 456 or other elements of the UE 400 also collect information about which APs the UE 400 is connected to, which also indicates which APs the UE 400 is not connected to. The NMS agent 456 of the UE 400 sends this information to the NMS 130 via the AP to which it is connected. In this way, the UE 400 not only sends information about the AP to which the UE 400 is connected, but also sends information about other APs that the UE 400 identifies and is not connected to, and their signal strengths. The AP in turn forwards this information to the NMS, including information about other APs that the UE 400 identifies in addition to itself. This additional level of granularity enables the NMS 130, and ultimately the network administrator, to better determine the Wi-Fi experience directly from the perspective of the client device.
[0095] In some examples, the NMS agent 456 further enriches the client device data utilized in the service level. For example, the NMS agent 456 can go beyond basic fingerprint identification and provide supplemental details for attributes such as device type, manufacturer, and different versions of the operating system. In detailed client attributes, the NMS 130 can display the radio hardware and firmware information of the UE 400 received from the NMS client agent 456. The more details that the NMS agent 456 can extract, the better the VNA / AI engine is at high-level device classification. The VNA / AI engine of the NMS 130 is constantly learning and becoming more accurate in its ability to distinguish between device-specific issues or broad device issues, for example, specifically identifying that a specific operating system version is affecting certain clients.
[0096] In some examples, the NMS agent 456 can cause the user interface 410 to display a prompt prompting the end user of the UE 400 to enable location permissions before the NMS agent 456 can report the device's location, client information, and network connection data to the NMS. The NMS agent 456 will then begin reporting connection data as well as location data to the NMS. In this way, the end user of the client device can control whether the NMS agent 456 can report client device information to the NMS.
[0097] Figure 5 is a block diagram illustrating an example network node 500 according to one or more techniques of the present disclosure. In one or more examples, the network node 500 implements an Figure 1A A device or server of the network 134, such as a switch 146, an AAA server 110, a DHCP server 116, a DNS server 122, a network server 128, etc., or a device or server supporting Figure 1B Another network device of one or more of the wireless network 106, wired LAN 175, or SD-WAN 177, or data center 179, for example, router 187.
[0098] In this example, the network node 500 includes a wired interface 502, such as an Ethernet interface, a processor 506, an input / output 508, such as a display, buttons, a keyboard, a keypad, a touch screen, a mouse, etc., and a memory 512 coupled together by a bus 514, and the various elements can exchange data and information on the bus 514. The wired interface 502 couples the network node 500 to a network, such as an enterprise network. Although only one interface is shown as an example, the network node can and typically has multiple communication interfaces and / or multiple communication ports. The wired interface 502 includes a receiver 520 and a transmitter 522.
[0099] The memory 512 stores executable software applications 532, an operating system 540, and data / information 530. The data 530 may include system logs and / or error logs storing event data (including behavioral data) for the network node 500. In examples where the network node 500 comprises a "third party" network device, the same entity does not own or have access to the AP or wired client device and the network node 500. Thus, in examples where the network node 500 is a third party network device, the NMS 130 does not receive, collect, or otherwise access network data from the network node 500.
[0100] In an example where network node 500 includes a server, network node 500 can receive data and information via receiver 520, e.g., including operation-related information such as registration requests, AAA services, DHCP requests, Simple Notification Service (SNS) lookups, and web page requests, and send data and information via transmitter 522 (e.g., including configuration information, authentication information, web page data, etc.).
[0101] In an example where the network node 500 includes a wired network device, the network node 500 can be connected to one or more APs or other wired client devices, such as an IoT device, via a wired interface 502. For example, the network node 500 can include multiple wired interfaces 502 and / or the wired interface 502 can include multiple physical ports for connecting to multiple APs or other wired client devices within a site via corresponding Ethernet cables. In some examples, each AP or other wired client device connected to the network node 500 can access a wired network via a wired interface 502 in the network node 500. In some examples, one or more APs or other wired client devices connected to the network node 500 can each obtain power from the network node 500 via a corresponding Ethernet cable and a Power over Ethernet (PoE) port of the wired interface 502.
[0102] In examples where network node 500 includes a session-based router that employs a stateful, session-based routing scheme, network node 500 can be configured to independently perform path selection and traffic engineering. Using session-based routing can enable network node 500 to avoid using a centralized controller (such as an SDN controller) to perform path selection, traffic engineering, and avoid using tunnels. In some examples, network node 500 can implement session-based routing as Security Vector Routing (SVR) provided by Juniper Networks, Inc. In examples where network node 500 includes a session-based router (e.g., SDN controller) that operates as a network gateway for a site of an enterprise network, network node 500 can implement session-based routing as a security vector router (SVR) provided by Juniper Networks, Inc. Figure 1A , 1B In the case of a router or gateway 187), the network node 500 may be located at the underlying physical WAN (e.g., Figure 1B SD-WAN 177) with one or more other session-based routers (e.g., Figure 1B The network node 500 operating as a session-based router can collect data at the peer path level and report the peer path data to the NMS 130.
[0103] In examples where network node 500 includes a packet-based router, network node 500 may employ a packet-based or flow-based routing scheme to forward packets according to defined network paths, such as those established by a centralized controller that performs path selection and traffic engineering. In examples where network node 500 includes a packet-based router operating as a network gateway for an enterprise network site (e.g., Figure 1A , Figure 1B In the case of a router or gateway 187A), the network node 500 may be located in the underlying physical WAN (e.g., Figure 1B The network node 500 may collect data at the tunnel level on the SD-WAN 177 with one or more other packet-based gateways operating as network portals to other sites in the enterprise network (e.g., the network node 500 operating as a packet-based router), and the tunnel data may be retrieved by the NMS 130 via an API or open configuration protocol, or the tunnel data may be reported to the NMS 130 via the NMS agent 544 or another module running on the network node 500.
[0104] The data collected and reported by the network node 500 may include periodically reported data and event-driven data. The network node 500 is configured to collect logical path statistics through bidirectional forwarding detection (BFD) detection and data extracted from messages and / or counters at the logical path (e.g., peer path or tunnel) level. In some examples, the network node 500 is configured to collect statistics and / or sample other data according to a first periodic interval (e.g., every 3 seconds, every 5 seconds, etc.). The network node 500 can store the collected and sampled data as path data, for example, in a buffer.
[0105] In some examples, the network node 500 alternatively includes an NMS agent 544. The NMS agent 544 may periodically create a statistical data package according to a second periodic interval (e.g., every 3 minutes). The collected and sampled data reported periodically in the statistical data package may be referred to as "oc-stats" in this article. In some examples, the statistical data package may also include detailed information about the client and the related client session connected to the network node 500. The NMS agent 544 may then report the statistical data package to the NMS 130 in the cloud. In other examples, the NMS 130 may request, retrieve or otherwise receive the statistical data package from the network node 500 via an API, an open configuration protocol, or another communication protocol. The statistical data package created by the NMS agent 544 or another module of the network node 500 may include a header identifying the network node 500 and the statistical data and data samples of each logical path from the network node 500. In other examples, NMS agent 544 reports event data to NMS 130 in the cloud in response to the occurrence of certain events at network node 500 when events occur, and / or NMS 130 may observe and record events on network node 500 when events occur. Event-driven data may be referred to herein as "oc events".
[0106] Figure 6 is a flowchart illustrating an example operation of a network management system according to one or more techniques of the present disclosure. Figure 1A Example operations are described with reference to NMS 130 and AP device 142 in FIG.
[0107] The NMS 130 obtains connection event data for one or more AP devices 142A at the site 106A, wherein each event included in the connection event data includes a connection or disconnection event for a connection session, such as connection sessions 162A, 162B between the AP 142A-1 and the NMS 130 provided by the service providers 160A, 160B, respectively (602). The connection sessions 162A, 162B may include a TCP connection session for a management path between the AP 142A-1 at the site 102A and the NMS 130. In some examples, the data path between the AP 142A-1 at the site and one or more of the cloud-based applications, application servers, and / or data centers includes the same path as the management path.
[0108] To obtain the connection event data, NMS 130 is configured to read the connection event data of AP device 142A at site 102A from the record created by NMS 130 or receive the connection event information reported by AP device 142B at site 102A. In some cases, NMS 130 determines the physical distance between site 102N and NMS 130, and filters out the connection event data from AP device 142N at site 102N when the physical distance exceeds a preset distance.
[0109] The NMS 130 detects a number of connection exchanges in the connection event data within the time window, wherein the connection exchange includes a change from a connection session 162A provided by a first SP 160A to a connection session 162B provided by a second SP 160B (604). The first SP 160A may provide a first connection type, such as broadband, and the second SP 160B may provide a second connection type different from the first connection type, such as LTE. Each event included in the connection event data may include an address of the SP 160A, 160B providing the connection session that experienced the connection or disconnection event. For each event included in the connection event data, the NMS 130 performs a reverse lookup of the address of the SP included in the event to determine the name and location of the SP providing the connection session that experienced the connection or disconnection event.
[0110] To detect a connection exchange including a change from a first connection session 162A provided by the first SP 160A to a second connection session 162B provided by the second SP 160B based on the connection event data, the NMS 130 is configured to detect a first event including a connection event of the first connection session 162A provided by the first SP 160A between the AP 142A-1 and the NMS 130, detect a second event including a disconnection event of the first connection session 162, detect a third event including a connection event of the second connection session 162B provided by the second SP 160B between the AP 142A-1 and the NMS 130, and detect a fourth event including a disconnection event of the second connection session 162. To detect the number of connection exchanges in the connection event data within a time window, the NMS 130 is configured to increment a counter for each connection exchange of the AP 142A-1 occurring between the first service provider 160A and the second service provider 160B within the time window. The time window includes a rolling time window.
[0111] Based on the number of detected connection exchanges that meet the threshold, the NMS 130 predicts the root cause of the connection exchange as a WAN problem (606). After each increment of the counter, the NMS 130 determines whether the current number of connection exchanges meets the threshold. In some examples, the NMS 130 can determine a severity associated with the number of detected connection exchanges that meet the threshold within a single time window or for each of two or more consecutive time windows, and modify the threshold based on the severity of the number of detected connection exchanges.
[0112] NMS 130 generates a notification of the predicted root cause of the connection exchange (608). In some examples, the notification can be sent to an administrator computing device, such as administrator device 111, for presentation to an administrator of the site. The notification of the predicted root cause of the connection exchange can include a recommendation to determine at least one of a WAN health metric or a health metric of one or more gateway devices of the WAN.
[0113] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. The various features described as modules, units, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices or other hardware devices. In some cases, various features of electronic circuits may be implemented as one or more integrated circuit devices, such as an integrated circuit chip or chipset.
[0114] If implemented in hardware, the present disclosure may be directed to an apparatus such as a processor or an integrated circuit device (e.g., an integrated circuit chip or chipset). Alternatively or additionally, if implemented in software or firmware, the techniques may be implemented at least in part by a computer-readable data storage medium that includes instructions that, when executed, cause the processor to perform one or more of the methods described above. For example, a computer-readable data storage medium may store such instructions for execution by a processor.
[0115] The computer-readable medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include computer data storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory, electrically erasable programmable read-only memory, flash memory, magnetic or optical data storage media, etc. In some examples, the article of manufacture may include one or more computer-readable storage media.
[0116] In some examples, computer-readable storage media may include non-transient media. The term "non-transient" may mean that the storage medium is not contained in a carrier or propagating signal. In some examples, non-transient storage media may store data that changes over time (e.g., in RAM or cache).
[0117] The code or instructions may be software and / or firmware executed by a processing circuit that includes one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described in the present disclosure may be provided within a software module or a hardware module.
Claims
1. A network management system NMS, comprising: Memory; as well as one or more processors in communication with the memory and configured to: obtaining connection event data for one or more network access server (NAS) devices at a site, wherein each event included in the connection event data comprises a connection or disconnection event of a connection session provided by a service provider between the NAS devices of the one or more NAS devices and the NMS; detecting a number of connection switches in the connection event data within a time window, wherein the connection switch comprises a change from a first connection session provided by a first service provider to a second connection session provided by a second service provider; predicting a root cause of the connection swap as a wide area network problem based on a detected number of the connection swaps satisfying a threshold; and A notification of the predicted root cause of the connection exchange is generated.
2. A network management system according to claim 1, wherein each event included in the connection event data includes an address of the service provider that provides the connection session that experiences the connection or disconnection event, and wherein the one or more processors are configured to perform a reverse lookup of the address of the service provider included in the event for each event included in the connection event data to determine the name and location of the service provider that provides the connection session.
3. The network management system of claim 1, wherein the first service provider provides a first connection type, and the second service provider provides a second connection type different from the first connection type.
4. The network management system of claim 1 , wherein to obtain the connection event data, the one or more processors are configured to read the connection event data for the one or more NAS devices at the site from a record created by the NMS, or to receive the connection event data reported by the one or more NAS devices at the site.
5. The network management system of claim 1 , wherein to detect the connection switch, the one or more processors are configured to: detecting a first event, the first event comprising a connection event of the first connection session provided by the first service provider between the NAS device and the NMS; detecting a second event comprising a disconnection event of the first connection session; detecting a third event, the third event comprising a connection event of the second connection session provided by the second service provider between the NAS device and the NMS; as well as A fourth event identifying a disconnect event of the second connection session is detected.
6. The network management system of any one of claims 1-5, wherein in order to detect the number of the connection exchanges in the connection event data within the time window, the one or more processors are configured to increment a counter for each connection exchange of the NAS device between the first service provider and the second service provider occurring during the time window. 7 . The network management system of claim 6 , wherein the one or more processors are configured to determine whether a current number of connection exchanges meets the threshold after each increment of the counter.
8. The network management system according to any one of claims 1 to 5, wherein the one or more processors are configured to: determining a severity associated with a number of the detected connection exchanges that satisfy the threshold in a single time window or for each of two or more consecutive time windows; and The threshold is modified based on the severity of the detected number of connection exchanges.
9. The network management system according to any one of claims 1 to 5, wherein each of the first connection session and the second connection session comprises a transmission control protocol connection session for a management path between the NAS device at the site and the NMS.
10. The network management system of claim 9, wherein a data path between the NAS device at the site and one or more of a cloud-based application, an application server, or a data center comprises the same path as the management path.
11. The network management system according to any one of claims 1 to 5, wherein the time window comprises a rolling time window.
12. The network management system of any of claims 1-5, wherein the notification of the predicted root cause of the connection switch includes a recommendation to determine at least one of a WAN health metric or a health metric of one or more gateway devices of the WAN.
13. The network management system according to any one of claims 1 to 5, wherein the one or more processors are configured to: determining a physical distance between the site and the NMS; and When the physical distance exceeds a preset distance, the connection event data for the one or more NAS devices at the site is filtered out.
14. A network management method, comprising: obtaining, by a network management system NMS, connection event data for one or more network access server NAS devices at a site, wherein each event included in the connection event data comprises a connection or disconnection event of a connection session provided by a service provider between the NAS devices of the one or more NAS devices and the NMS; detecting, by the NMS, a number of connection switches in the connection event data within a time window, wherein the connection switch comprises a change from a first connection session provided by a first service provider to a second connection session provided by a second service provider; Based on the number of the detected connection exchanges that meet a threshold, predicting, by the NMS, a root cause of the connection exchange as a wide area network problem; as well as A notification of the predicted root cause of the connection switch is generated by the NMS.
15. A network management method according to claim 14, wherein each event included in the connection event data includes the address of the service provider that provides the connection session that experiences the connection or disconnection event, and wherein the method further includes, for each event included in the connection event data, performing a reverse lookup on the address of the service provider included in the event to determine the name and location of the service provider that provides the connection session.
16. The network management method according to claim 14, wherein detecting the connection exchange comprises: detecting a first event, the first event comprising a connection event of the first connection session provided by the first service provider between the NAS device and the NMS; detecting a second event comprising a disconnection event of the first connection session; detecting a third event, the third event comprising a connection event of the second connection session provided by the second service provider between the NAS device and the NMS; as well as A fourth event identifying a disconnect event of the second connection session is detected.
17. The network management method of any one of claims 14-16, wherein detecting the number of connection exchanges in the connection event data within the time window comprises incrementing a counter for each connection exchange of the NAS device between the first service provider and the second service provider occurring during the time window.
18. The network management method according to any one of claims 14 to 16, further comprising: determining a severity associated with a number of detected said connection exchanges that satisfy said threshold in a single time window or for each of two or more consecutive time windows; as well as The threshold is modified based on the severity of the detected number of connection exchanges.
19. The network management method according to any one of claims 14 to 16, further comprising: determining a physical distance between the site and the NMS; as well as When the physical distance exceeds a preset distance, the connection event data of the one or more NAS devices at the site are filtered out.
20. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured as the network management system according to any one of claims 1 to 13, or to be configured to execute the network management method according to any one of claims 14 to 19.
Citation Information
Patent Citations
Link status monitoring based on packet loss detection
US10200264B2
Stateful load balancing in a stateless network
US10277506B2
Network packet flow controller with extended session management
US10432522B2
Intent-based analytics
US10756983B2
Method for conveying AP error codes over BLE advertisements
US10862742B2