Virtual network assistant with active analysis and association engine using ML models
Through the active analysis and association engine (PACE) in the virtual network assistant (VNA), the unsupervised machine learning model is used to solve the problem of inefficient network diagnosis in complex computer networks, and automated diagnosis and resource optimization are achieved.
Patent Information
- Application Number
- CN202510360073.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-24
- Filing Date
- 2021-07-28
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to efficiently diagnose and solve network problems in complex computer networks, especially under high load conditions, resulting in waste of resources and low diagnostic efficiency.
Through the Virtual Network Assistant (VNA), the active analysis and association engine (PACE) is performed by dynamically building and applying unsupervised machine learning (ML)-based models, reducing the resources required for network diagnosis and automatically identifying abnormal behaviors that require further analysis.
Automatic network diagnosis is realized, accurately identifying transient problems and abnormal problems that require in-depth analysis, reducing resource waste and improving diagnostic efficiency.
Smart Images

Figure CN120224239A_ABST
Abstract
Description
[0001] Division Application Instructions
[0002] This application is a divisional application of a Chinese patent application with an application date of July 28, 2021, an application number of 202110857663.5, and a title of "Virtual Network Assistant with Active Analysis and Correlation Engine Using ML Model".
[0003] This application claims the benefit of U.S. Application No. 17 / 303,222, filed on May 24, 2021, and U.S. Provisional Patent Application No. 63 / 177,253, filed on April 20, 2021, the entire contents of each of which are incorporated herein by reference. Technical Field
[0004] This disclosure generally relates to computer networks, and more particularly, to machine learning-based diagnostics of computer networks and network systems. Background Art
[0005] Wireless access networks utilize a network of wireless access points (APs), which are physical electronic devices that enable other devices to wirelessly connect to a wired network using various wireless networking protocols and technologies (such as one or more of wireless local area networking protocols compliant with the IEEE 802.11 standard (i.e., "WiFi"), Bluetooth / Bluetooth Low Energy (BLE), mesh networking protocols (such as ZigBee), or other wireless networking technologies). Many different types of devices (such as laptops, smartphones, tablet computers, wearable devices, appliances, and Internet of Things (IoT) devices) incorporate wireless communication technology and can be configured to connect to a wireless access point when the device is within range of a compatible wireless access point in order to access the wired network.
[0006] Wireless access networks and computer networks in general are complex systems that can experience transient and / or permanent problems. Some problems may result in a significant degradation of system performance, while other problems may resolve themselves without significantly affecting the system-level performance as perceived by the user. Some problems may be expected and acceptable under high load, and once the load subsides, self-healing mechanisms (such as retries, etc.) may make the problem disappear. Summary of the Invention
[0007] Generally, the present disclosure describes techniques that enable a Virtual Network Assistant (VNA) to execute a Proactive Analytics and Correlation Engine (PACE) that is configured to dynamically build and apply unsupervised machine learning (“ML-based”) models for reducing or minimizing resources spent on network diagnostics. As described herein, the techniques enable the Proactive Analytics and Correlation Engine to apply an unsupervised ML-based model to collected network event data to determine whether the network event data represents an expected transient network error that can self-correct or an abnormal behavior that needs to be further analyzed by the Virtual Network Assistant to support resolution of an underlying fault in the network system.
[0008] In addition, these techniques achieve adaptive closed-loop tuning of the unsupervised ML-based network model by leveraging real-time network data partitioned into time series subgroups of sliding windows for dynamically computing expected ranges (minimum / maximum expected occurrences) for various types of network events over a defined time period. The ML-based models applied by the VNA are trained to predict an occurrence level for a network event based on actual network event data used as training data, which is augmented with dynamically determined expected ranges for different types of network events. The PACE of a Network Management System (NMS) applies the (multiple) ML-based models to network event data received from the network system (excluding the most recent observation time frame) and operates to predict the occurrence level expected to be seen during the current observation time frame for various types of network events and the estimated (predicted) minimum and maximum thresholds for each type of network event, i.e., the predicted tolerance range for the number of occurrences of each type of network event. After determining that the true observations of the network event data during the current observation period deviate outside the range set by the minimum and / or maximum thresholds estimated (predicted) by the model during that period, the Proactive Analytics and Correlation Engine flags those network events as indicating abnormal behavior, thus triggering a more detailed root cause analysis.
[0009] The techniques of the present disclosure provide one or more technical advantages and practical applications. For example, the techniques implement automated Virtual Network Assistants that can determine which network problems should be analyzed and which problems should be considered transient problems that can be resolved on their own and thus ignored without spending additional computational resources.
[0010] To ensure that complex computer networks meet the needs of their user base, network administrators seek to quickly resolve any problems that may arise during system operation. On the other hand, analyzing the network and trying to find the root cause of every problem can result in a waste of computational resources because the system may over-analyze the root cause of a problem that has already been rectified, e.g., through a retry mechanism, before or immediately after the results of the network analyzer become available.
[0011] Further, to achieve certain technical efficiencies, the technology implements an automated virtual network assistant based on an unsupervised ML-based model, thereby reducing and / or eliminating the time-consuming task of labeling each message flow and statistic as representing a "good / normal" message flow or a "bad / failed" message flow.
[0012] In one example, the present disclosure relates to a method that includes: receiving network event data indicative of the operational behavior of a network, where the network event data defines a series of network events of one or more event types; dynamically determining, for each event type and based on the network event data, corresponding minimum (MIN) and maximum (MAX) thresholds that define an expected occurrence range for the network events; the method further includes: constructing an unsupervised machine learning model based on the network event data and the dynamically determined minimum and maximum thresholds for each event type, without the need to label each network event in the network events of the network event data; and after constructing the unsupervised machine learning model, using the machine learning model to process additional network event data to determine a predicted occurrence count of network events for each event type among the event types. The method further includes identifying one or more network events among the network events as indicative of abnormal network behavior based on the predicted occurrence count and the dynamically determined minimum and maximum thresholds for each event type.
[0013] In another example, the present disclosure relates to a network management system (NMS) for managing one or more access point (AP) devices in a wireless network. The NMS includes a memory that stores network event data received from the AP devices, where the network event data indicates the operational behavior of the wireless network and where the network event data defines a series of network events of one or more event types that vary over time. The NMS is configured to apply an unsupervised machine learning model to the network event data to determine, for the most recent observation period among observation periods: (i) a predicted occurrence count of network events for each event type among the event types, and (ii) an estimated minimum (MIN) threshold and an estimated maximum (MAX) threshold for each event type, where the MIN threshold and the MAX threshold define an expected occurrence range for network events of the corresponding event type; and identifying one or more network events among the network events as indicative of abnormal network behavior based on the estimated minimum threshold and the estimated maximum threshold and the actual network event data for the most recent observation period among the observation periods. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1is a block diagram of an example network system where a Virtual Network Assistant (VNA) executes a Proactive Analytics and Correlation Engine (PACE) that is configured to dynamically build and apply unsupervised ML-based models for network diagnostics.
[0015] Figure 2 is a block diagram of an example access point device in accordance with one or more techniques of the present disclosure.
[0016] Figure 3 is a block diagram of an example network management system having a VNA configured to execute PACE, which is configured to dynamically build and apply unsupervised ML-based models for network diagnostics.
[0017] Figure 4 is a block diagram of an example network node (such as a router or switch) in accordance with one or more techniques of the present disclosure.
[0018] Figure 5 is a block diagram of an example user equipment device in accordance with one or more techniques of the present disclosure.
[0019] Figure 6 illustrates an example network event table.
[0020] Figure 7 illustrates a time window used to measure the variability associated with each network event counter.
[0021] Figure 8 illustrates a flowchart of a process for estimating the expected variability and bounds of network event counters.
[0022] Figure 9 illustrates a time series of values of a VPE counter including dynamic bounds.
[0023] Figure 10 illustrates a time series of values of a determined VPE counter including dynamic bounds.
[0024] Figure 11 shows an example diagram depicting the training of a behavior model of a system.
[0025] Figure 12A and Figure 12B is a graph depicting a function of VPE error
[0026] Figure 13 shows an example histogram of prediction errors.
[0027] Figure 14 illustrates a flowchart for building a network event behavior model, including ensuring that the model excludes anomalous behavior.
[0028] Figure 15 FIG. is a flowchart illustrating a process triggered when a VPE prediction error exceeds a dynamic boundary segment. DETAILED DESCRIPTION
[0029] As described herein, commercial premises such as offices, hospitals, airports, stadiums or retail stores typically install complex wireless network systems throughout the premises, including a network of wireless access points (APs), to provide wireless network services to one or more client devices (or simply "clients") at the site. Clients may include, for example, smart phones or other mobile devices, Internet of Things (IoT) devices, etc. As mobile clients move throughout the premises, these mobile clients may automatically switch from one wireless access point to another wireless access point within range in order to provide seamless network connectivity throughout the premises for the user. Additionally, if a particular wireless access point provides poor coverage, the client device may automatically attempt to connect to a different wireless access point with better coverage.
[0030] In many examples, wireless network service providers implement systems to monitor and collect one or more network performance metrics to monitor network behavior and measure the performance of the wireless network at the site. For example, service level expectation (SLE) metrics can be used to measure various aspects of wireless network performance. SLE metrics seek to measure and understand network performance from the perspective of the end user experience on the network. One example SLE metric is the coverage metric, which tracks the number of user minutes in which the received signal strength indicator (RSSI) of a client measured by the access point to which the client is connected is below a configurable threshold. Another example SLE metric is the roaming metric, which tracks the percentage of successful roaming between two access points within a specified threshold for a client. Other example SLE metrics can include connection time, throughput, successful connection, capacity, AP uptime, and / or any other metric that can indicate one or more aspects of wireless network performance. The thresholds can be customized and configured by the wireless service provider to define the service level expectations at the site.
[0031] According to one or more techniques described herein, a virtual network assistant (VNA) executes a proactive analysis and correlation engine (PACE) that is configured to dynamically build and apply an unsupervised ML-based model for network diagnosis, i.e., a machine learning model trained using unlabeled training data. As described herein, the techniques enable the proactive analysis and correlation engine of the virtual network assistant to detect network problems that need to be resolved and support the resolution of identified faults in the network system.
[0032] As an example, the present disclosure describes example embodiments in which a network management system (NMS) of a complex network system (e.g., a wireless network system) implements an active analysis and correlation engine to determine whether a problem detected in the network system is expected as part of normal system operation or is atypical. If the NMS determines that the problem is typical under the observed conditions, then assuming the problem can be resolved by self-healing mechanisms such as restart or auto-reconfiguration mechanisms, the NMS can be configured to ignore these problems. However, if the NMS determines that the problem is atypical under the observed conditions, then the NMS can be configured to automatically invoke more complex and computationally costly network analysis to determine the root cause of the problem and automatically perform remediation, such as restarting or reconfiguring one or more network components to restore a satisfactory system-level experience (SLE).
[0033] In the various examples described herein, the technology is capable of building an unsupervised ML model based on network event data collected by the NMS for the network. For example, according to the technology described herein, the NMS can automatically generate and retrain an unsupervised ML model for the active analysis and correlation engine based on network events extracted from historically observed messages and / or statistics for the network. Then, the active analysis and correlation engine of the NMS can apply the ML model to data streams and / or logs of newly collected data of various network event types (e.g., statistics, messages, SLE metrics, etc., referred to herein as "PACE" event data or event types) to detect whether the currently observed network event data and the incoming data stream indicate normal system operation or whether the incoming network event data indicates an atypical system behavior event or trend corresponding to a failing network that needs mitigation.
[0034] As described, when the active analysis and correlation engine applies the ML model to network event data indicating a need for mitigation, the NMS may invoke a more complex root cause network analysis component of a virtual network assistant (VNA) to identify the root cause of the abnormal system behavior and, if possible, trigger an automated or semi-automated corrective action. In this way, the active analysis and correlation engine (PACE) can build and apply an ML model based on a particular complex network, where PACE is deployed as a mechanism for quickly and effectively determining whether to perform further resource-intensive analysis on the incoming stream of network event data collected from elements within the complex network system (e.g., in real time).
[0035] Further, in addition to identifying which issues need attention, some of the examples described herein can be configured to monitor messages exchanged within a complex network system as well as a number of operation counters and statistics. During normal operation, the ratios between the values of different counters and statistics can be assumed to be within a specific range of acceptable values, referred to herein as the {Min,Max} range. As described in more detail below, one technical advantage of the techniques described herein is that the NMS can be configured to build an unsupervised ML model such that the ML model automatically determines and adjusts the dynamic {Min,Max} range of acceptable values for network event data (e.g., specific statistics, counters, metrics, etc.) representing normal operation, thereby improving the accuracy and reliability of the ML model used to determine whether to trigger a more in-depth root cause analysis of the network event data.
[0036] Figure 1 is a block diagram of an example network system 100 in which a Virtual Network Assistant (VNA) executes a Proactive Analysis and Correlation Engine (PACE) that is configured to dynamically build, apply, and retrain an unsupervised ML-based model for network diagnosis based on network data collected in real time.
[0037] The example network system 100 includes a plurality of sites 102A - 102N at which a network service provider manages one or more wireless networks 106A - 106N, respectively. Although in Figure 1 each site 102A - 102N is shown as including a single wireless network 106A - 106N respectively, in some examples, each site 102A - 102N may include multiple wireless networks, and the present disclosure is not limited in this regard.
[0038] Each site 102A - 102N includes a plurality of access points (APs), generally referred to as AP 142. For example, site 102A includes a plurality of APs 142A-1 to 142A-N. Similarly, site 102N includes a plurality of APs 142N-1 to 142N-N. Each AP 142 can be any type of wireless access point, including but not limited to commercial or enterprise APs, routers, or any other device capable of providing wireless network access.
[0039] Each site 102A - 102N also includes a plurality of client devices, otherwise known as user equipment devices (UEs), generally referred to as UEs 148, to represent the various wireless-enabled devices within each site. For example, a plurality of UEs 148A-1 to 148A-N are currently located at site 102A. Similarly, a plurality of UEs 148N-1 to 148N-N are currently located at site 102N. Each UE 148 can be any type of wireless client device, including but not limited to mobile devices (such as smart phones, tablet computers, or laptop computers), personal digital assistants (PDAs), wireless terminals, smart watches, smart rings, or other wearable devices. The UE 148 can also include IoT client devices, such as printers, security devices, environmental sensors, or any other device configured to communicate over one or more wireless networks.
[0040] The example network system 100 also includes various networking components for providing networking services within a wired network. By way of example, these include an Authentication, Authorization, and Accounting (AAA) server 110 for authenticating users and / or UEs 148, a Dynamic Host Configuration Protocol (DHCP) server 116 for dynamically allocating network addresses (such as IP addresses) to UEs 148 after authentication, a Domain Name System (DNS) server 122 for resolving domain names to network addresses, a plurality of servers 128 (such as web servers, database servers, file servers, etc.), and a Network Management System (NMS) 136. As Figure 1 shown, the various devices and systems of network 100 are coupled together via one or more networks 134 (such as the Internet and / or an enterprise intranet). Each of the servers 110, 116, 122, and / or 128, the APs 142, the UEs 148, the NMS 136, and any other server or device attached to or forming part of the network system 100 can include a system log or error log module, where each of these devices records the state of the device, including normal operating states and error conditions.
[0041] In Figure 1 the example, the NMS 136 is a cloud-based computing platform that manages the wireless networks 106A - 106N at one or more of the sites 102A - 102N. As further described herein, the NMS 136 provides an integrated suite of management tools and implements the various techniques of the present disclosure.
[0042] According to the techniques described herein, NMS 136 monitors SLE metrics received from wireless networks 106A - 106N at each of the sites 102A - 102N, respectively, and manages network resources, such as APs 142 at each site, to provide a high-quality wireless experience to end users, IoT devices, and clients at the sites. Generally, NMS 136 can provide a cloud-based platform for network SLE data acquisition, monitoring, activity logging, reporting, predictive analytics, network anomaly identification, and alert generation.
[0043] For example, NMS 136 can include a Virtual Network Assistant (VNA) 133 that implements an event processing platform for providing real-time insights and simplified troubleshooting for IT operations, and automatically takes corrective actions or provides recommendations to proactively address wireless network issues. The VNA 133 can include, for example, an event processing platform that is configured to process hundreds or thousands of concurrent event streams from sensors and / or agents associated with APs 142 and / or nodes within network 134. For example, according to various examples described herein, the VNA 133 of NMS 136 can include an underlying analysis and network error identification engine and an alert system. The underlying analysis engine of the VNA 133 can apply historical data and models to the inbound event streams to compute assertions, such as identified anomalies or predicted occurrences of events that constitute network error conditions. Further, the VNA 133 can provide real-time alerts and reports to notify administrators of any predicted events, anomalies, trends, and can perform root cause analysis and automate or assist in error remediation.
[0044] Further example details of operations implemented by the VNA 133 of the NMS 136 are described in U.S. Application Serial No. 14 / 788,489, filed June 30, 2015, and titled "Monitoring Wireless Access Point Events", U.S. Application Serial No. 16 / 835,757, filed March 31, 2020, and titled "Network System Fault Resolution Using a Machine Learning Model", U.S. Application Serial No. 16 / 279,243, filed February 19, 2019, and titled "Systems and Methods for a Virtual Network Assistant", U.S. Application Serial No. 16 / 237,677, filed December 31, 2018, and titled "Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault Detection", U.S. Application Serial No. 16 / 251,942, filed January 18, 2019, and titled "Method for Spatio-Temporal Modeling", and U.S. Application Serial No. 16 / 296,902, filed March 8, 2019, and titled "Method for Conveying AP Error Codes Over BLE Advertisements", the entire contents of all of which are incorporated herein by reference.
[0045] In some examples, the VNA 133 of the NMS 136 can apply machine learning techniques to identify the root cause of an error condition detected or predicted from an event data stream. If the root cause can be automatically resolved, then the VNA 133 invokes one or more corrective actions to correct the root cause of the error condition, thereby automatically improving the underlying SLE metric and also automatically improving the user experience. Further example details of the root cause analysis and automatic correction techniques performed by the NMS 136 are described in U.S. Application Serial No. 14 / 788,489, filed June 30, 2015, and titled "Monitoring Wireless Access Point Events", U.S. Application Serial No. 16 / 835,757, filed March 31, 2020, and titled "Network System Fault Resolution Using a Machine Learning Model", U.S. Application Serial No. 16 / 279,243, filed February 19, 2019, and titled "Systems and Methods for a Virtual Network Assistant", U.S. Application Serial No. 16 / 237,677, filed December 31, 2018, and titled "Methods and Apparatus for Facilitating Fault Detection and / or Predictive Fault Detection", U.S. Application Serial No. 16 / 251,942, filed January 18, 2019, and titled "Method for Spatio-Temporal Modeling", and U.S. Application Serial No. 16 / 296,902, filed March 8, 2019, and titled "Method for Conveying AP Error Codes Over BLE Advertisements", the entire contents of all of the applications being incorporated herein by reference.
[0046] In operation, NMS 136 observes, collects, and / or receives event data 139, which can take the form of data extracted from, for example, messages, counters, and statistics. According to one particular implementation, the computing device is part of the network management server 136. According to other implementations, NMS 136 can include one or more computing devices, dedicated servers, virtual machines, containers, services, or other forms of an environment for performing the techniques described herein. Similarly, the computing resources and components implementing VNA 133 and PACE 135 can be part of NMS 136, can execute on other servers or execution environments, or can be distributed to nodes (e.g., routers, switches, controllers, gateways, etc.) within the network 134.
[0047] According to one or more techniques of the present disclosure, an active analysis and correlation engine (PACE) 135 of a virtual network assistant that dynamically constructs, trains, applies, and retrains unsupervised ML model(s) 137 to event data 139 determines whether the collected network event data represents abnormal behavior that needs to be further analyzed by VNA 133 to support root cause analysis and fault resolution. More specifically, PACE 135 of the NMS applies ML model(s) 137 to network event data 139 received from network system 100, which does not include network events for a most recent observed time frame, and the ML model 137 operates to predict the occurrence levels that are expected to be seen for each type of network event during the current observed time frame. Additionally, based on the network event data 139, the ML model 137 predicts the estimated (predicted) minimum and maximum thresholds for each type of network event during the current observation period, i.e., the predicted tolerance range for the number of occurrences of each type of network event during that period. After determining that the true (actual) observations of the network event data 139 during the current observation time period deviate beyond the minimum and / or maximum thresholds predicted by the ML model 137 during that period, PACE 135 of VNA 133 marks those network events as indicating abnormal behavior, thereby triggering VNA 133 to perform a root cause analysis on those events.
[0048] The techniques of the present disclosure provide one or more advantages. For example, the techniques enable the automated virtual network assistant 133 to accurately determine which potential network problems should be subject to a more in-depth root cause analysis and which problems should be considered noise or transient problems that can be resolved in the normal course and thus can be ignored. Further, to achieve certain technical efficiencies, the techniques implement an automated virtual network assistant based on an unsupervised machine learning (ML) model, thereby reducing and / or eliminating the time-consuming task of labeling each message flow and statistic as representing a "good / normal" message flow or a "bad / failed" message flow. Additionally, the techniques support the automatic retraining of the ML-based model to adapt to changing network conditions, thereby eliminating false positives that might otherwise occur and lead to over-allocation of resources associated with root cause analysis.
[0049] Although the techniques of the present disclosure are described in this example as being performed by the NMS 136, it should be understood that the techniques described herein can be performed by any other computing device(s), system(s), and / or server(s), and the present disclosure is not limited in this regard. For example, one or more computing devices configured to perform the functionality of the techniques of the present disclosure can reside in a dedicated server or be included in any other server in addition to or apart from the NMS 136, or can be distributed across the network 100 and may or may not form part of the NMS 136.
[0050] In some examples, a network node (e.g., a router or switch within the network 134) or even an access point 142 can be configured to locally build, train, apply, and retrain unsupervised ML model(s) 137 based on locally collected SLE metrics to determine whether the collected network event data should be discarded or whether the data represents abnormal behavior that needs to be forwarded to the NMS 136 for further root cause analysis by the VNA 350 Figure 2 ) to support the identification and resolution of faults.
[0051] Figure 2 is a block diagram of an example access point (AP) device 200 configured according to one or more techniques of the present disclosure. Figure 2 The example access point 200 shown in Figure 1 can be used to implement any AP 142 as shown and described herein. The access point 200 can include, for example, a Wi-Fi, Bluetooth, and / or Bluetooth Low Energy (BLE) base station or any other type of wireless access point.
[0052] In Figure 2In the example of, the access point 200 includes a wired interface 230, wireless interfaces 220A - 220B, one or more processors 206, a memory 212, and a user interface 210 that are coupled together via a bus 214. Various components can exchange data and information through this bus. The wired interface 230 represents a physical network interface and includes a receiver 232 and a transmitter 234 for sending and receiving network communications, such as packets. The wired interface 230 directly or indirectly couples the access point 200 to Figure 1 the (one or more) networks 134. The first wireless interface 220A and the second wireless interface 220B represent wireless network interfaces and respectively include receivers 222A and 222B. Each receiver includes a receiving antenna through which the access point 200 can receive wireless signals from wireless communication devices (such as Figure 1 the UE 148). The first wireless interface 220A and the second wireless interface 220B also respectively include transmitters 224A and 224B. Each transmitter includes a transmitting antenna through which the access point 200 can transmit wireless signals to wireless communication devices (such as Figure 1 the UE 148). In some examples, the first wireless interface 220A can include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz) and the second wireless interface 220B can include a Bluetooth interface and / or a Bluetooth Low Energy (BLE) interface.
[0053] (One or more) processors 206 are programmable hardware-based processors configured to execute software instructions, such as software instructions used to define a software or computer program, stored in a computer-readable storage medium (such as the memory 212), such as a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical drive) or a memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory. This computer-readable storage medium stores instructions to cause one or more processors 206 to perform the techniques described herein.
[0054] The memory 212 includes one or more devices configured to store programming modules and / or data associated with the operation of the access point 200. For example, the memory 212 can include a computer-readable storage medium, such as a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical drive) or a memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory. This computer-readable storage medium stores instructions to cause one or more processors 206 to perform the techniques described herein.
[0055] In this example, the memory 212 stores executable software, including an application programming interface (API) 240, a communication manager 242, configuration settings 250, a device status log 252, and a data store 254. The device status log 252 includes a list of events specific to the access point 200. The events can include logs of both normal events and error events, such as, for example, memory status, reboot events, crash events, Ethernet port status, upgrade failure events, firmware upgrade events, configuration changes, etc., as well as the time and date stamps for each event. The log controller 255 determines the logging level of the device based on instructions from the NMS 136. The data 254 can store any data used and / or generated by the access point 200, including data collected from the UE 148, such as data used to compute one or more SLE metrics, which is transmitted by the access point 200 for cloud-based management of the wireless network 106A by the NMS 136.
[0056] The communication manager 242 includes program code that, when executed by the processor(s) 206, allows the access point 200 to communicate with the UE 148 and / or the network(s) 134 via any of the interface(s) 230 and / or 220A - 220C. The configuration settings 250 include any device settings for the access point 200, such as radio settings for each of the wireless interfaces 220A - 220C. These settings can be configured manually or can be remotely monitored and managed by the NMS 136 to optimize the performance of the wireless network on a regular (e.g., hourly or daily) basis.
[0057] The input / output (I / O) 210 represents physical hardware components capable of interacting with a user, such as buttons, displays, etc. Although not shown, the memory 212 generally stores executable software for controlling the user interface regarding inputs received via the I / O 210.
[0058] As described herein, the AP device 200 can measure SLE - related data (i.e., network event data) from the status log 252 and report it to the NMS 136. The SLE - related data can include various parameters indicating the performance and / or status of the wireless network. The parameters can be measured and / or determined by one or more UE devices and / or one or more APs 200 in the wireless network. The NMS 136 determines one or more SLE metrics based on the SLE - related data received from the APs in the wireless network and stores the SLE metrics as event data 139( Figure 1)。According to one or more techniques of the present disclosure, the PACE 135 of NMS136 analyzes SLE metrics (i.e., event data 139) associated with a wireless network to dynamically construct, train, apply, and retrain one or more unsupervised ML models 137 to determine whether the collected network event data represents abnormal behavior that needs to be further analyzed by the VNA 133 to support root cause analysis and fault resolution.
[0059] Figure 3 FIG. shows an example network management system (NMS) 300 configured according to one or more techniques of the present disclosure. The NMS 300 can be used to implement, for example, Figure 1 the NMS136 in. In this example, the NMS 300 is responsible for monitoring and managing one or more wireless networks 106A-106N at sites 102A-102N, respectively. In some examples, the NMS 300 receives data collected by the AP 200 from the UE 148, such as data used to compute one or more SLE metrics, and analyzes the data for cloud-based management of the wireless networks 106A-106N. In some examples, the NMS 300 can be Figure 1 part of another server shown in or part of any other server.
[0060] The NMS 300 includes a communication interface 330, one or more processors 306, a user interface 310, a memory 312, and a database 318. The various elements are coupled together via a bus 314 through which the various elements can exchange data and information.
[0061] The one or more processors 306 execute software instructions, such as software instructions that define a software or computer program, stored in a computer-readable storage medium (such as the memory 312), such as a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical drive) or a memory (such as a flash memory or a RAM) or any other type of volatile or non-volatile memory, the computer-readable storage medium storing instructions for causing the one or more processors 306 to perform the techniques described herein.
[0062] The communication interface 330 can include, for example, an Ethernet interface. The communication interface 330 couples the NMS 300 to a network and / or the Internet, such as any network(s) 134 and / or any local area network as shown in Figure 1 The communication interface 330 includes a receiver 332 and a transmitter 334 through which the NMS 300 communicates to / from the AP 142, the servers 110, 116, 122, 128, and / or any other device or system that forms part of the network 100, such as Figure 1receive / transmit data and information from / to any of those shown in FIG. The data and information received by the NMS 300 can include, for example, SLE-related or event log data received from the access point 200, which is used by the NMS 300 to remotely monitor the performance of the wireless networks 106A-106N. The NMS can also transmit data to any network device (such as the AP 142 at any network site 102A-102N) via the communication interface 330 to remotely manage the wireless networks 106A-106N.
[0063] The memory 312 includes one or more devices configured to store programming modules and / or data associated with the operation of the NMS 300. For example, the memory 312 can include a computer-readable storage medium, such as a non-transitory computer-readable medium, including a storage device (such as a disk drive or an optical drive) or a memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory, which stores instructions for causing one or more processors 306 to execute the techniques described herein.
[0064] In this example, the memory 312 includes the API 320, the SLE module 322, the Virtual Network Assistant (VNA) / AI engine 350, the Radio Resource Management (RRM) engine 360, and the Root Cause Analysis engine 370. The NMS 300 can also include any other programming modules, software engines, and / or interfaces configured for the remote monitoring and management of the wireless networks 106A-106N (including the remote monitoring and management of any AP 142 / 200).
[0065] The SLE module 322 is capable of setting and tracking thresholds for SLE metrics for each network 106A-106N. The SLE module 322 further analyzes the SLE-related data collected by the APs, such as any AP 142 from the UEs in each of the wireless networks 106A-106N. For example, the APs 142A-1 to 142A-N collect SLE-related data from the UEs 148A-1 to 148A-N currently connected to the wireless network 106A. This data is transmitted to the NMS 300, which is executed by the SLE module 322 to determine one or more SLE metrics for each of the UEs 148A-1 to 148A-N currently connected to the wireless network 106A. The one or more SLE metrics can also be aggregated to each AP at the site to gain insights into the contribution of each AP to the wireless network performance at the site. The SLE metrics track whether the service level meets the thresholds configured for each SLE metric. Each metric can also include one or more classifiers. If the metric does not meet the SLE threshold, the failure can be attributed to one of the classifiers to further understand where the failure occurred.
[0066] Example SLE metrics that can be determined by NMS 300 and their classifiers are shown in Table 1.
[0067] Table 1
[0068]
[0069]
[0070] The RRM engine 360 monitors one or more metrics for each of the sites 106A - 106N to understand and optimize the RF environment at each site. For example, the RRM engine 360 can monitor coverage and capacity SLE metrics for the wireless network 106 at site 102 to identify potential problems with SLE coverage and / or capacity in the wireless network 106, and adjust the radio settings of the access points at each site to address the identified problems. For example, the RRM engine can determine channel and transmit power distribution across all APs 142 in each of the networks 106A - 106N. For example, the RRM engine 360 can monitor events, power, channels, bandwidth, and the number of clients connected to each AP. The RRM engine 360 can also automatically change or update the configuration of one or more APs 142 at site 106 with the aim of improving coverage and capacity SLE metrics, and thus providing an improved wireless experience for users.
[0071] The VNA / AI engine 350 analyzes the data received from the APs 142 / 200 and its own data to identify when an undesirable exception state is encountered in one of the wireless networks 106A - 106N. For example, the VNA / AI engine 350 may use the root cause analysis module 370 to identify the root cause of any undesirable or exception state. In some examples, the root cause analysis module 370 utilizes artificial intelligence - based techniques to help identify the root cause of any poor SLE metrics at one or more of the wireless networks 106A - 106N. Additionally, the VNA / AI engine 350 may automatically invoke one or more corrective actions designed to address the identified root cause(s) of one or more poor SLE metrics. Examples of corrective actions that may be automatically invoked by the VNA / AI engine 350 may include, but are not limited to: invoking the RRM 360 to restart one or more APs, adjusting / modifying the transmission power of a specific radio in a specific AP, adding an SSID configuration to a specific AP, changing the channel on an AP or a set of APs, etc. Corrective actions may also include restarting switches and / or routers, invoking the download of new software to an AP, switch, or router, etc. These corrective actions are given for example purposes only, and the present disclosure is not limited in this regard. If an automatic corrective action is not available or does not adequately address the root cause, then the VNA / AI engine 350 may proactively provide a notification, including recommended corrective actions to be taken by IT personnel to resolve the network error.
[0072] According to one or more techniques of the present disclosure, the PACE 335 of the virtual network assistant that dynamically constructs, trains, applies, and retrains unsupervised ML models 337 on event data (SLE metrics 316) to determine whether the collected network event data represents anomalous behavior that needs to be further analyzed by the root cause analysis 370 of the VNA 350 to support the identification and resolution of faults.
[0073] The techniques of the present disclosure provide one or more advantages. For example, the techniques enable the automated virtual network assistant 350 to accurately determine which potential network problems should be subject to a more in - depth root cause analysis 370 and which problems should be considered noise or transient problems that can be resolved in the normal course and thus can be ignored. Further, to achieve certain technical efficiencies, the techniques implement an automated virtual network assistant based on unsupervised machine learning (ML) models 337, thereby reducing and / or eliminating the time - consuming task of labeling each message flow and statistic as representing a "good / normal" message flow or a "bad / failed" message flow. Additionally, the techniques support the automatic retraining of ML - based models to adapt to changing network conditions, thereby eliminating false positives that may otherwise occur and lead to over - allocation of resources associated with root cause analysis.
[0074] Figure 4 An example user equipment (UE) device 400 is shown. Figure 4 The example UE device 400 shown in can be used to implement any UE 148 as described and shown herein with respect to Figure 1 The UE device 400 can include any type of wireless client device, and the present disclosure is not limited in this regard. For example, the UE device 400 can include a mobile device (such as a smart phone, a tablet computer, or a laptop computer), a personal digital assistant (PDA), a wireless terminal, a smart watch, a smart ring, or any other type of mobile or wearable device. The UE 400 can also include any type of IoT client device, such as a printer, a security sensor or device, an environmental sensor, or any other connected device configured to communicate over one or more wireless networks.
[0075] According to one or more techniques of the present disclosure, one or more SLE parameter values (i.e., data used by the NMS 136 to compute one or more SLE metrics) are received from each UE 400 in the wireless network. For example, the NMS 136 receives one or more SLE parameter values from Figure 1 the UE 148 in the networks 106A - 106N of. In some examples, the NMS 136 receives the SLE parameter values from the UE 148 on a continuous basis, and the NMS can compute one or more SLE metrics for each UE on a periodic basis defined by a first predetermined time period (such as every 10 minutes or other predetermined time period).
[0076] The UE device 400 includes a wired interface 430, wireless interfaces 420A - 420C, one or more processors 406, a memory 412, and a user interface 410. The various elements are coupled together via a bus 414 through which the various elements can exchange data and information. The wired interface 430 includes a receiver 432 and a transmitter 434. If needed, the wired interface 430 can be used to couple the UE 400 to Figure 1 the (one or more) networks 134 of. The first wireless interface 420A, the second wireless interface 420B, and the third wireless interface 420C respectively include receivers 422A, 422B, and 422C, each receiver including a receiving antenna via which the UE 400 can receive wireless signals from wireless communication devices such as Figure 1 the AP 142 of, Figure 2 the AP 200 of, other UEs 148, or other devices configured for wireless communication. The first wireless interface 420A, the second wireless interface 420B, and the third wireless interface 420C respectively further include transmitters 424A, 424B, and 424C, each transmitter including a transmitting antenna via which the UE 400 can transmit to wireless communication devices such asFigure 1 AP 142 of Figure 2 AP 200, other UEs 138, and / or other devices configured for wireless communication) transmit wireless signals. In some examples, the first wireless interface 420A may include a Wi-Fi 802.11 interface (e.g., 2.4 GHz and / or 5 GHz) and the second wireless interface 420B may include a Bluetooth interface and / or a Bluetooth Low Energy interface. The third wireless interface 420C may include, for example, a cellular interface through which the UE device 400 may connect to a cellular network.
[0077] (Multiple) processors 406 execute software instructions, such as software instructions used to define a software or computer program, stored in a computer-readable storage medium (such as memory 412), such as a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical drive) or memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory, the computer-readable storage medium storing instructions to cause one or more processors 406 to perform the techniques described herein.
[0078] Memory 412 includes one or more devices configured to store programming modules and / or data associated with the operation of the UE 400. For example, memory 412 may include a computer-readable storage medium, such as a non-transitory computer-readable medium, including a storage device (e.g., a disk drive or an optical drive) or memory (such as flash memory or RAM) or any other type of volatile or non-volatile memory, the computer-readable storage medium storing instructions to cause one or more processors 406 to perform the techniques described herein.
[0079] In this example, memory 412 includes an operating system 440, applications 442, a communication module 444, configuration settings 450, and a data storage device 454. The data storage device 454 may include, for example, a status / error log that includes a list of events and / or SLE-related data specific to the UE 400. Depending on the logging level based on instructions from a network management system, the events may include logs of both normal events and error events. The data storage device 454 may store any data used and / or generated by the UE 400, such as data used to compute one or more SLE metrics, the data collected by the UE 400 and transmitted to any AP 138 in the wireless network 106 for further transmission to the NMS 136.
[0080] The communication module 444 includes program code that, when executed by the processor(s) 406, enables the UE 400 to communicate using any one of the wired interface(s) 430, wireless interfaces 420A - 420B, and / or cellular interface 450C. The configuration settings 450 include any device settings set for the UE 400 for each of the wireless interface(s) 420A - 420B and / or cellular interface 420C.
[0081] Figure 5 is a block diagram illustrating an example network node 500 configured according to the techniques described herein. In one or more examples, the network node 500 implements a device or server attached to Figure 1 the network 134, such as a router, switch, AAA server, DHCP server, DNS server, VNA, web server, etc., or network devices such as routers, switches, etc. In some embodiments, Figure 4 the network node 400 is Figure 1 the server 110, 116, 122, 128 or Figure 1 the router / switch of the network 134.
[0082] In this example, the network node 500 includes components of a communication interface 502 (such as an Ethernet interface), a processor 506, an input / output 508 (such as a display, buttons, keyboard, keypad, touch screen, mouse, etc.), a memory 512, and a component 516, such as components of a hardware module, such as components of a circuit, through which various elements can exchange data and information, coupled together via a bus 509. The communication interface 502 couples the network node 500 to a network, such as an enterprise network. Although only one interface is shown by way of example, those skilled in the art should recognize that network nodes can and typically do have multiple communication interfaces. The communication interface 502 includes a receiver 520 through which the network node 500 (such as a server) can receive data and information, such as including operation-related information, such as registration requests, AAA services, DHCP requests, Simple Notification Service (SNS) lookups, and web requests. The communication interface 502 includes a transmitter 522 through which the network node 500 (such as a server) can send data and information, such as including configuration information, authentication information, web data, etc.
[0083] Memory 512 stores executable software applications 532, an operating system 540, and data / information 530. The data 530 includes system logs and / or error logs that store SLE metrics (event data) for node 500 and / or other devices (such as wireless access points) based on a logging level according to instructions from a network management system. In some examples, network node 500 may forward SLE metrics to a network management system (e.g., Figure 1 NMS 136) of Figure 2 for analysis as described herein. Alternatively or additionally, network node 500 may provide a platform for executing PACE 135 to locally build, train, apply, and retrain (multiple) unsupervised ML models 337 based on data 530 (SLE metrics) to determine whether collected network event data should be discarded or whether the data represents abnormal behavior that needs to be forwarded to NMS 136 for further root cause analysis of VNA 350 (
[0084] PACE Event
[0085] Figure 6 is a simplified example table 600 of network events collected and used by PACE (such as PACE 135) according to the techniques described herein. NMS 136 extracts events from messages received from access points 142 and / or network nodes (such as routers and switches of network 134). Messages notifying NMS 136 of the occurrence of network events may arrive at any time after any part of the network has experienced a network event (such as network events associated with activities such as DNS, DHCP, ARP, etc.).
[0086] Example message 600 includes an ID that includes an index number for each network event listed in the rows of column 620. The simplified table includes 15 network events, but the number of events may be much larger, where the number of events depends on the number of events included in the training database. Column 610 provides indexes for the events, which can be used to simplify references to specific events. Column 630 of event dictionary table 600 provides text that can be displayed to support IT technicians or system administrators in servicing the system.
[0087] Column 640 includes data specifying the network event type and helps classify the events into specific groups of related events, and column 650 provides a more detailed text description of the nature of each network event. Information from annotation column 650 may be useful for technicians who may need to service the system but are not used by automated systems described in more detail below.
[0088] Estimating the Expected Variability and Bounds of Network Event Counters
[0089] Figure 7 Illustrated are exemplary time windows 700A, 700B, 700C through 700m+1 that are used by PACE 135 to dynamically measure the variability associated with the counter values for each different type of network event in real time (i.e., as event data 139 is collected within network system 100). In this example, the first time window of time window sequence 700A begins at time t0 and lasts for a duration of w seconds. Similarly, successive time windows begin at time t 0+w starts at time t and continues until time t 0+2w During each time window, PACE 135 monitors the arriving network events and maintains a count of the number of occurrences of each type of network event during the particular time window. In some examples, PACE 135 creates and stores the counts as a vector of network events (referred to herein as a vector of PACE events or VPEs):
[0090] VPE(t) = [c1, c2, c3, … c n Equation 1
[0091] in:
[0092] VPE(t) - network event vector at time t,
[0093] c i - the number of first network events i that occurred during said time
[0094] Window t,
[0095] t - timestamp indicating the start of the time window
[0096] i - index of network event, such as Figure 6 The index of column 610
[0097] n – number of network events.
[0098] As referenced below Figure 8 As explained in more detail in flowchart 800 of FIG. 1 , since the time origin t0 is arbitrarily chosen, PACE 135 checks and determines if the origin of the time window starts at a slightly different time, then the counters c1 to c n In this way, PACE 135 dynamically measures the variability of counter values for a time series of event data, such as a real-time stream of event data received from APs and network nodes of network system 100, using a sliding window of overlapping time windows 700a-700m offset by time increments.
[0099] Figure 8 FIG. is a flowchart of an example process 800 performed by PACE 135 for estimating the expected variability and acceptable bounds of network event counters. The specific values of each event counter parameter / element depend on the specific starting point of time t0, where the time window is set to start. For example, referring to Figure 7 , each event counter in the event counter is used to establish the network event vector described in Equation 1. The event counters measured by PACE 135 within the time window of the time window sequence 700a will be different from those measured within the time window of the time window sequence 700b (which starts after an incremental number of seconds) or within the time window of the time window sequence 700(m + 1) (which starts before an incremental number of seconds).
[0100] The process 800 performed by PACE 135 for processing event data 139 begins at operation 805 and continues to operation 810, where PACE 135 determines initial time window parameters such as the starting point t0, the window duration (e.g., 2 weeks), the starting time increment (e.g., 5 minutes), and the range (limit) within which the starting time t0 varies. PACE 135 can determine these values based on configuration data provided by a system administrator, data scientist, or other user. Although the following process refers to specific elements c of the VPE i is described, it should be understood that PACE 135 performs the same process for each of the n elements of the PACE vector (i.e., for each counter).
[0101] Once PACE 135 has determined the initial parameters in operation 810, PACE 135 proceeds to step 815, where the system measures, accumulates, and stores the total number for the network event counters for each time window in the time window sequence. An example of such multivariate time series data generated and stored by PACE 135 is illustrated in Figure 9 . In some examples, the measured and observed PACE counters are stored in a table, such as the table illustrated in Figure 10 .
[0102] In the first iteration through step 820 (which is illustrated by row 1050 in Figure 10 ), only the initial values of Table 10 exist, and as such, PACE 135 will set the measurement parameter c i (which will be c i Max in step 825) to the value of the corresponding measured VPE value for each c i within the corresponding time window. Similarly, in the first iteration through step 830, only c iThe initial parameters exist, and as such, PACE 135 sets the CiMin illustrated in column 1025 to the corresponding measured VPE parameter c illustrated by row 1050 within the corresponding time window t0 in operation 835. i These values are assigned in Figure 10 as shown.
[0103] According to the example implementation, the VPE counter value obtained when the origin of the time window is at t0 is used as C i * , and the VPE vector constructed from these counter values is called VPE*.
[0104] PACE 135 proceeds to step 840, where a new starting time point is set by incrementing or decrementing the starting time point by an increment of seconds. In operation 845, PACE 135 determines whether the new starting time of the window is still within the limits for changing the starting time of the time window. Or, in other words, whether the time origin is still within the time period [t0 - m increment,..., t0,..., t0 + m increment]. If so, then PACE 135 loops back to operation 815 and measures the VPE value in each new time window with a starting point different from the previous time window. Examples of these measurements are illustrated in Figure 10 row 1055, and the measured value or count for each element of VPE is [c1’, c2’,..c n ’]. Assuming c1’ > c1, PACE 135 sets c1’ to c1 max, assuming c2 > c2’, the method sets c2’ to c2 min, and assuming c n ’ > c1, the method sets c n ’ to c n max. Similarly, for the time window illustrated by row 1060, assuming c1” > c1’, PACE 135 sets c1” to c1 max, assuming c2” falls between c2 and c2’, PACE135 does not change the entry in the table, and assuming c n ” < c n , PACE 135 sets c n ” to c1 min.
[0105] PACE 135 continues to iterate through operations 815, 820, 825, 840, and 845 until the shift of the starting point of the time window covers a predetermined time period. This duration is based on the threshold for changing the starting time of the time window set in the initial operation 810.
[0106] Once the entire range of the starting time t0 is covered, PACE 135 moves to operation 850. For example, Figure 9shows a time period covering up to t 0+m增量 where m is a predetermined number. A similar time period can be used to shift the start time backward to t 0-m增量 (not shown in the figure). In step 850, PACE 135 determines the variability of the counts for each network event type (i.e., SLE parameters), specifically, determines the corresponding ci Max and ci Min values for each VPE counter in each time window based on the values set in the previous step.
[0107] PACE 135 proceeds to operation 855, where VPE*, VPE Max, and VPE Min are stored, so as to use the determined time-based variability for network events to enhance the unlabeled network event data 139 to generate enhanced training data for training the ML model 137, which is used to predict the estimated counts of network events and define the estimated minimum and maximum thresholds that define the corresponding ranges of expected counts for each network type.
[0108] In one example implementation, PACE 135 ends at operation 860. According to another example implementation, PACE 135 can be configured to repeat method 800 every W seconds and continuously generate a time series of the corresponding vectors VPE*, VPE Min, and VPE Max for adaptively reconstructing the ML model 137 in view of real-time network event data.
[0109] Figure 9 Illustrates a time series of the values of the VPE counter obtained by PACE 135 according to process 800 discussed above. In this example, the x-axis 910 provides time and in one example implementation provides the time t0 for the corresponding VPE counter values. The y-axis 920 provides the value of a specific c j value. For simplicity, only the value of a single counter c j is provided. Specifically, for each origin time along the x-axis 910, the figure illustrates the values of VPE* 930 and dynamically determined VPE Max 940 and VPE Min 950.
[0110] Training System Behavior Models
[0111] Figure 11 is a block diagram of component 1100 of PACE 135 that operates to train the behavior ML model 137. As described above, the time series of the historical values of the VPE counter 1110 generated by PACE 135 is used as an input to the dynamic boundary determination module 1115. As explained in more detail above, each vector of the PACE counter generated by PACE 135 corresponds to a set of sliding time windows separated by w seconds. At this time, the dynamic boundary determination module outputs the value Cj *, C j Max and C j Min. (C j *'s value is actually the same as C j ). As explained in the following references Figure 12A and Figure 12B These values are used as inputs to module 1150 to derive a function of the VPE prediction error.
[0112] In operation, PACE 135 applies the system behavior model 1120 (i.e., the ML model 137) to the received event data 139 (referred to as the historical VPE counter 1110), and this event data produces an estimated (e.g., "predicted") value for the current value of the VPE as output 1130:
[0113] [VPE t-k , VPE t-k+1 , …VPE t-1 -> Behavior model -> Predicted VPE t Equation 2
[0114] Where
[0115] [VPE t-k ,VPE t-k+1 ,…VPE t is a time series of the values of the historical VPE counter, the behavior model is Figure 11 module 1120, and
[0116] Predicted VPE t is the predicted value of the VPE counter based on the historical values of the VPE counter.
[0117] In some example implementations, the dimension of the estimated (output) VPE vector 1130 can be the same as the value of the input VPE vector 1110. According to another example implementation, the dimension of the estimated (output) VPE vector is less than the dimension of the input vector. For example, assume the input vector consists of the following two values: the number of clients and the number of clients that failed to associate with the AP. Thus, the input VPE has a dimension of 2. According to one of the example implementations, the estimated VPE may only contain the number of clients that failed to associate, so the predicted VPE has a dimension of 1.
[0118] The current VPE and the estimated VPE are used as inputs to the subtraction module 1140. It should be noted that when the dimensions of the VPE and the estimated VPE are different, only the elements common to the two vectors are utilized by the subtraction module 1140.
[0119] The output of the subtraction module 1140 is an error vector representing the VPE prediction error as follows:
[0120] VPEt Error = VPE t – Predicted VPE t Equation 3
[0121] Where:
[0122] VPE t Error - VPE prediction error
[0123] VPE t – Network element vector at time t
[0124] Predicted VPE t – Predicted vector of network elements at time t
[0125] In this way, the VPE prediction error represents the prediction (expected) vector of network elements predicting VPE t (including MAX and MIN components) and the observed vector VPE of network elements t The difference between. The VPE prediction error is used as an input to module 1150, which produces a vector as a function of the VPE prediction error, and PACE 135 uses this vector to drive the adaptation of the network event behavior model 1120 ( Figure 1 the ML model 139) based on the actually observed network event data 139.
[0126] Figure 12A and Figure 12B are graphs of different functions of the illustrated VPE error. For simplicity, only one dimension of the VPE error vector is described. The x-axis 1210 provides the VPE error. The function of the VPE error is provided on the y-axis 1220. The origin of the x-axis is at the point where the predicted VPE value is equal to VPE* for the corresponding time. The dynamic thresholds VPE Min 1230 and VPE Max 1240 are marked on the x-axis, where VPE Min is less than VPE*, and VPE Max is greater than VPE*.
[0127] As can be seen from Figure 12A when the VPE error between VPE Min and VPE Max follows the first function 1250, and when the VPE error is greater than VPE Max or less than VPE Min, the VPE error follows the second function 1260.
[0128] Figure 12B Illustrates a specific rectified linear unit (ReLU) function. The figure uses the same x and y axes 1210 and 1220 and the same numerical symbols for VPE* 1225, VPE min 1230, and VPE Max 1240. The VPE error function is set to zero when the prediction error is as follows:
[0129] VPE Min < VPE error < VPE Max Equation 4
[0130] And similarly, when the VPE error is greater than VPE Max or less than VPE Min, the VPE error function is set to a linear function of the VPE error. Equation 5 below illustrates an example of such a function.
[0131]
[0132] In either case, the function of the VPE error drives the adaptation of the parameters of the ML model 137 by the PACE 135. In a particular example implementation, the system behavior model represented by the ML model 137 is an AI-driven model such as a long short-term memory (LSTM). Other implementations may utilize other AI and machine learning (ML) algorithms to adapt the parameters of the system behavior model 1120 ( Figure 1 of the ML model 137) so as to minimize the predicted VPE error when network event data 139 is received and processed according to the techniques described herein.
[0133] According to one example implementation, the parameters of the system behavior model are continuously adapted based on the new values of the counters being streamed to the NMS 136. According to another example implementation, the process of creating the system behavior model (adapting the parameters of the model based on the counters calculated for the historical time series of network event data) is performed periodically (e.g., once every two weeks) based on the recorded network event data. Given that in some deployments, the parameters within the system behavior model (ML 137) may shift rather slowly, if at all, updating the behavior model periodically by the PACE 135 can save computational resources and still provide an adequate representation of the network system 100.
[0134] In some examples, to reduce the computational resources required, the PACE 135 may use only a subset of the VPEs rather than the entire set.
[0135] In this manner, PACE 135 constructs and adapts ML model 137 (system behavior model) to capture the normal operation ratios between various network events. For example, the number of DNS failures relative to the total number of clients and the total number of failures and successes. Specifically, in one aspect of the system behavior model, PACE 135 updates the model such that the model learns the "normal" number of DNS failures expected for a given total number of clients and total number of successes. More specifically, generally, NMS 138 and (in particular) event data 139 store information about each individual network event, such as failed attempts by clients, such as attempts to roam, attempts to obtain authentication, attempts to obtain an IP address, etc. PACE 135 applies and adapts ML model 137 as described herein to train the model to accurately predict whether network event data (e.g., failed mobility, failed authentication, etc.) represents a network event within the "normal" range and, as such, expected behavior (i.e., expected transient failures due to current operating conditions) or whether one or more of these failures are abnormal and require further analysis and / or mitigation by VNA 133.
[0136] System Behavior Model Update and Validation Using Error Histograms
[0137] According to a specific implementation, PACE 135 of NMS 136 constructs a system behavior model (ML model 137) based on event data 139 collected from multiple sites 102. Such a system behavior model is referred to herein as a global system behavior model. According to one implementation, once the global model is constructed, its parameters are fixed, and the VPE data (or just a subset of the data) used to construct the global model is run through the model, thereby recording the prediction error. Similarly, according to another implementation, PACE 135 records the value of the prediction error only when the value of the prediction error is outside the VPE, outside the dynamically determined range [VPE Min, VPE Max].
[0138] Figure 14 is a flowchart 1400 illustrating an example process by which an active analysis and correlation engine (e.g., Figure 1 PACE 135) adaptively updates a network event ML-based behavior model (e.g., ML model 137). For purposes of illustration, the example process will be described with respect to Figure 1 PACE 135.
[0139] In this example, PACE 135 starts at operation 1405 and proceeds to operation 1410, where PACE 135 receives or otherwise collects a new time series of network event data for a most recent time period (e.g., for the most recent two weeks). At operation 1415, PACE 135 divides the information from the time series into two subsets based on timestamps, also referred to herein as time series subgroups, each subgroup having a different date / time range. PACE 135 uses the first subset (e.g., the first two-thirds of the collected network event data) to build a PACE behavior model (ML model 137) in operation 1420 and uses the second subset (e.g., the remaining one-third of the data) to test the built behavior model and build a prediction error histogram / probability function in operation 1425 and in operation 1430.
[0140] In some examples, PACE 135 generates and utilizes a prediction error probability function for at least two different technical advantages: a) ensuring that the network events used to build the PACE behavior model are associated with normal behavior, and b) ensuring that the parameters tuned for the PACE behavior model provide an appropriate representation of the underlying behavior (including ratios) of various network events. As explained in more detail below, the validation process performed by PACE 135 utilizes a histogram of prediction errors (such as the prediction error at the output of subtraction module 1140). According to yet another implementation, PACE 135 builds a histogram based on the output of a function of VPE error module 1150. An example of such a prediction error histogram is shown in Figure 13 FIG. The x-axis 1310 depicts the prediction error (or, in another specific implementation, the amount by which the prediction exceeds a dynamically determined range [VPE Min, VPE Max]). The y-axis 1320 provides the probability of a particular error value occurring during operation across all sites using the second subset of network event data (e.g., the subset used to test the effectiveness of the PACE behavior model). According to one aspect of the present invention, the fact that most errors have small values provides validation for the successful construction of the system model.
[0141] In operation 1435, PACE 135 compares the histogram / probability of the VPE error from the newly built model with the error histogram / probability obtained as part of the training of the previous behavior model. If in decision operation 1440, PACE 135 determines that the new error probability distribution / histogram is not similar to the previous error distribution (e.g., via using the Kullback-Leibler divergence measure described below), then PACE 135 proceeds to operation 1445 and excludes the outlier data points / events from the training data. In some examples, PACE 135 may generate an output to notify the user, administrator, or data scientist about the event to guide the user in excluding the outlier data points. PACE 135 loops back to operation 1415 and repeats the training process using the cleaned data excluding the outlier events, thereby adaptively retraining the ML model 137 as needed based on the real-time network event data.
[0142] Once PACE 135 determines that the new prediction error is similar to the previous prediction error distribution function (decision operation 1440), PACE 135 terminates the training process and proceeds to operation 1450 and starts using the newly built VPE behavior model to process subsequent event data. For example, when the network event system behavior model is updated periodically, such as once every two weeks, the currently built histogram of the prediction error is compared with the previous prediction error, such as the prediction error histograms from two weeks, four weeks, six weeks, etc. ago. According to yet another implementation, if the two histograms are highly similar or correlated, then the new network event behavior model is validated and the parameters of the newly built system are used. PACE 135 completes the retraining and new build cycle in operation 1455.
[0143] In one specific implementation, PACE 135 performs the comparison of operation 1440 by determining the correlation between the past and the histogram and the current prediction error. If the correlation is greater than a specific threshold, such as 0.8, then PACE 135 determines that the histogram is similar enough. According to another example implementation, the method uses information theory to determine the similarity between the histograms representing the error probabilities. Specifically, the method employs the Kullback-Leibler divergence (also known as relative entropy or KLD) to determine a measure of how different one probability distribution is from a second reference probability distribution based on the past histogram.
[0144] Identifying Behavioral Anomalies Using System Behavior Models
[0145] Figure 15FIG. 1500 is a flow chart of an example process performed by virtual network assistant 133 for performing root cause analysis of network event data triggered by PACE 135, i.e., when PACE 135 applies an ML model 137 and calculates a VPE prediction error that is outside a dynamically determined range (boundary) for one or more types of network events, indicating that a more in-depth analysis of abnormal network behavior associated with these network events needs to be performed. For purposes of example, the example process will be described with respect to Figure 1 VNA 133 and PACE 135.
[0146] In this example, PACE 135 of VNA 133 begins at operation 1505 and proceeds to operation 1510, where a system behavior model is constructed and parameters are fixed, as described herein. At operation 1515, PACE 135 feeds real-time VPE information (vectorized input generated from network event data 139) into ML model 137 and determines a VPE prediction error at step 1520. PACE 135 proceeds to decision operation 1525 and checks whether the VPE prediction error based on the predicted VPE generated by ML model 137 is within the dynamically determined range [VPE Min, VPE Max] for each parameter.
[0147] If PACE 135 determines that the VPE prediction error falls within the dynamically determined range, then PACE 135 loops back to operation 1515 and similarly processes any newly received network event data 139. In this way, as long as the prediction error is within the dynamically determined range, the behavior of the entire network system 100 is considered normal, even though individual clients in the system may report failures, such as ARP failures, DHCP failures, or DNS failures. For example, in a system with a large number (e.g., 1000 clients), it is normal to experience a few (e.g., 1 or 2) failures such as those described above.
[0148] However, if the VPE prediction error exceeds the dynamically determined prediction error range established by [VPE Min, VPE Max], then PACE 135 determines that an abnormality in network behavior has been detected, i.e., corresponding failures or other network events are occurring at a frequency outside the expected range. PACE 135 proceeds to operation 1530 and determines the type of network event exhibiting the abnormal behavior. For example, PACE 135 may use Figure 6 type information 640 of table 600 of
[0149] In operation 1535, in response to the identification of one or more abnormal network events by PACE 135, NMS 136 invokes the Virtual Network Assistant 133 to perform a more in-depth analysis of the event data 139 to determine the root cause of the detected abnormality. In decision operation 1540, the VNA 133 determines whether a remedial action can be invoked based on the root cause analysis, such as restarting or reconfiguring one or more APs 142, network nodes, or other components or by outputting a scripted action for an administrator to follow. If a remedial corrective action is available, then at operation 1545, the VNA 133 invokes the remedial action, such as restarting a particular device, a component of a device, a module of a device, etc. In either case, if the remedial action is available or unavailable, the VNA 133 proceeds to operation 1550 and outputs a report / alert to notify the technician of the identified problem and / or the corrective action taken by the VNA 133 to automatically remedy the underlying root cause of the problem. After invoking the VNA 133, the PACE 135 loops back to operation 1515 and continues to process the real-time event data 139 received from the network system 100.
[0150] The techniques described herein can be implemented using software, hardware, and / or a combination of software and hardware. Various examples relate to apparatuses, such as mobile nodes, mobile wireless terminals, base stations, such as access points, communication systems. Various examples also relate to methods, such as methods of controlling and / or operating communication devices (such as wireless terminals (UEs), base stations, control nodes, access points, and / or communication systems). Various examples also relate to non-transitory machines, such as computers, readable media, such as ROM, RAM, CDs, hard disks, etc., which include machine-readable instructions for controlling the machine to implement one or more steps of the method.
[0151] It should be understood that any specific order or hierarchy of steps in the disclosed processes is an example of an example method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged while remaining within the scope of the present disclosure. The appended method claims present the elements of the various steps in a sample order and are not meant to be limited to the particular order or hierarchy presented.
[0152] In various examples, the devices and nodes described herein are implemented using one or more modules to perform steps corresponding to one or more methods, such as signal generation, transmission, processing, and / or reception steps. Thus, in some examples, various features are implemented using modules. Such modules can be implemented using software, hardware, or a combination of software and hardware. In some examples, each module is implemented as an individual circuit, where the device or system includes separate circuits for implementing the functions corresponding to each described module. Many of the methods or method steps described above can be implemented using machine-executable instructions (such as software, included in a machine-readable medium such as a memory device, e.g., RAM, floppy disk, etc.) to control a machine, such as a general-purpose computer with or without additional hardware, to implement all or part of the methods described above in, for example, one or more nodes. Thus, among other things, various examples relate to a machine-readable medium, such as a non-transitory computer-readable medium, including machine-executable instructions for causing a machine (such as a processor and associated hardware) to perform one or more steps of the (multiple) methods described above. Some examples relate to a device including a processor configured to implement one, multiple, or all steps of one or more methods of an example aspect.
[0153] In some examples, one or more processors (such as a CPU) of one or more devices (such as a communication device (such as a wireless terminal (UE)) and / or an access node) are configured to perform the steps of the methods described as being performed by the device. The configuration of the processor can be implemented by using one or more modules (such as software modules) to control the processor configuration and / or by including hardware (such as a hardware module) in the processor to perform the listed steps and / or control the processor configuration. Thus, some but not all examples relate to a communication device having a processor, such as a user equipment, the processor including modules corresponding to each step of the various described methods performed by the device in which the processor is included. In some but not all examples, the communication device includes modules corresponding to each step of the various described methods performed by the device in which the processor is included. The modules can be implemented purely in hardware, such as as a circuit, or can be implemented using software and / or hardware or a combination of software and hardware.
[0154] Some examples relate to a computer program product including a computer-readable medium including code for causing a computer or computers to implement various functions, steps, actions, and / or operations (such as one or more of the steps described above). In some examples, the computer program product can and sometimes does include different code for each step to be performed. Thus, the computer program product can and sometimes does include code for each individual step of a method (such as a method of operating a communication device (such as a wireless terminal or node)). The code can be in the form of machine (such as a computer) executable instructions stored on a computer-readable medium such as RAM (Random Access Memory), ROM (Read-Only Memory), or other types of storage devices. In addition to relating to a computer program product, some examples relate to a processor configured to implement one or more of the various functions, steps, actions, and / or operations described above. Thus, some examples relate to a processor configured to implement some or all of the steps of the methods described herein, such as a CPU, a graphics processing unit (GPU), a digital signal processing (DSP) unit, etc. The processor can be used, for example, in the communication devices or other devices described in this application.
[0155] In view of the foregoing description, many additional variations of the methods and apparatuses of the various examples described above will be apparent to those skilled in the art. Such variations will be considered to be within the scope of this disclosure. The methods and apparatuses can and in various examples are used in conjunction with BLE, LTE, CDMA, orthogonal frequency division multiplexing (OFDM), and / or various other types of communication technologies that can be used to provide a wireless communication link between an access node and a mobile node. In some examples, the access node is implemented as a base station that uses OFDM and / or CDMA to establish a communication link with a user equipment device (such as a mobile node). In various examples, the mobile node is implemented as a laptop computer, a personal data assistant (PDA), or other portable device, including receiver / transmitter circuitry and logic and / or routines for implementing the method.
[0156] In the detailed description, numerous specific details are set forth in order to provide a thorough understanding of some examples. However, those of ordinary skill in the art will understand that some examples can be practiced without these specific details. In other instances, well-known methods, procedures, components, units, and / or circuits have not been described in detail so as not to obscure the discussion.
[0157] Some examples can be used in conjunction with a variety of devices and systems, such as user equipment (UE), mobile device (MD), wireless station (STA), wireless terminal (WT), personal computer (PC), desktop computer, mobile computer, laptop computer, notebook computer, tablet computer, server computer, handheld computer, handheld device, personal digital assistant (PDA) device, handheld PDA device, in-vehicle device, out-of-vehicle device, hybrid device, vehicle-mounted device, non-vehicle-mounted device, mobile or portable device, consumer device, non-mobile or non-portable device, wireless communication station, wireless communication device, wireless access point (AP), wired or wireless router, wired or wireless modem, video device, audio device, audio-video (A / V) device, wired or wireless network, wireless local area network, wireless video area network (WVAN), local area network (LAN), wireless LAN (WLAN), personal area network (PAN), wireless PAN (WPAN), etc.
[0158] Some examples can be used in combination with the following: devices and / or networks operating according to existing Wireless Gigabit Alliance (WGA) specifications (Wireless Gigabit Alliance, WiGig MAC and PHY Specification Version 1.1, April 2011, Final Specification) and / or future versions and / or their derivatives; devices and / or networks operating according to existing IEEE 802.11 standards (IEEE 802.11-2012, IEEE Standard for Information technology--Telecommunications and information exchange between systems Local and metropolitan area networks--Specific requirements Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specifications, March 29, 2012; IEEE 802.11ac-2013 (“IEEE P802.11ac-2013, IEEE Standard for Information Technology-Telecommunications and Information Exchange Between Systems-Local and Metropolitan Area Networks-Specific Requirements-Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specifications-Amendment 4: Enhancements for Very High Throughput for Operation in Bands below 6GHz”, December 2013); IEEE 802.11ad (“IEEE P802.11ad - 2012, "IEEE Standard for Information Technology - Telecommunications and Information Exchange Between Systems - Local and Metropolitan Area Networks - Specific Requirements - Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specifications - Amendment 3: Enhancements for Very High Throughput in the 60GHz Band” (December 28, 2012); IEEE - 802.11REVmc (“IEEE 802.11 - REVmcTM / D3.0, June 2014 draft standard for Information technology - Telecommunications and information exchange between systems Local and metropolitan area networks Specific requirements; Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specification”); IEEE802.11 - ay (P802.11ay Standard for Information Technology - Telecommunications and Information Exchange Between Systems Local and Metropolitan Area Networks - Specific Requirements Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specifications - Amendment: Enhanced Throughput for Operation in License - Exempt Bands Above 45GHz)), IEEE 802.Operate with 11 - 2016 and / or future versions and / or their derivatives; devices and / or networks that operate according to the existing Wi-Fi Alliance (WFA) Peer-to-Peer (P2P) specification (Wi-Fi P2P technical specification, version 1.5, August 2014) and / or future versions and / or their derivatives; devices and / or networks that operate according to existing cellular specifications and / or protocols (e.g., 3rd Generation Partnership Project (3GPP), 3GPP Long Term Evolution (LTE)) and / or future versions and / or their derivatives; units and / or devices that are part of the above networks or operate using any one or more of the above protocols.
[0159] Some examples can be used in combination with the following: one-way and / or two-way radio communication systems, cellular radiotelephone communication systems, mobile phones, cellular phones, wireless phones, Personal Communication System (PCS) devices, PDA devices incorporating wireless communication devices, mobile or portable Global Positioning System (GPS) devices, devices incorporating GPS receivers or transceivers or chips, devices incorporating RFID elements or chips, Multiple Input Multiple Output (MIMO) transceivers or devices, Single Input Multiple Output (SIMO) transceivers or devices, Multiple Input Single Output (MISO) transceivers or devices, devices with one or more internal antennas and / or external antennas, Digital Video Broadcast (DVB) devices or systems, multi-standard radio devices or systems, wired or wireless handheld devices (e.g., smartphones), Wireless Application Protocol (WAP) devices, etc.
[0160] Some examples can be used in combination with one or more types of wireless communication signals and / or systems (e.g., Radio Frequency (RF), Infrared (IR), Frequency Division Multiplexing (FDM), Orthogonal FDM (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), FDM Time Division Multiplexing (TDM), Time Division Multiple Access (TDMA), Multi-User MIMO (MU-MIMO), Space Division Multiple Access (SDMA), Extended TDMA (E-TDMA), General Packet Radio Service (GPRS), Extended GPRS, Code Division Multiple Access (CDMA), Wideband CDMA (WCDMA), CDMA 2000, Single-Carrier CDMA, Multi-Carrier CDMA, Multi-Carrier Modulation (MDM), Discrete Multi-Tone (DMT), Bluetooth, Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee TM, Ultra-wideband (UWB), Global System for Mobile Communications (GSM), 2G, 2.5G, 3G, 3.5G, 4G, fifth generation (5G) or sixth generation (6G) mobile networks, 3GPP, Long Term Evolution (LTE), LTE-Advanced, Enhanced Data Rates for GSM Evolution (EDGE), etc. are used in combination. Other examples can be used for various other devices, systems, and / or networks.
[0161] Some illustrative examples can be used in combination with WLAN (Wireless Local Area Network) (e.g., Wi-Fi network). Other examples can be used in combination with any other suitable wireless communication network (e.g., wireless local area network, "piconet", WPAN, WVAN, etc.).
[0162] Some examples can be used in combination with wireless communication networks that communicate in the bands of 2.4 GHz, 5 GHz, and / or 60 GHz. However, other examples can be implemented using (any) other suitable wireless communication bands (e.g., the extremely high frequency (EHF) band (millimeter wave (mmWave) band), e.g., a band between 20 GHz and 300 GHz, WLAN band, WPAN band, bands within the bands according to the WGA specification, etc.).
[0163] Although only some simple examples of various device configurations are provided above, it should be understood that many variations and permutations are possible. Additionally, the technology is not limited to any specific channel, but generally applies to (any) frequency range / (any) channel. Additionally, as discussed, the technology may be useful in unlicensed spectrum.
[0164] Although the examples are not limited in this regard, discussions using terms such as, for example, "processing", "computing", "operation", "determining", "establishing", "analyzing", "checking", etc. can refer to the (operations) and / or (processes) of a computer, computing platform, computing system, communication system or subsystem, or other electronic computing device that manipulates data of physical (e.g., electronic) quantities within the registers and / or memory of the computer and / or transforms the data into other data of physical quantities similarly represented within the registers and / or memory of the computer or other information storage medium that can store instructions to perform the operations and / or processes.
[0165] Although the examples are not limited in this regard, as used herein, the terms "plurality" and "a plurality" can include, for example, "a number" or "two or more." The terms "plurality" and "a plurality" can be used throughout the specification to describe two or more components, devices, elements, units, parameters, circuits, etc. For example, "a plurality of stations" can include two or more stations.
[0166] It may be advantageous to set forth definitions of certain words and phrases used throughout this document: The terms "include" and "comprise" and their derivatives mean including but not limited to; the term "or" is inclusive and means and / or; the phrases "associated with" and "associated therewith" and their derivatives may mean including, being included within, interconnecting with, interconnected with, containing, being contained within, connected to or connecting with, coupled to or coupling with, being compatible with, cooperating with, interleaving, juxtaposing, being adjacent to, being bound or binding to, having, possessing, etc.; and the term "controller" means any device, system, or part thereof that controls at least one operation, and such device can be implemented in hardware, circuitry, firmware, software, or some combination of at least two of them. It should be noted that the functionality associated with any particular controller can be centralized or distributed, whether local or remote. Definitions of certain words and phrases are provided throughout this document, and one of ordinary skill in the art should understand that in many cases, if not most cases, such definitions apply to the prior and future use of such defined words and phrases.
[0167] Examples have been described with respect to communication systems and protocols, techniques, means, and methods for performing communication in, for example, a wireless network or generally in any communication network operating using (a) any communication protocol. Examples of such are home or access networks, wireless home networks, wireless corporate networks, etc. However, it should be understood that, in general, the systems, methods, and techniques disclosed herein will equally apply to other types of communication environments, networks, and / or protocols.
[0168] For purposes of explanation, numerous details are set forth in order to provide a thorough understanding of the technology. However, it should be understood that the disclosure may be practiced in many ways beyond the specific details set forth herein. Further, while the examples illustrated herein show various components of the system juxtaposed, it should be understood that the various components of the system may be located at remote portions of a distributed network (such as a communication network, nodes), within a domain host and / or the Internet or a dedicated secure, non-secure, and / or encrypted system, and / or within a network operation or management device internal or external to the network. As an example, a domain host may also be used to refer to any device, system, or module that manages and / or configures any one or more aspects of a network or communication environment and / or one or more of the (multiple) transceivers and / or stations and / or (multiple) access points described herein or communicates therewith.
[0169] Accordingly, it should be understood that the components of the system may be combined into one or more devices, or split among devices (such as transceivers, access points, stations, domain hosts, network operation or management devices, nodes), or juxtaposed on a particular node of a distributed network (such as a communication network). As should be understood from the following description, and for reasons of computational efficiency, the components of the system may be arranged anywhere within the distributed network without affecting its operation. For example, the various components may be located at a domain host, node, domain management device (such as a MIB), network operation or management device, (multiple) transceivers, stations, (multiple) access points, or some combination thereof. Similarly, one or more functional portions of the system may be distributed between a transceiver and an associated computing device / system.
[0170] In addition, it should be understood that the various links (including the lines of any communication channel / element / connecting element) may be wired or wireless links or any combination thereof, or any other known or later developed element capable of providing and supplying and / or communicating data to and from the connected elements. As used herein, the term module may refer to any known or later developed hardware, circuitry, software, firmware, or combination thereof capable of performing the functionality associated with that element. As used herein, the terms determine, operate, and calculate and their variants may be used interchangeably and include any type of methodology, process, technique, mathematical operation, or protocol.
[0171] Further, while some of the examples described herein relate to a transmitter portion of a transceiver performing certain functions or a receiver portion of a transceiver performing certain functions, the disclosure is intended to include corresponding and complementary transmitter-side or receiver-side functionality, respectively, in both the same transceiver and / or one or more other transceivers, and vice versa.
[0172] Examples regarding enhanced communication are described. However, it should be understood that, in general, the systems and methods herein will be equally applicable to any type of communication system that utilizes any one or more protocols (including wired communication, wireless communication, power line communication, coaxial cable communication, fiber optic communication, etc.) in any environment.
[0173] Example systems and methods regarding IEEE 802.11 and / or and / or Low-energy transceivers and associated communication hardware, software, and communication channels are described. However, to avoid unnecessarily obscuring the present disclosure, the following description omits well-known structures and devices that may be shown in block diagram form or otherwise summarized.
[0174] Although the flowcharts described above have been discussed with respect to a particular sequence of events, it should be understood that changes to that sequence can occur without materially affecting the operation of the examples. Additionally, the example techniques illustrated herein are not limited to the specifically illustrated examples, but can also be utilized with other examples, and each described feature can be individually and separately claimed.
[0175] The systems described above can be implemented on wireless telecommunications devices / systems (such as IEEE 802.11 transceivers, etc.). Examples of wireless protocols that can be used with this technology include IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, IEEE 802.11ac, IEEE 802.11ad, IEEE 802.11af, IEEE 802.11ah, IEEE 802.11ai, IEEE 802.11aj, IEEE 802.11aq, IEEE 802.11ax, Wi-Fi, LTE, 4G, WirelessHD, WiGig, WiGi, 3GPP, Wireless LAN, WiMAX, DensiFi SIG, UnifiSIG, 3GPP LAA (Licensed Assisted Access), etc.
[0176] Additionally, systems, methods, and protocols can be implemented to improve one or more of a dedicated computer, a programmed microprocessor or microcontroller, and (multiple) peripheral integrated circuit elements, ASICs, or other integrated circuits, digital signal processors, hardwired electronic or logic circuits (such as discrete element circuits), programmable logic devices (such as PLDs, PLAs, FPGAs, PALs, modems, transceivers, any comparable devices, etc.). Generally, any device capable of implementing a state machine can benefit from the various communication methods, protocols, and technologies according to the disclosure provided herein, and the state machine can in turn implement the methodologies illustrated herein.
[0177] Examples of processors as described herein can include, but are not limited to, at least one of the following: 800 and 801, with 4G LTE integration and 64-bit computing 610 and 615, with 64-bit architecture A7 processor M7 motion coprocessor series Core TM processor series, Intel processor series i5-4670K and i7-4770K 22nm Haswell i5-3570K 22nm Ivy Bridge FX TM processor series FX-4300, FX-6300, and FX-8350 32nm Vishera Kaveri processor, Texas Jacinto C6000 TM Automotive infotainment processor, Texas OMAP TM Automotive-grade mobile processor Cortex TM -M processor Cortex-A and ARM926EJ-S TM processor AirForce BCM4704 / BCM4703 wireless networking processor, AR7100 wireless network processing unit, other industry-equivalent processors, and can perform computing functions using any known or future-developed standards, instruction sets, libraries, and / or architectures.
[0178] In addition, the disclosed method can be easily implemented in software using an object or object-oriented software development environment that provides portable source code that can be used on various computer or workstation platforms. Alternatively, the disclosed system can be implemented partially or fully in hardware using standard logic circuits or VLSI designs. Whether software or hardware is used to implement the system according to the example depends on the speed and / or efficiency requirements of the system, the particular functions, and the particular software or hardware system or microprocessor or microcomputer system being utilized. The communication systems, methods, and protocols illustrated herein can be easily implemented in hardware and / or software by those of ordinary skill in the applicable art using any known or later-developed systems or architectures, devices, and / or software based on the functional descriptions provided herein and having general knowledge in the fields of computers and telecommunications.
[0179] In addition, the disclosed technology can be easily implemented in software and / or firmware that can be stored on a storage medium to enhance the performance of a general-purpose computer programmed in cooperation with a controller and memory, a special-purpose computer, a microprocessor, etc. In these cases, the system and method can be implemented as a program embedded in a personal computer, such as a JAVA.RTM. applet or a CGI script, as a resource resident on a server or computer workstation, as a routine embedded in a special-purpose communication system or system component, etc. The system can also be implemented by physically incorporating the system and / or method into a software and / or hardware system, such as the hardware and software system of a communication transceiver.
[0180] Therefore, it is apparent that systems and methods for enhancing and improving the session user interface have been provided at least. Many alternatives, modifications, and variations will be or will become apparent to those of ordinary skill in the applicable art. Accordingly, the present disclosure is intended to embrace all such alternatives, modifications, equivalents, and variations that fall within the spirit and scope of the present disclosure.
Claims
1. A network management system, comprising: a memory; and one or more processors coupled to the memory and configured to: apply an unsupervised machine learning model to network event data to determine a predicted occurrence measure of network events of a first event type among one or more event types during an observation period; determine a prediction error based on the observed network event data during the observation period, the prediction error indicating a difference between the predicted occurrence measure and an observed occurrence measure of network events of the first event type; and detect abnormal network behavior based on the prediction error being outside a threshold occurrence range for network events of the first event type.
2. The system according to claim 1, wherein the unsupervised machine learning model has been trained using unlabeled training data, the unlabeled training data including network event data observed during a plurality of overlapping time windows having different time origins and augmented with a threshold occurrence range for network events of the one or more event types.
3. The system according to claim 2, wherein the one or more processors are further configured to determine a minimum threshold and a maximum threshold defining the threshold occurrence range for network events of the one or more event types based on the network event data observed during the plurality of overlapping time windows having different time origins.
4. The system according to claim 1, wherein, in order to determine the minimum threshold and the maximum threshold, the one or more processors are further configured to: for each of the plurality of overlapping time windows, determine an occurrence measure of network events of each event type among the one or more event types; for each event type, determine the maximum threshold as the maximum occurrence measure of network events of the corresponding event type during any of the plurality of overlapping time windows; and for each event type, determine the minimum threshold as the minimum occurrence measure of network events of the corresponding event type during any of the plurality of overlapping time windows.
5. The system according to any one of claims 1 to 4, wherein the one or more processors are further configured to: perform a root cause analysis of the observed network event data based on detecting the abnormal network behavior; and determine a remediation action based on one or more identified root causes of the abnormal network behavior.
6. A method for network management, comprising: applying, by a computing system, an unsupervised machine learning model to network event data to determine a predicted occurrence measure of network events of a first event type among one or more event types during an observation period; determining, by the computing system, a prediction error based on the observed network event data during the observation period, the prediction error indicating a difference between the predicted occurrence measure and an observed occurrence measure of network events of the first event type; and The computing system detects anomalous network behavior based on the prediction error being outside a threshold occurrence range of network events for the first event type.
7. The method of claim 6, wherein the unsupervised machine learning model has been trained using unlabeled training data, the unlabeled training data including network event data that is observed during a plurality of overlapping time windows having different time origins and is augmented with threshold occurrence ranges of network events for the one or more event types.
8. The method of claim 7, further comprising determining a minimum threshold and a maximum threshold that define the threshold occurrence range of network events for the one or more event types based on the network event data observed during the plurality of overlapping time windows having different time origins.
9. The method of claim 8, wherein determining the minimum threshold and the maximum threshold further comprises: For each of the plurality of overlapping time windows, determining an occurrence measurement of network events for each of the one or more event types; For each event type, determining the maximum threshold as the maximum occurrence measurement of network events for the corresponding event type during any of the plurality of overlapping time windows; And For each event type, determining the minimum threshold as the minimum occurrence measurement of network events for the corresponding event type during any of the plurality of overlapping time windows.
10. The method of any one of claims 6 to 9, further comprising: Performing a root cause analysis of the observed network event data based on detecting the anomalous network behavior; And Determining a remedial action based on one or more identified root causes of the anomalous network behavior.
11. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to: Apply an unsupervised machine learning model to network event data to determine a predicted occurrence measurement of network events of a first event type among one or more event types during an observation period; Determine a prediction error based on the observed network event data during the observation period, the prediction error indicating a difference between the predicted occurrence measurement and an observed occurrence measurement of network events of the first event type; And Detect anomalous network behavior based on the prediction error being outside a threshold occurrence range of network events for the first event type.
12. The computer-readable medium of claim 11, wherein the unsupervised machine learning model has been trained using unlabeled training data, the unlabeled training data including network event data that is observed during a plurality of overlapping time windows having different time origins and is augmented with threshold occurrence ranges of network events for the one or more event types.
13. The computer-readable medium according to claim 12, wherein the instructions further cause the one or more programmable processors to determine a minimum threshold and a maximum threshold that define a threshold occurrence range for network events for the one or more event types based on the network event data observed during the plurality of overlapping time windows having different time origins.
14. The computer-readable medium according to claim 11, wherein to determine the minimum threshold and the maximum threshold, the instructions further cause the one or more programmable processors to: For each of the plurality of overlapping time windows, determine an occurrence measurement for network events for each of the one or more event types; For each event type, determine the maximum threshold as the maximum occurrence measurement of the network events of the corresponding event type during any of the plurality of overlapping time windows; And For each event type, determine the minimum threshold as the minimum occurrence measurement of the network events of the corresponding event type during any of the plurality of overlapping time windows.
15. The computer-readable medium according to any one of claims 11 to 14, wherein the instructions further cause the one or more programmable processors to: Perform a root cause analysis of the observed network event data based on detecting the abnormal network behavior; and Determine a remedial action based on one or more identified root causes of the abnormal network behavior.
Citation Information
Patent Citations
Methods and apparatus for facilitating fault detection and / or predictive fault detection
US10958585B2
Systems and methods for a virtual network assistant
US10985969B2
Monitoring wireless access point events
US20170005886A1
Method for spatio-temporal monitoring
US20200236008A1
Method for conveying AP error codes over BLE advertisements
US20200287782A1