Mechanism for intelligent and comprehensive monitoring system using peer-to-peer agents in a network
By introducing peer-to-peer agents into the network, collaboration between hardware and software monitoring devices is achieved, solving the problems of device isolation and information aggregation, and providing global network monitoring and intelligent fault detection capabilities.
Patent Information
- Application Number
- CN202311321156.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-16
- Filing Date
- 2023-10-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-10-12
AI Technical Summary
In existing technologies, network monitoring devices lack visibility, and the isolation between devices makes it impossible to collaborate, making it difficult to achieve global network monitoring. Furthermore, the use of different monitoring solutions by different devices makes it difficult to aggregate and present information.
By achieving peer-to-peer integration between hardware and software monitoring devices in the network, collaborative monitoring is carried out using peer-to-peer agents, devices communicate directly and execute pre-configured conditions to trigger actions, providing a comprehensive intelligent monitoring system.
It enables collaborative monitoring among network devices, provides a global network view, improves fault detection and root cause analysis capabilities, and simplifies information aggregation and presentation.
Smart Images

Figure CN118677800B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Network monitoring of enterprise systems can be performed using a combination of hardware and software-based solutions. Test agents in a cloud-based configuration can communicate directly with their cloud backend using the same network that the test agent is to monitor, and the test agent can then perform monitoring tasks and report results to the cloud backend. However, some limitations can include a lack of visibility of various test agents, different capabilities of devices associated with test agents, and the use of different monitoring solutions for different test agents. BRIEF DESCRIPTION OF DRAWINGS
[0002] Figure 1 An environment showing a conventional monitoring model with different endpoints according to the prior art is shown;
[0003] Figure 2 An environment showing a mechanism to facilitate an intelligent and comprehensive monitoring system using peer agents according to an aspect of the present application is shown;
[0004] Figure 3 An environment showing a mechanism to facilitate an intelligent and comprehensive monitoring system using peer agents according to an aspect of the present application is shown; Figure 2 including communication between peer devices in the event of uplink failure from a device;
[0005] Figure 4 An exemplary user dashboard with displayed results from an intelligent and comprehensive monitoring system according to an aspect of the present application is shown;
[0006] Figure 5A A flowchart showing a method to illustrate a mechanism to facilitate an intelligent and comprehensive monitoring system using peer agents in a network according to an aspect of the present application is shown;
[0007] Figure 5B A flowchart showing a method to illustrate a mechanism to facilitate an intelligent and comprehensive monitoring system using peer agents in a network according to an aspect of the present application is shown; and
[0008] Figure 6 A computer system and apparatus to facilitate a mechanism to facilitate an intelligent and comprehensive monitoring system using peer agents in a network according to an aspect of the present application is shown.
[0009] In the drawings, like reference numerals refer to same elements throughout. DETAILED DESCRIPTION
[0010] Aspects of the invention address the limitations of siloed monitoring devices by providing a system that utilizes peer-to-peer integration between different types of hardware and software monitoring devices in a network. The described aspects can also perform detection of preconfigured conditions (e.g., monitored network metrics or test results) that can trigger predetermined corresponding actions, thereby creating a comprehensive intelligent monitoring system on the network.
[0011] Network monitoring of enterprise systems can be performed using a combination of hardware and software based solutions. Hardware based solutions can include dedicated hardware platforms that perform continuous network monitoring. One example can be a purpose built device that is placed in strategic locations within a network to provide network and user based monitoring to ensure uptime, performance, and quality of service (QOS), and to guarantee SLAs (service level agreements). Software based solutions can include client software products that measure network metrics, for example, application level software products that run on user provided host hardware that also provide monitoring of desired metrics from the perspective of a given device.
[0012] In a cloud based configuration, test agents can typically communicate directly with their cloud backend using the same network that the test agent is monitoring. The test agent can then perform monitoring tasks and report results to the cloud backend. Test agents can be associated with endpoint devices ("endpoints") and non-endpoint devices (as defined below), and can be further based on hardware, software, or a combination of hardware and software. One example of a hardware based test agent can be a user experience insight (UXI) sensor that can be deployed as a hardware sensor that tracks the health and performance of a network. Based on configurations received from a cloud or other server, the UXI sensor can track a wireless (or wired) network by continuously monitoring the network and reporting measurements, results, and any problems it can encounter or identify in the network to a UXI backend. An example of a software based test agent can be an equivalent software monitoring agent that runs within a virtualized container. The software agent can execute or be deployed on a compatible host, such as an access point (AP), a hardware switch, or a client software application running on a laptop computer.
[0013] However, for both endpoint devices and non-endpoint devices, there are some limitations in using current hardware and software based solutions. One limitation can be that a single device in a network can lack visibility into other devices or clients in the network and can perform its monitoring and operations as a leaf node device, essentially isolated from other devices on the network. Any issues that the single device can discover can be unique to that single device (e.g., firewall policies) or can involve network wide issues (e.g., issues related to software defined wide area network (SD-WAN) / gateway).
[0014] Another limitation can be that each device can perform domain specific tasks. For example, a hardware based endpoint can monitor the quality of a wireless network, while a software based endpoint can perform another type of task related to a host running software (e.g., a switch or access point (AP)). In existing architectures, peer-driven actions and collaboration on network monitoring can require central (e.g., cloud based) orchestration rather than direct communication between peer devices. If an event or condition of interest related to a particular device occurs, such as a fault condition, the inability to establish a connection with a central entity (e.g., the cloud) can result in the inability to resolve the event or condition of interest. In some cases, the event or condition can occur and end before the cloud connection is established, which can result in the correct orchestration never being achieved. In other cases, the inability to establish the cloud connection can result in the event or condition never being detected or resolved, which can also result in the correct orchestration never being achieved.
[0015] Another limitation can involve using different monitoring solutions for different devices. To handle different devices on a network, an end user must individually configure different monitoring solutions. Tracking different information (e.g., metrics, results, and statuses) associated with these varying monitoring solutions can be burdensome and result in a cumbersome and inefficient overall monitoring solution for the network. Furthermore, it can be difficult to piece together information received from different types of devices (and corresponding different monitoring solutions) in order to present data to a user in a useful or cohesive manner. This different information can also result in a user facing greater challenges in performing root cause analysis and fault detection.
[0016] Aspects of instant application address these limitations by providing a system that leverages peer-to-peer integration between different types of hardware and software monitoring devices in a network to provide a comprehensive intelligent monitoring system across the network. The described aspects facilitate devices that are able to communicate and collaborate to provide a collective overview or representation of the network, i.e., by forming a collective ecosystem of elements including endpoint devices and non-endpoint devices. The described system aspects can provide a comprehensive intelligent monitoring system on, for example, a customer network, and can further provide user experience insights at different points of the network infrastructure. By using the collaborative features of the described aspects, the system can provide network metrics and other related information at, from, or through any device in the network, including endpoint devices and non-endpoint devices (as described below). Moreover, the system can perform intelligent monitoring using a set of conditions and corresponding actions. These conditions / actions can be user-defined or system-defined. Thus, instead of individual devices operating in isolation (as described below with respect to Figure 1 the described aspects can provide an intelligent and comprehensive monitoring system using peer-to-peer agents and condition / action configurations.
[0017] The terms "endpoint" and "endpoint device" are used interchangeably in this disclosure and refer to a physical or virtual device that can connect to and exchange information with other devices via or over a network. An endpoint device can be an "edge node" or "leaf node" in a network topology. Examples of endpoint devices can include, but are not limited to, desktop computers, laptop computers, mobile devices, tablets, printers, access points, embedded devices, servers, virtual machines, thin clients, sensors, actuators, point-of-sale terminals, and smart meters.
[0018] The term "non-endpoint device" is used in this disclosure to refer to a physical or virtual device that can connect to and exchange information with other devices via or over a network. A non-endpoint device can be an "interior node" or "branch node" in a network topology. Examples of non-endpoint devices can include, but are not limited to, switches, access switches, gateways, routers, and one or more endpoint devices described herein.
[0019] The terms "peer" and "peer device" are used interchangeably in this disclosure and refer to devices between which there can be a direct communication infrastructure or other communication link.
[0020] The term“test agent” is used in this disclosure to refer to a component, module, unit, or program that can be associated with an endpoint device and a non-endpoint device, and can be further based on hardware, software, or a combination of hardware and software. The test agent can perform various functions or operations, including but not limited to testing, monitoring, and communication of information, such as metrics, results, and status of network-related and device-related information.
[0021] Comparison of traditional monitoring models using different endpoints and using a peer agent
[0022] Figure 1 An environment 100 using a conventional monitoring model with different endpoints according to the prior art is shown. The environment 100 can include: a device 104 associated with a user 106 and a display 108; and devices 110. The device 104 can be a cloud server, or can represent a cloud-based service or server set. The devices 110 can include various devices, including but not limited to: an endpoint device; a non-endpoint device; an intermediary device; an edge or leaf node; an internal or branch node; an access point; a switch; an access switch; a gateway; and a router. The device 104 and the devices 110 can communicate with each other via a network 120. For example, the devices 110 can include: an access point 112 in communication with the device 104 via a link 122; a laptop computer 114 in communication with the device 104 via a link 124; a switch 116 in communication with the device 104 via a link 126; and a monitoring sensor 118 in communication with the device 104 via a link 128.
[0023] In the environment 100, during operation, each of the devices 112-118 can only operate in isolation from the other devices 112-118, and can only report its own results via a particular system in order to be accessed via a cloud, backend, or other server and subsequently displayed on the display 108. If a problem occurs with a respective link to the device 104 (e.g., the link 126 of the switch 116), the device associated with the respective link (e.g., the switch 116) can be unable to perform necessary actions, such as: retrieving a configuration file from the device 104; recording network metrics monitored by the switch 116; and reporting any device-specific or general network problems observed by the switch 116 based on the location of the switch 116 in the overall network topology of the devices 110.
[0024] Further, these metrics, issues, or results can be reported by each individual isolated device and only accessible through a separate user dashboard or backend system. For example, display 108 can include information that must be accessed via four separate systems, including: access point 112 ("system l"): monitored network information 170; laptop 114 ("system 2"): monitored network information 172; switch 116 ("system 3"): monitored network information 174; and monitoring sensor 118 ("system 4"): monitored network information 176. The information (170-176) from the four separate systems can be displayed in a non-cohesive manner that can only provide information about each device, but not provide an overall view of the entire network.
[0025] Thus, environment 100 depicts how each of devices 112-118 can perform its own monitoring in isolation from the other devices, where user review of the monitored information can also be performed in isolation from different backend systems.
[0026] The described aspects provide integration of peer-to-peer agents that can provide collaborative monitoring of a network. Figure 2 An environment 200 for an intelligent and comprehensive monitoring system using peer-to-peer agents is shown in accordance with one aspect of the present application. Environment 200 can include: a device 204 associated with a user 206 and a display 208; and devices 210. Device 204 can be a cloud server, or can represent a cloud-based service or server set. Devices 210 can include various devices, including but not limited to: an endpoint device; a non-endpoint device; an intermediary device; an edge or leaf node; an internal or branch node; an access point; a switch; an access switch; a gateway; and a router. Device 204 and devices 210 can communicate with each other via a network 220. For example, devices 210 can include: an access point 112 in communication with device 204 via a link 222; a laptop 214 in communication with device 204 via a link 224; a switch 216 in communication with device 204 via a link 226; and a monitoring sensor 218 in communication with device 204 via a link 228. Display 208 can include information as described below with respect to Figure 3 the described information.
[0027] In environment 100, each of devices 112-118 can operate in a collaborative manner to provide an intelligent and comprehensive monitoring system in operation, in contrast to the isolated monitoring described above with respect to environment 100. Figure 1 Figure 2 Device 210 can perform several operations that facilitate this collaborative monitoring. Laptop computer 214 can be used to illustrate exemplary devices within device 210. During operation, laptop computer 214 can deploy a software-based network monitoring agent. Laptop computer 214 can discover multiple peer devices on the same local network without receiving cloud orchestration instructions. That is, based on the deployed software agent, laptop computer 214 can use protocols such as Multicast Domain Name System (mDNS), Zigbee, Bluetooth Mesh, and Broadcast Ethernet to discover or identify peer devices. Laptop computer 214 or any device 210 can also use other protocols to discover its peers. Furthermore, other radio frequency (RF) based communications (e.g., Zigbee) can be used for communication between devices on different networks. Exemplary results of the discovery process may include: laptop computer 214 discovering at least access point 212 and switch 216 as its peers, with communication occurring via links 232 and 234, respectively; access point 212 discovering at least laptop computer 214 and switch 216, with communication occurring via links 232 and 236, respectively; and switch 216 discovering at least access point 212, laptop computer 214, and monitoring sensor 218, with communication occurring via links 236, 234, and 238, respectively.
[0028] Because this discovery process may be initiated by a software-based agent deployed on the device, it allows devices (including endpoints) to be discovered and their potential peers identified without the need for centralized orchestration instructions, such as cloud orchestration instructions from device 204. Once peers are identified, the system can use the peer agent to provide collaborative monitoring, for example, when a link from one device (214) to the server (204) fails, as described below. Figure 3 As stated above.
[0029] Each corresponding device (e.g., laptop computer 214) can obtain a device-specific profile from server 204. The profile may include or indicate network metrics to be monitored by the corresponding device. The profile may also include one or more condition / action pairs, such as triggers or conditions and corresponding actions to be performed. Each condition / action may specify or indicate user-defined or system-configured elements, such as rules, thresholds, network metrics to be monitored, messages or notifications to be sent to one or more other devices, procedures to be initiated, etc.
[0030] Figure 3 One aspect according to this application is shown. Figure 2 Environment 200 includes communication between peer devices in the event of an uplink failure from a device. Figure 3If the laptop 214 determines successful direct communication with the device 204, the laptop 214 can obtain its configuration file from the device 204 (as described above). However, if the laptop 214 determines unsuccessful direct communication with the device 204 (e.g., the link 224 from the laptop 214 to the device 204 fails, as indicated by the bold "X" 302), the laptop 214 can select a peer from its previously identified list of peers. The laptop 214 can select a peer based on, for example, an ordering of the discovered plurality of peer devices; and a current network metric associated with one or more of the discovered plurality of peer devices. For example, the laptop 214 can maintain a list of its peers in a local cache. The list can be prioritized, ranked, or ordered according to a hop count or latency associated with each of its identified peers. The laptop 214 can use the list to determine an alternative communication path for receiving data from and sending data to the server (e.g., the device 204). That is, the laptop 214 can attempt to find a peer with a complete uplink or route from the list via which it can obtain the configuration file (or send its stored monitored network metrics and results, as described below).
[0031] The laptop 214 can communicate with the device 204 via the selected peer. For example, the laptop 214 can issue a request for the configuration file to its peer device switch 216 (via the communication 304), which can issue the request to the device 204 (via the communication 306). The device 204 can return the requested configuration file to the laptop 214 using the same communications 306 and 304 through the switch 216. Thus, the described aspects can provide for autonomous edge configuration for mutual network commissioning and monitoring.
[0032] Once the laptop 214 obtains its configuration file, the laptop 214 can monitor the network metrics indicated in the obtained configuration file. The laptop 214 can be configured to report results related to the monitored network metrics at periodic intervals, based on a predetermined time, or in response to detecting a user-defined or system-configured condition specified in the configuration file. Detecting, determining, or triggering a user-defined or system-configured condition can cause the laptop 214 to perform a particular corresponding action, such as obtaining additional network metrics. Exemplary condition / action pairs are described below.
[0033] If the laptop 214 determines successful direct communication with the device 204, the laptop 214 can send data associated with the monitored metrics and actions to the device 204 (similar to the direct communication described above for obtaining the configuration file from the device 204). However, if the laptop 214 determines unsuccessful direct communication with the device 204 (e.g., the link 224 from the laptop 214 to the device 204 fails, as indicated by the bold "X" 302), the laptop 214 can select a peer (e.g., the switch 216) from its previously identified list of peers. The laptop 214 can select the peer based on the factors described above with respect to Figure 2 the laptop 214 can communicate with the device 204 via the selected peer. In this example, the laptop 214 can send data associated with the monitored metrics and actions (if any) to the server via the selected peer (e.g., the switch 216). Note that although the same peer (i.e., the switch 216) is depicted as the selected device for both unsuccessful direct communication scenarios (i.e., obtaining the configuration file and sending the monitored network metrics), the selected peer can be any of the laptop 214's previously identified peers.
[0034] The transmitted data associated with the monitored metrics, and in some cases, actions, can include test result metrics such as latency, round trip time, path convergence, and routes taken. These metrics (i.e., "total test cases") can be confirmed or validated by the actual elements or devices involved in the data transmission. Because each element or device can validate the entire test case, the system can account for the network as a whole or in its entirety. For example, by providing packet fragmentation metrics on the appropriate traffic flows, low throughput results at the edge clients can have further depth. In addition, the system can validate QoS policies whereby traffic emanating from the edge device or test agent can be marked and throughput thresholds can be tested. By providing the representation at the source of the traffic and through the transmission infrastructure, the described aspects can not only result in a deeper understanding of various tests including route testing, bandwidth testing, latency, and round trip testing, but also a deeper understanding of the status or validation of the implementation of various policies.
[0035] The system can display (e.g., in a graphical user interface) the results of the tests and the validation of the policies (e.g., the results of the tests and the validation of the policies can be displayed in a graphical user interface as described above with respect to FIG. 2). For example, the system can display the results of the tests and the validation of the policies in a graphical user interface as described above with respect to FIG. 2. Figure 2 and Figure 3various information, including aggregated network monitoring data associated with devices (e.g., 214), and information related to transmitted data associated with monitored metrics and actions, if any. The system can display this information in conjunction with a network topology (e.g., as a geographic map with labeled physical locations of devices, groups of devices, networks, or groups of networks). For example, the displayed information can include network information 270, which can include network type 272, sensor status and information 274, and historical information 276; a summary of ongoing issues 278; and a visual representation of the network / topology 280, which can include communication type 282 and device / network information 284. The displayed information can also include a number associated with a device, group of devices, network, or group of networks. The displayed information can also include a link or connection between two devices in a network; and a link or connection between two networks within an entire network.
[0036] The displayed information can include interactive elements 286 that allow the user 206 to view one or more of the following: statistical information or status associated with a link or connection; a rate of received or transmitted data; a rate of dropped packets; whether an external service is available; whether an unexpected captive portal or proxy exists; whether a power outage is detected; and whether a response is received from a Dynamic Host Configuration Protocol (DHCP) server. The system can display this information on a user dashboard, as described below with respect to Figure 4
[0037] As described above, the configuration file can include condition / action pairs. That is, any metric or test case can correspond to a condition that, when detected, can trigger a proactive or immediate action based on a pre-determined threshold (e.g., as defined in the configuration file). One example of a condition / pair can be a condition in which a device detects a failed uplink to a server or cloud backend, which triggers a corresponding action of checking the uplink of one or more of the device’s peers. Another example can be a device discovering a problem with a peer (e.g., the condition is that the rate of dropped packets on a switch exceeds a pre-determined threshold), which triggers a corresponding action of running a CLI command on the switch to get relevant counters or other interface metrics. In another example, the device can detect a condition of a failed uplink on one or more of its peers, which triggers a corresponding action of performing a test on a certain port of the determined peer and storing the results on the device itself. Yet another example can be a device verifying reachability of a critical internal server from different hosts / paths in the network, where the verification results can define various actions to be taken by the device or other devices to perform actions to resolve any issues related to the verification results.
[0038] In these examples, the device can detect a condition corresponding to a monitored network metric associated with the device itself. For a corresponding trigger action, the device can directly perform a corresponding first action, or can notify a peer device to perform a corresponding second action. The device can also detect a condition corresponding to a monitored network metric associated with a peer device. For a corresponding trigger action, the device can directly perform a corresponding third action, or can notify the peer device or another peer device to perform a corresponding fourth action. For example, a device can detect a situation where a test result for a particular route falls below a predetermined level, which triggers a route analysis from a particular switch (i.e., a peer device). In other words, when a device detects that a condition of monitoring a network metric is satisfied, the device can notify its peer device (switch) to perform a particular operation related to the route analysis. In some aspects, the device can also notify a different peer device to perform a predetermined action corresponding to the condition.
[0039] Further, the described aspects can aggregate the monitored network metrics using the above-described techniques, and can further transmit the aggregated information to or apply the aggregated information to an external monitoring entity, e.g., in parallel, to allow the external monitoring assistant Figure 2 and Figure 3 to access the device 204 and apply analysis to the aggregated information.
[0040] The described aspects can further enrich any network problems discovered by the client test application via problem sharing and validation routines, which can be executed on parallel network infrastructure components, e.g., switches, access points, or other network nodes, which can be responsible for handling traffic emanating from the host client. For example, a first agent can be installed on a UXI sensor, and the UXI sensor can run on a user device and monitor a certain network. If the first agent discovers a problem with the particular network that the user device is monitoring and wishes to provide a complete report of the discovered problem immediately (i.e., in real-time or upon discovery), the first agent can need to request a second agent (e.g., a UXI network analysis engine (NAE)) running on a local switch to check the configuration or traffic pattern by running a set of command line interface (CLI) commands. Such communication can provide intelligent and comprehensive monitoring reports by aggregating information from different peer devices at the time of the problem occurrence and by presenting / displaying the aggregated information to a network administrator or other end user in a coherent manner. Example displays are described below in connection with Figure 4 .
[0041] Thus, as Figures 2-4 illustrated, the described aspects can perform intelligent and comprehensive network monitoring via a plurality of “monitoring systems,” where each device or group of devices can be considered a monitoring system, rather than by aboutFigure 1 The plurality of individual endpoints or devices shown perform isolated network monitoring individually.
[0042] Example dashboard results and user interactions
[0043] Figure 4 An exemplary user dashboard 400 with display results from an intelligent and comprehensive monitoring system is shown in accordance with an aspect of the present application. The user dashboard can be displayed on a display screen of a device associated with a user (or a server accessible or authenticated for use by the user), e.g., as described above with respect to the display 208 of the Figure 2 and Figure 3 The user dashboard 400 can include an information display portion 402 and a visual representation of a network topology 460. The information display portion 402 can include a network 404 portion that outlines information about respective network types, including: a number of sensors currently operating and a status of the respective network type; and a historical status of the respective network for a particular number of previous time intervals (e.g., in hours). The network 404 portion can include a row indicating the respective network type, with columns indicating: the network type 410; a column with a visual representation and a number of sensors currently operating (sensors now 430); and a column with a historical status (last 24H ongoing 440). The network type (410) can include: Wi-Fi 412; Ethernet 414; captive portal 416; Dynamic Host Configuration Protocol (DHCP) 418; Domain Name System (DNS) 420; and gateway 422.
[0044] Each row can include the above-described information. For example, the row 406 for Wi-Fi 412 can include: a number of sensors currently operating as a value 656 (element 432); a bar or other visual representation (element 434) indicating different statuses relative to a total number of sensors currently operating, e.g.: a filled-in pattern can indicate a relative number of offline sensors or sensors from which no signal is currently being detected; a diagonally-striped fill pattern can indicate a relative number of online sensors or sensors from which data is currently being received; and a non-filled or blank pattern can indicate a relative number of sensors that are currently being commissioned or undergoing diagnostics.
[0045] Row 406 of Wi-Fi 412 can also include a visual representation of the status of the entire respective network over a recent historical period (e.g., 24 hours) (element 442), where a bar graph can represent each hour, and the shading or color of the bar graph (not indicated) can represent the overall health of the respective network for that hour, as well as a number related to the overall status of the respective network as value 108 (element 444), which can indicate, for example, the number of sensor problems or errors over the past 24 hours, the number of resets or reboots related to the respective network, and the like.
[0046] Information display portion 402 can also include an ongoing section 450 that outlines the number of issues that can currently be occurring or that have not been resolved. Row 452 can include information related to certain issues and can list the number of sensors and networks involved in the issue. For example, as shown in row 452: the ongoing issue "low received bit rate" can involve 30 sensors and 4 networks; the ongoing issue "low received bit rate" can involve 30 sensors and 4 networks; the ongoing issue "external service unavailable" involves 14 sensors and 5 networks; the ongoing issue "unexpected captive portal or proxy" can involve 15 sensors and 2 networks; the ongoing issue "power outage detected" can involve 15 sensors; and the ongoing issue "DHCP server not responding" can involve 7 sensors and 5 networks. Other issues can also be indicated (row 454).
[0047] A visual representation of the network topology 460 can be displayed as a map, where the physical location of each network is marked by a label (e.g., labels 462 and 464). Different types of communication links between networks can be indicated with different types of arrows, such as a double-sided arrow (e.g., 470) for indicating a first type of communication, and a thick straight line without an arrow (e.g., 478 and 480) for indicating a second type of communication link. The number or value displayed in each label can correspond to various values, and each label can be colored or indicated in different visual ways. Examples of various types of values can include: the number of devices in the respective network at the marked location; the number of sub-networks within the respective network; the number of sensors associated with the respective network; and the number of current outages at the respective network. Various colors (not shown) or other visual indicators (e.g., stripes, patterns, shading, highlighting, etc.) can be used to indicate the type of value being displayed. In some aspects, a user can use an input device (e.g., by hovering over a label with a mouse, or entering a particular key stroke or pattern on a keyboard) or a finger gesture (e.g., by swiping, hovering, tapping, or holding a portion of a touch-sensitive display screen) to view or change the view of the displayed labels.
[0048] The user can also use interactive elements (not shown) on the user dashboard 400 to view specific or detailed information about a particular sensor, device, or communication link within a respective network, e.g., by clicking on tab 462 and further clicking on a hierarchical organizational view of a respective network and related information, which can be displayed in section 402, on another scrollable area of section 402, or as an overlay on section 402.
[0049] Accordingly, the user dashboard 400 can provide the user with a high-level site-to-site view of an entire network (e.g., a customer network spanning multiple geographic locations). The user dashboard 400 can also provide the user with an in-depth overview of the complete network and provide statistical information about various links (e.g., via the information display section 402 and as described above). The various capabilities of user control via the dashboard 400 can enhance the performance of the comprehensive intelligent monitoring system described herein.
[0050] Example method for facilitating an intelligent and comprehensive monitoring system
[0051] Figure 5A A flowchart 500 is shown that illustrates a method of mechanisms that facilitate an intelligent and comprehensive monitoring system using peer-to-peer agents in a network, in accordance with an aspect of the present application. During operation, the system deploys a network monitoring agent on a device in a network (operation 502). The system discovers, without receiving cloud orchestration instructions, a plurality of peer devices on the same local network via the device (operation 504). If the system determines successful direct communication with a server (decision 506), then the system obtains a network monitoring configuration file for the device from the server, where the configuration file indicates network metrics to monitor and conditions associated with the network metrics (operation 508). If the system determines unsuccessful direct communication with the server (decision 506), then the system obtains the network monitoring configuration file for the device from the server via a first peer device (operation 510). The system can determine or select the first peer device based on the exemplary criteria described above.
[0052] The system monitors the network metrics indicated in the configuration file (operation 512). If a condition corresponding to a monitored network metric is satisfied or triggered (decision 514), then the system performs a predetermined action (operation 516), and operation continues at Figure 5B If a condition corresponding to a monitored network metric is not satisfied or triggered (decision 514), then operation continues at Figure 5B If a condition corresponding to a monitored network metric is not satisfied or triggered (decision 514), then operation continues at
[0053] Figure 5BA flow diagram 520 is shown that illustrates a method that facilitates mechanisms of an intelligent and comprehensive monitoring system that uses peer agents in a network, in accordance with an aspect of the present application. During operation, if the system determines successful direct communication with a server (decision 522), then the system sends data associated with monitored metrics and actions to the server (operation 524). If the system determines unsuccessful direct communication with the server (decision 522), then the system sends data to the server via a second peer device (operation 526). The system can determine or select the first peer device based on the exemplary criteria described above. The system allows the server to display aggregated network monitoring data associated with the device and peer devices in the network in conjunction with a network topology (operation 528). The system displays aggregated network monitoring data associated with the device, peer devices in the network, and other devices in the network on a screen of a device associated with the server (operation 530). Operation returns.
[0054] Thus, by deploying network monitoring agents to devices in a network, and by allowing peer discovery so that peers are used as devices that communicate configuration information and measurement results or monitoring data with a central cloud server (or other central server or backend), and further by using intelligent monitoring (e.g., the monitored or triggered conditions / actions described above), the described aspects can facilitate a peer-coordinated, intelligent, and comprehensive monitoring of a network.
[0055] Example system and apparatus for facilitating an intelligent and comprehensive monitoring system
[0056] Figure 6 A computer system 600 and a device 640 are shown that facilitate mechanisms of an intelligent and comprehensive monitoring system that uses peer agents in a network, in accordance with an aspect of the present application. The computer system 600 includes a processor 602, a memory 604, and a storage device 606. The memory 604 can include volatile memory (e.g., RAM) used as a main storage facility for the computer system 600, and can be used to store one or more memory pools. In addition, the computer system 600 can be coupled to peripheral input / output (I / O) user devices 610 (e.g., a display device 612, a keyboard 614, and a pointing device 616). The computer system 600 can correspond to the device 104 of Figure 1 , and the peripheral I / O user devices 610 can correspond to the display 108 of Figure 1 . The computer system 600 can communicate with multiple devices (or devices) such as the device 640 and the device 660 via communication links 680 and 682, respectively. The devices 640 and 660 can correspond to the device 110 of Figure 1 . The storage device 606 can store an operating system 618, a content processing system 620, and data 628.
[0057] The content processing system 620 can include instructions that, when executed by the computer system 600, can cause the computer system 600 to perform the methods and / or processes described in the present disclosure. In particular, the content processing system 620 can include instructions for issuing data packets to and / or receiving data packets from other network nodes through a computer network (communication unit 622). The data packets can include configuration information, network metrics or statistics, data associated with a device or peer device, information associated with conditions / actions in a configuration file, and data related to the operations described herein.
[0058] The content processing system 620 can also include instructions for aggregating data received from one or more devices or apparatuses (e.g., 640 / 660) (data aggregation unit 624). The content processing system 620 can include instructions for displaying, managing, updating, and modifying data displayed or to be displayed on a screen using the I / O device 610 (display management unit 626). The content processing system 620 can also include instructions for processing communications with external entities, such as entities that can retrieve aggregated network monitoring data or metrics stored by the computer system 600 and perform further analysis on the retrieved data (communication unit 622).
[0059] The data 628 can include any data required as input or generated as output by the methods and / or processes described in the present disclosure. In particular, the data 634 can store at least: data; network metrics; device related information; and information related to displaying network metrics or network topology information.
[0060] The apparatus 640 (as an example of a plurality of devices / apparatuses that can also include, for example, 660) can include a plurality of units or components that can communicate with one another, via wired, wireless, quantum optical, or electrical communication channels. The apparatus 640 can be implemented using one or more integrated circuits, and can include more units / components than those shown Figure 6fewer or more elements or devices. In addition, device 600 can be integrated in a computer system, or implemented as one or more separate devices capable of communicating with other computer systems and / or devices. Specifically, device 600 can include elements 642-656 for performing the functions or operations as described herein, including: a communication element 642 for communicating with one or more other devices or computer systems; a proxy deployment element 644 for deploying a network monitoring proxy on a device in a network; a peer discovery element 646 for discovering, by a device, a plurality of peer devices on a same local network without receiving cloud orchestration instructions; a profile management element 648 for obtaining, from a server or from the server via a first peer device, a network monitoring profile for the device based on a determination made by the communication element 642 as to whether there is a directed communication with the server, (where the profile indicates network metrics to monitor and conditions associated with the network metrics); a network metric monitoring element 650 for monitoring the network metrics indicated in the profile; a condition determination element 652 for determining whether a condition corresponding to a monitored network metric is satisfied; an action management element 654 for performing a predetermined action; the communication element 642 is also for transmitting data associated with the monitored metrics and actions to the server or to the server via a second peer device based on a determination made by the communication element 642 as to whether there is a directed communication with the server; and a display management element 656 for managing display of the data by the server in conjunction with a network topology to display aggregated network monitoring data associated with the device and peer devices in the network.
[0061] In general, the disclosed aspects provide a method, non-transitory computer readable storage medium, system, and apparatus for facilitating peer collaboration monitoring of a network. In one aspect, the system deploys a network monitoring agent on a device in a network. The system discovers, by the device, a plurality of peer devices on a same local network without receiving cloud orchestration instructions. In response to successful direct communication with a server, the system obtains, from the server, a network monitoring configuration file for the device, where the configuration file indicates network metrics to monitor and conditions associated with network metrics. In response to unsuccessful direct communication with the server, the system obtains the configuration file from the server via a first peer device. The system monitors the network metrics indicated in the configuration file. In response to the conditions corresponding to the monitored network metrics being satisfied, the system performs a predetermined action. In response to successful direct communication with the server, the system sends, to the server, data associated with the monitored metrics and the action. In response to unsuccessful direct communication with the server, the system sends, to the server, the data via a second peer device, thereby allowing the server to display, in combination with a network topology, aggregated network monitoring data associated with the device and the peer devices in the network.
[0062] In a variant of this aspect, when the conditions corresponding to the monitored network metrics are satisfied in association with the device, the action includes at least one of: the device directly performing a first action; and the device notifying a peer device to perform a second action.
[0063] In another variant of this aspect, when the conditions corresponding to the monitored network metrics are satisfied in association with a peer device, the action includes at least one of: the device directly performing a third action; and the device notifying the peer device or another peer device to perform a fourth action.
[0064] In another variant, the conditions corresponding to the monitored network metrics being satisfied is based on at least one of: a system configuration condition for a system configured action or a user defined action; and a user defined condition for a system configured action or a user defined action.
[0065] In another variant, the system deploys the network monitoring agent on a plurality of devices in the same local network, where the plurality of devices includes the peer devices. The system monitors, by respective devices, network metrics based on configuration files for the respective devices. The system sends, by respective devices, the monitored network metrics to the server or to the server via one or more peer devices of the respective devices.
[0066] In another variant, the device, the peer devices in the same local network, and the other devices comprise at least one of: an endpoint device; an intermediary device; an edge or leaf node; an internal or branch node; an access point; a switch; an access switch; a gateway; and a router.
[0067] In another variant, the system discovers the peer devices of the device based on a protocol, the protocol comprising at least one of: a multicast domain name system (mDNS) protocol; a Zigbee protocol; a Bluetooth mesh protocol; and a broadcast Ethernet protocol.
[0068] In another variant, the system displays, on a screen of a device associated with the server, the aggregated network monitoring data associated with the device, the peer devices in the network, and the other devices in the network. The aggregated network monitoring data is displayed in conjunction with the network topology. The displayed aggregated network monitoring data indicates at least one of: a visual representation of the network topology; a physical location of a device or group of devices in the network; a number associated with the device or the group of devices; and a link or connection between two devices in the network.
[0069] In another variant, the display further comprises interactive user elements that allow a user of the device associated with the server to view one or more of: statistical information or status associated with the link or connection; a rate of receiving or transmitting data; a rate of dropped packets; whether an external service is available; whether there is an unexpected captive portal or proxy; whether a power outage is detected; and whether a response is received from a dynamic host configuration protocol (DHCP) server.
[0070] In another variant, the system sends, via the server, data associated with the monitored metrics and the actions to an external monitoring entity.
[0071] In another variant, the system stores information associated with the discovered peer devices in a local cache of the device.
[0072] In another variant, the system determines the first peer device and the second peer device based on at least one of: an ordered sequence of the discovered peer devices; and a current network metric associated with one or more of the discovered peer devices.
[0073] In another aspect, a non-transitory computer-readable storage medium stores instructions that, when executed by a computer, cause the computer to perform the above-described method, including with respect to Figure 2 、 Figure 3 、 Figure 4 ,Figure 5A and Figure 5B .
[0074] In yet another aspect, an apparatus comprises: a proxy deployment unit to deploy a network monitoring proxy on a device in a network; a peer discovery unit to discover, by the device, a plurality of peer devices on a same local network without receiving cloud orchestration instructions; a communication unit to determine successful or unsuccessful direct communication with a server; a configuration file management unit to: obtain, from the server, a network monitoring configuration file for the device in response to the communication unit determining successful direct communication with the server, wherein the configuration file indicates a network metric to monitor and a condition associated with the network metric; and obtain the configuration file from the server via a first peer device in response to the communication unit determining unsuccessful direct communication with the server; a network metric monitoring unit to monitor the network metric indicated in the configuration file; a condition determination unit to determine that the monitored network metric is satisfied; an action management unit to perform a predetermined action in response to the condition determination unit determining that the condition corresponding to the monitored network metric is satisfied; the communication unit is further to: transmit, to the server, data associated with the monitored metric and the action in response to determining successful direct communication with the server; and transmit the data to the server via a second peer device in response to determining unsuccessful direct communication with the server; and a display management unit to allow the server to display, in conjunction with a network topology, aggregated network monitoring data associated with the device and the peer devices in the network.
[0075] The above description is provided for the purpose of enabling any person skilled in the art to make and use the aspects and examples, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects and applications without departing from the spirit and scope of the disclosure. Thus, the aspects described herein are not intended to be limited to the aspects shown, but are to be accorded the widest scope consistent with the principles and features disclosed herein.
[0076] Moreover, the above descriptions of aspects are provided as examples only and are not intended to limit or restrict the aspects to the forms disclosed. Many modifications, additions, and permutations of the described aspects will become apparent to those skilled in the art once given the above descriptions. Accordingly, the aspects described herein are intended to embrace all such alternatives, modifications, and variations as fall within the scope of the appended claims.
Claims
1. A method for facilitating peer collaboration monitoring of a network, the method comprising: deploying a network monitoring agent on a device in a network; discovering, by the device in the network, a plurality of peer devices on a same local network without receiving cloud orchestration instructions; in response to successful direct communication with a server, obtaining a network monitoring profile for the device in the network from the server, wherein the profile indicates a network metric to monitor and a condition associated with the network metric; in response to unsuccessful direct communication with the server, obtaining the profile from the server via a first peer device; monitoring the network metric indicated in the profile; in response to the condition corresponding to the monitored network metric being satisfied, performing a predetermined action; in response to successful direct communication with the server, sending data associated with the monitored metric and the action to the server; and in response to unsuccessful direct communication with the server, sending the data to the server via a second peer device, the sent data allowing the server to display, in combination with a network topology, aggregated network monitoring data associated with the device in the network and the plurality of peer devices.
2. The method of claim 1, wherein when the condition corresponding to the monitored network metric is satisfied in association with the device in the network, the action comprises at least one of: the device in the network directly performing a first action; and the device in the network notifying a peer device to perform a second action.
3. The method of claim 1, wherein when the condition corresponding to the monitored network metric is satisfied in association with a peer device, the action comprises at least one of: the device in the network directly performing a third action; and the device in the network notifying the peer device or another peer device to perform a fourth action.
4. The method of claim 1, wherein the condition corresponding to the monitored network metric being satisfied is based on at least one of: a system configuration condition for a system configured action or a user defined action; and a user defined condition for a system configured action or a user defined action.
5. The method of claim 1, further comprising: deploying the network monitoring agent on a plurality of devices in the same local network, wherein the plurality of devices includes the plurality of peer devices; monitoring, by respective devices, network metrics based on a profile for the respective device; and sending, by the respective devices, the monitored network metrics to the server or to the server via one or more peer devices of the respective devices.
6. The method of claim 1, wherein the device in the network, the plurality of peer devices in the same local network, and other devices include at least one of: an endpoint device; an intermediary device; an edge or leaf node; an internal or branch node; an access point; a switch; a gateway; and a router, and wherein the switch includes an access switch. 7. The method of claim 1, further comprising discovering the plurality of peer devices in the network based on a protocol, the protocol comprising at least one of: a multicast domain name system (mDNS) protocol; a Zigbee protocol; a Bluetooth mesh protocol; and a broadcast Ethernet protocol.
8. The method of claim 1, further comprising: displaying, on a screen of a device associated with the server, the aggregated network monitoring data associated with devices in the network, the plurality of peer devices in the network, and other devices in the network, wherein the aggregated network monitoring data is displayed in conjunction with the network topology, and wherein the displayed aggregated network monitoring data indicates at least one of: a visual representation of the network topology; a physical location of a device or group of devices in the network; a number associated with a device or group of devices; and a link or connection between two devices in the network.
9. The method of claim 8, wherein the screen further displays interactive user elements that allow a user of the device associated with the server to view one or more of: statistical information or status associated with the link or connection; a rate of receiving or transmitting data; a rate of dropped packets; whether an external service is available; whether an unexpected captive portal or proxy exists; whether a power outage is detected; and whether a response is received from a dynamic host configuration protocol (DHCP) server.
10. The method of claim 1, further comprising: sending, via the server, the data associated with the monitored metrics and the actions to an external monitoring entity.
11. The method of claim 1, further comprising: storing information associated with the discovered plurality of peer devices in a local cache of a device in the network.
12. The method of claim 1, further comprising determining the first peer device and the second peer device based on at least one of: an ordered sequence of the discovered plurality of peer devices; and a current network metric associated with one or more of the discovered plurality of peer devices.
13. A non-transitory computer-readable storage medium comprising instructions executable by a computer to: deploy a network monitoring agent on a device in a network; discover, by the device in the network, a plurality of peer devices on a same local network without receiving cloud orchestration instructions; in response to successful direct communication with a server, obtain, from the server, a network monitoring configuration file for the device in the network, wherein the configuration file indicates network metrics to monitor and conditions associated with network metrics; in response to unsuccessful direct communication with the server, obtain, via a first peer device, the configuration file from the server; monitor the network metrics indicated in the configuration file; in response to the conditions corresponding to the monitored network metrics being satisfied, perform a predetermined action; in response to successful direct communication with the server, sending data associated with the monitored metric and the action to the server; and in response to unsuccessful direct communication with the server, sending the data to the server via a second peer device, the sent data allowing the server to display, in combination with a network topology, aggregated network monitoring data associated with devices in the network and the plurality of peer devices.
14. The non-transitory computer-readable storage medium of claim 13, wherein when the condition corresponding to the monitored network metric is satisfied in association with a device in the network, the action comprises at least one of: the device in the network directly performing a first action; and the device in the network notifying a peer device to perform a second action; and wherein when the condition corresponding to the monitored network metric is satisfied in association with a peer device, the action comprises at least one of: the device in the network directly performing a third action; and the device in the network notifying the peer device or another peer device to perform a fourth action.
15. The non-transitory computer-readable storage medium of claim 14, wherein the condition corresponding to the monitored network metric being satisfied is based on at least one of: a system configuration condition for a system configured action or a user defined action; and a user defined condition for a system configured action or a user defined action.
16. The non-transitory computer-readable storage medium of claim 14, wherein the instructions further comprise instructions to: deploy the network monitoring agent on a plurality of devices in the same local network, wherein the plurality of devices comprises the plurality of peer devices; monitor, by a respective device, a network metric based on a configuration file for the respective device; and send, by the respective device, the monitored network metric to the server or to the server via one or more peer devices of the respective device.
17. The non-transitory computer-readable storage medium of claim 14, wherein the instructions further comprise instructions to: display, on a screen of a device associated with the server, the aggregated network monitoring data associated with devices in the network, the plurality of peer devices in the network, and other devices in the network, wherein the aggregated network monitoring data is displayed in combination with the network topology, wherein the displayed aggregated network monitoring data indicates at least one of: a visual representation of the network topology; a physical location of a device or group of devices in the network; a number associated with a device or group of devices; and a link or connection between two devices in the network.
18. The non-transitory computer-readable storage medium of claim 17, wherein an interactive user element is further displayed on the screen, the interactive user element allowing a user of the device associated with the server to view one or more of: statistical information or status associated with the link or connection; a rate of receiving or transmitting data; a rate at which packets are dropped; whether external services are available; whether there is an unexpected captive portal or proxy; whether a power outage is detected; and whether a response is received from a Dynamic Host Configuration Protocol (DHCP) server.
Citation Information
Patent Citations
Cyber security system with adaptive machine learning features
CN109246072A
Systems and methods for broadband communication link performance monitoring
CN111670562A