A network device exception monitoring method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA CONSTRUCTION BANK
- Filing Date
- 2024-12-31
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]目前,网络监控往往基于单设备单线路的监控来定位故障,而对于多设备多线路的监控能力不足,不能从整体上评估网络领域的运行情况,导致异常检测效率低下,且由于不能快速发现网络故障,导致业务中断情况频发,用户体验感较差
[0036]基于本说明书提供的一种网络设备异常监控方法,接收针对目标网络中的目标设备的异常检测请求;根据所述目标设备在所述目标网络中的拓扑连接关系,确定所述目标设备对应的第一设备组;其中,所述第一设备组中包含多个与所述目标设备存在设备稳定性关联关系的第一设备;根据所述第一设备组中的多个第一设备的端口的性能数据,确定所述第一设备组的运行状态数据;获取所述目标设备在所述目标网络中所在的第一线路,并获取所述第一线路对应的第一线路组;其中,所述第一线路组包含多条与所述第一线路存在线路稳定性关联关系的第二线路;根据所述第一线路组中的多条第二线路的线路运行情况,确定所述第一线路组的运行状态数据;根据所述第一设备组的运行状态数据和/或所述第一线路组的运行状态数据,确定针对所述目标设备的异常检测结果。这样,一方面,在目标网络中存在多个设备以及多条线路的情况下,可以根据与目标设备存在设备稳定性关联关系的第一设备构建的第一设备组,以及与目标设备所在的第一线路存在线路稳定性关联关系的第二线路构建的第一线路组的运行状态数据,对目标设备是否存在异常进行检测,可以提高设备异常检测效率。另一方面,可以从业务角度出发,以组(即设备组和线路组)为单位,根据目标设备对应的第一设备组和第一线路组的运行状态数据,对目标设备进行异常监控,可以在保证高效的异常检测效率的同时,避免由于设备或线路存在异常而导致的业务中断情况的发生,保证了业务连续性。
Smart Images

Figure CN119788547B_ABST
Abstract
Description
Technical Field
[0001] This manual belongs to the field of network device monitoring technology, and in particular relates to a method and device for monitoring network device anomalies. Background Technology
[0002] Currently, network monitoring is often based on monitoring single devices and single lines to locate faults. However, it lacks the ability to monitor multiple devices and multiple lines, and cannot comprehensively assess the operation of the network. This results in low efficiency in anomaly detection, and the inability to quickly detect network faults leads to frequent business interruptions and a poor user experience.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This specification provides a method and apparatus for monitoring network device anomalies. From a business perspective, it monitors target devices in groups based on the operating status data of the first device group and the first line group corresponding to the target device. This method can ensure high efficiency in anomaly detection while avoiding service interruptions caused by device or line anomalies, thus ensuring service continuity.
[0005] This manual provides a method for monitoring network device anomalies, including:
[0006] Receive anomaly detection requests for target devices in the target network;
[0007] Based on the topological connection relationship of the target device in the target network, a first device group corresponding to the target device is determined; wherein, the first device group contains multiple first devices that have a device stability association relationship with the target device;
[0008] Based on the performance data of the ports of multiple first devices in the first device group, the operating status data of the first device group is determined;
[0009] Obtain the first line where the target device is located in the target network, and obtain the first line group corresponding to the first line; wherein, the first line group includes multiple second lines that have a line stability correlation with the first line;
[0010] Based on the operational status of multiple second lines in the first line group, determine the operational status data of the first line group;
[0011] Based on the operating status data of the first equipment group and / or the operating status data of the first line group, determine the anomaly detection result for the target equipment.
[0012] In one embodiment, determining the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group includes:
[0013] Based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device, the load balancing status of the first device group is determined.
[0014] Based on the load balancing status, determine the operating status data of the first device group.
[0015] In one embodiment, determining the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group includes:
[0016] Obtain abnormal detection results for multiple first devices in the first device group within a preset detection period;
[0017] Based on the performance data of the ports of multiple first devices in the first device group and the anomaly detection results, the operating status data of the first device group is determined.
[0018] In one embodiment, determining the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group includes:
[0019] The performance data of the ports of multiple first devices in the first device group are aggregated, and the operating status data of the first device group is determined based on the aggregation results.
[0020] In one embodiment, the method further includes:
[0021] Based on the anomaly detection results, determine whether the target device has an anomaly;
[0022] If it is determined that the target device is abnormal, an abnormality handling strategy for the target device is determined based on the operating status data of the first device group and / or the operating status data of the first line group.
[0023] In one embodiment, determining the anomaly handling strategy for the target device based on the operating status data of the first device group includes:
[0024] If, based on the operating status data of the first device group, it is determined that the first device group has an unbalanced load, the anomaly handling strategy for load balancing the first device group is determined based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device.
[0025] In one embodiment, determining the anomaly handling strategy for the target device based on the operating status data of the first line group includes:
[0026] If, based on the operating status data of the first line group, it is determined that the number of second lines with signal interruption in the first line group is greater than a preset threshold, the abnormal handling strategy for lines that have a line stability correlation with the first line group is determined.
[0027] This manual provides a network device anomaly monitoring device, including:
[0028] The anomaly detection request module is used to receive anomaly detection requests for target devices in the target network;
[0029] The first group of determining modules is used to determine the first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein, the first device group includes multiple first devices that have a device stability association relationship with the target device;
[0030] The first state determination module is used to determine the operating state data of the first device group based on the performance data of the ports of multiple first devices in the first device group.
[0031] The second group of determining modules is used to obtain the first line where the target device is located in the target network, and to obtain the first line group corresponding to the first line; wherein, the first line group includes multiple second lines that have a line stability correlation with the first line;
[0032] The second state determination module is used to determine the operating state data of the first line group based on the operating status of multiple second lines in the first line group.
[0033] The anomaly detection module is used to determine the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
[0034] This specification also provides an electronic device, including a processor and a memory for storing processor-executable instructions, wherein the processor, when executing the instructions, implements a method for monitoring network device anomalies.
[0035] This specification also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed, implement a method for monitoring network device anomalies.
[0036] Based on the network device anomaly monitoring method provided in this specification, the method involves receiving an anomaly detection request for a target device in a target network; determining a first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein the first device group includes multiple first devices that have a device stability association with the target device; determining the operating status data of the first device group based on the port performance data of the multiple first devices in the first device group; obtaining the first line where the target device is located in the target network, and obtaining the first line group corresponding to the first line; wherein the first line group includes multiple second lines that have a line stability association with the first line; determining the operating status data of the first line group based on the line operation status of the multiple second lines in the first line group; and determining the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group. Thus, on the one hand, when there are multiple devices and multiple lines in the target network, the method can detect whether the target device has anomalies based on the operating status data of the first device group constructed from the first devices that have a device stability association with the target device, and the first line group constructed from the second lines that have a line stability association with the first line where the target device is located, thereby improving the efficiency of device anomaly detection. On the other hand, from a business perspective, the target device can be monitored for anomalies by group (i.e., equipment group and line group) based on the operating status data of the first equipment group and the first line group corresponding to the target device. This can ensure high efficiency in anomaly detection while avoiding business interruptions caused by equipment or line anomalies, thus ensuring business continuity. Attached Figure Description
[0037] To more clearly illustrate the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a flowchart illustrating a method for monitoring network device anomalies, provided in one embodiment of this specification.
[0039] Figure 2 This is a schematic diagram of the electronic device structure provided in one embodiment of this specification;
[0040] Figure 3 This is a schematic diagram of the structural composition of a network device anomaly monitoring device provided in one embodiment of this specification;
[0041] Figure 4 This is a flowchart illustrating an embodiment of a network device first device group anomaly monitoring method provided in this specification;
[0042] Figure 5 This is a flowchart illustrating another method for monitoring anomalies in the first line group of a network device, provided in one embodiment of this specification. Detailed Implementation
[0043] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0044] See Figure 1 As shown in the embodiments of this specification, a method for monitoring network device anomalies is provided, wherein the method is specifically applied to the server side. In specific implementation, the method may include the following:
[0045] S101: Receive an anomaly detection request for a target device in the target network;
[0046] S102: Based on the topological connection relationship of the target device in the target network, determine the first device group corresponding to the target device; wherein, the first device group includes multiple first devices that have a device stability association relationship with the target device;
[0047] S103: Determine the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group;
[0048] S104: Obtain the first line where the target device is located in the target network, and obtain the first line group corresponding to the first line; wherein, the first line group includes multiple second lines that have a line stability correlation with the first line;
[0049] S105: Determine the operating status data of the first line group based on the operating status of multiple second lines in the first line group;
[0050] S106: Determine the anomaly detection result for the target device based on the operating status data of the first equipment group and / or the operating status data of the first line group.
[0051] The device stability association can refer to the collaborative working relationship (i.e., high availability relationship) established between multiple devices in a network system through mechanisms such as primary / backup, load balancing, or redundancy design. For example, taking the device stability association as an association built through the primary / backup mechanism, assuming the target device is the primary device, the first device with the device stability association relationship with the target device can be the backup device of the target device. That is, when the target device is unavailable, the first device can be switched to the primary device and take over the business being processed by the target device, so as to ensure the continuity of business processing through device stability association.
[0052] The aforementioned first device group can be a high-availability device group constructed from first devices that have a high-availability relationship with the target device (i.e., device stability association). In other words, the first device group can be a logical set composed of multiple first devices with redundancy, master-slave failover, or load balancing relationships with the target device. Taking the first device group as an example of devices with a master-slave relationship with the target device, the first devices can maintain coordination with the target device through mechanisms such as heartbeat detection and state synchronization to ensure that the first devices can quickly take over services and guarantee network high availability in the event of a target device failure. Specifically, the target device can be a core switch, and the first devices can be backup switches. Thus, multiple backup switches can form a high-availability device group (i.e., the first device group) corresponding to the core switch. The first devices in the high-availability device group can ensure the stable operation of the data center.
[0053] The aforementioned first line group can refer to a high-availability line group composed of multiple redundant lines (i.e., second lines) corresponding to the first line. The second lines can dynamically collaborate with the first line through performance indicators such as latency and bandwidth utilization, ensuring that the second lines can automatically take over data transmission tasks when the first line is interrupted or its performance fails to meet requirements. High-availability line groups can be used to guarantee the reliability of network connections for services. For example, a dedicated transmission line group can be constructed based on multiple physical links (i.e., second lines) corresponding to the first line to support high-speed data transmission across data centers.
[0054] The performance data of the ports of the first device mentioned above may include key network device performance data such as traffic, error messages, packet loss, number of ports, temperature, and fan status.
[0055] The operating status data of the first equipment group can be used to characterize the operating status of the first equipment group (such as normal, fault, standby, etc.).
[0056] The operation status of the second line can include the connection status (e.g., unobstructed, disconnected, etc.) and connection performance status, while the operation status data of the first line group can be used to characterize the connection status of the first line group (e.g., unobstructed, disconnected, jittering, etc.).
[0057] In some embodiments, determining the first device group corresponding to the target device based on the topological connection relationship of the target device in the target network may specifically include:
[0058] Based on the topological connection relationship of the target device in the target network, obtain multiple devices in the target network that have a cooperative relationship with the target device, and based on preset device configuration information, obtain multiple devices in the target network that have a primary / backup relationship with the target device.
[0059] The acquired multiple devices are identified as first devices, and a first device group corresponding to the target device is constructed based on the multiple first devices.
[0060] Specifically, by identifying the target device's connection location within the network topology, multiple devices with collaborative relationships (such as service collaboration or load balancing collaboration) can be determined. Furthermore, by combining device configuration information such as network configuration files or management platform records, multiple devices with primary / backup relationships with the target device can be identified. These identified devices can be designated as primary devices, meaning they are devices associated with the target device's stability. These primary devices can then work with the target device to share the high availability requirements of network services.
[0061] In some embodiments, obtaining the first line group corresponding to the first line may specifically include:
[0062] S1: Obtain multiple lines with the same data processing nodes as the first line, and according to preset line configuration information, obtain multiple lines in the target network that have a primary / backup relationship and / or bandwidth sharing relationship with the first line;
[0063] S2: The acquired multiple lines are identified as second lines, and a first line group corresponding to the first line is constructed based on the multiple second lines.
[0064] Specifically, firstly, multiple lines with the same data processing nodes as the first line can be obtained through the physical link information between the first line and other lines; secondly, multiple lines with redundancy characteristics (such as primary / backup switching, multi-path transmission, or bandwidth sharing mechanism) as the first line can be obtained based on the preset line configuration information, and the obtained multiple lines can be identified as the second line.
[0065] Based on the above embodiments, on the one hand, when there are multiple devices and multiple lines in the target network, the operation status data of the first device group (constructed by the first device with a device stability correlation with the target device) and the first line group (constructed by the second line with a line stability correlation with the first line where the target device is located) can be used to detect whether the target device has any abnormalities, thereby improving the efficiency of device abnormality detection. On the other hand, from a business perspective, the target device can be monitored for abnormalities based on the operation status data of the first device group and the first line group corresponding to the target device, using groups (i.e., device groups and line groups) as units. This can ensure high efficiency in abnormality detection while avoiding service interruptions caused by device or line abnormalities, thus ensuring service continuity.
[0066] In some embodiments, the method for determining the operating status data of the first device group based on the performance data of the ports of the plurality of first devices in the first device group may further include the following:
[0067] S1: Determine the load balancing status of the first device group based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device;
[0068] S2: Determine the operating status data of the first device group based on the load balancing status.
[0069] The aforementioned load balancing status refers to the allocation and utilization of resources (such as port traffic, computing power, bandwidth, etc.) among the devices within the first device group. Load balancing status characterizes whether the workload of the multiple devices in the first device group remains even when carrying services, and whether some devices are overloaded or idle. Uneven load distribution may cause performance problems, such as overload of individual devices or waste of some resources, thereby affecting the overall efficiency and reliability of the system.
[0070] In some embodiments, determining the load balancing status of the first device group based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device may specifically include:
[0071] The system acquires the performance data of ports of multiple first devices in the device group, as well as the maximum and minimum values of the performance data of the target device's port, and determines the load balancing status of the first device group based on the difference between the maximum and minimum values.
[0072] Furthermore, since there may be a primary / backup relationship between the target device and multiple first devices, the load balancing status of the first device group can be determined based on the performance data of the ports of multiple first devices, the performance data of the ports of the target device, the preset load weight corresponding to each first device, and the load weight corresponding to the target device.
[0073] Based on the above embodiments, by monitoring the load balancing status, the target devices can be monitored through the operating status of the first device group, which can avoid resource waste and single-point overload problems, and enhance the stability and high availability of services. In addition, it helps to simplify fault location and management, improve operation and maintenance efficiency, and support dynamic expansion capabilities, ensuring that network resources can be flexibly adjusted when business demand increases, thereby guaranteeing the long-term stable operation and efficient operation of the system.
[0074] In some embodiments, the method for determining the operating status data of the first device group based on the performance data of the ports of the plurality of first devices in the first device group may further include the following:
[0075] S1: Obtain the abnormal detection results for multiple first devices in the first device group within a preset detection period;
[0076] S2: Determine the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group and the anomaly detection results.
[0077] In some embodiments, obtaining abnormal detection results for multiple first devices in the first device group within a preset detection period may specifically include:
[0078] Within a preset detection period, status data from multiple devices in the first device group are collected and analyzed to detect anomalies in the device group's performance metrics (such as port traffic, packet loss rate, latency, and CPU utilization). Anomaly detection can be achieved using various methods, such as threshold comparison, historical trend analysis, or machine learning model prediction, to accurately identify potentially abnormal devices or behaviors within the device group.
[0079] Based on the above embodiments, by obtaining the anomaly detection results of multiple first devices in the first device group, not only can the abnormal devices be accurately located, but potential problems can also be discovered through correlation analysis, thereby improving the accuracy and timeliness of fault warning.
[0080] In some embodiments, the method for determining the operating status data of the first device group based on the performance data of the ports of the plurality of first devices in the first device group may further include the following:
[0081] The performance data of the ports of multiple first devices in the first device group are aggregated, and the operating status data of the first device group is determined based on the aggregation results.
[0082] Specifically, by aggregating port performance data (such as traffic, packet loss rate, latency, and bandwidth utilization) from multiple devices in the first device group, the system can comprehensively analyze the overall operating status of the device group. Aggregation processing can include calculating statistical indicators such as average, maximum, minimum, and sample variance to assess the load balancing and resource utilization within the first device group. Furthermore, by comprehensively comparing historical trends and real-time dynamics of this performance data, potential anomalies in the first device group can be identified, and ports that may be faulty or overloaded can be identified. Based on the results of the aggregation processing, operating status data for the first device group can be generated, including load balancing status, health assessment results, and overall resource availability.
[0083] Based on the above embodiments, aggregating port performance data allows for a holistic understanding of the device group's operational status, avoiding the limitations of single-device data analysis and facilitating the rapid detection of performance anomalies and problem localization. This operational status data improves the resource utilization of the device group, ensures balanced load distribution, and reduces overload risk and failure rate, thereby further enhancing network stability and high availability. Furthermore, this aggregated analysis reduces the complexity of manual troubleshooting and improves operational efficiency.
[0084] In some embodiments, the method may further include the following:
[0085] S1: Determine whether the target device has an anomaly based on the anomaly detection results;
[0086] S2: If it is determined that the target device is abnormal, an abnormality handling strategy for the target device is determined based on the operating status data of the first device group and / or the operating status data of the first line group.
[0087] The aforementioned operational status data may include performance indicators (such as traffic, latency, packet loss rate, etc.) of other devices in the first device group and / or the availability and current load of backup lines in the first line group.
[0088] Specifically, based on operational status data, the optimal anomaly handling strategy can be dynamically selected. This could include switching traffic or tasks from the target device to backup devices within the first device group, rerouting data traffic to backup lines within the first line group, or triggering operational alarms to arrange manual intervention. Furthermore, by combining historical anomaly records and current operational status, the potential impact can be predicted and preventative measures can be taken to further reduce the disruption of anomalies to business operations.
[0089] Based on the above embodiments, by combining the operational status data of the first device group and the first line group to determine the anomaly handling strategy, optimal resource allocation and dynamic adjustment can be achieved, ensuring uninterrupted service and reducing the risk of anomaly propagation. This method not only improves the automation level and response speed of anomaly handling but also effectively reduces the complexity of manual intervention, further enhancing the high availability and stability of network device monitoring and providing strong support for continuous network operation.
[0090] In some embodiments, the method for determining anomaly handling strategies for the target device based on the operating status data of the first device group may further include the following:
[0091] If, based on the operating status data of the first device group, it is determined that the first device group has an unbalanced load, the anomaly handling strategy for load balancing the first device group is determined based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device.
[0092] Specifically, when a load imbalance is identified in the first device group, the specific high-load and low-load devices are identified by calculating indicators such as traffic, bandwidth utilization, latency, and packet loss rate for each port based on the port performance data of multiple first devices in the first device group and the target device. Simultaneously, by combining the resource allocation rules within the first device group and real-time service requirements, targeted load balancing anomaly handling strategies can be formulated. For example, traffic allocation weights can be dynamically adjusted to redirect some traffic from high-load devices to low-load devices, or backup devices can be activated to share the traffic load.
[0093] Furthermore, based on predictive models, the impact of load balancing adjustments on the overall performance of the equipment group can be assessed, and the optimal handling scheme can be adopted to reduce service interruptions or performance fluctuations. The execution of this strategy can be completed automatically through intelligent scheduling algorithms, or manual intervention can be triggered when necessary.
[0094] Based on the above embodiments, by analyzing the load balancing problem of the first device group and formulating precise anomaly handling strategies, the utilization efficiency of resources within the first device group can be significantly improved, avoiding performance degradation or failure of individual devices due to overload. Simultaneously, load balancing optimization can ensure the even distribution of service traffic, reduce network latency and data loss, and improve the overall network stability and high availability. Furthermore, automated load balancing reduces the intervention of maintenance personnel, improves the efficiency and intelligence of network management, and provides crucial support for efficient operation and maintenance in complex network environments.
[0095] In some embodiments, the method for determining anomaly handling strategies for the target device based on the operating status data of the first line group may further include the following:
[0096] If, based on the operating status data of the first line group, it is determined that the number of second lines with signal interruption in the first line group is greater than a preset threshold, the abnormal handling strategy for lines that have a line stability correlation with the first line group is determined.
[0097] Specifically, when the number of interrupted second lines in the first line group exceeds a preset threshold, the operational status data of the first line group is analyzed. This includes the load capacity of normal lines, the availability of backup lines, and historical fault information and causes of interruption (such as physical disconnection, signal interference, or equipment failure). Combining this data, the overall stability of the first line group can be assessed, and other lines with stability correlations to the interrupted lines can be identified, such as lines sharing critical nodes, belonging to the same physical link, or having insufficient redundancy design. In response, anomaly handling strategies can be developed, such as prioritizing the use of backup lines to replace the interrupted lines to carry service traffic, dynamically adjusting traffic allocation to prevent overload of related lines, or issuing high-level alarms to prompt maintenance personnel to conduct emergency investigations and repairs of the relevant line group. If necessary, multi-line redundancy switching technology can also be used to ensure uninterrupted service traffic, thereby mitigating the impact of line interruptions on network stability.
[0098] Based on the above embodiments, by conducting in-depth analysis of signal interruption situations and their associated lines in the first line group and formulating anomaly handling strategies, the fault tolerance and service recovery efficiency of the first line group can be significantly improved. The implementation of anomaly handling strategies can effectively reduce the impact of interrupted lines on associated lines and overall services, ensuring the continuity of critical services.
[0099] Furthermore, it can proactively detect potential line stability risks, reducing the likelihood of further escalation of faults and improving network reliability. In addition, anomaly handling reduces the complexity of manual intervention, improving the efficiency and accuracy of network operations and maintenance.
[0100] As can be seen from the above, the network device anomaly monitoring method provided in this specification receives an anomaly detection request for a target device in a target network; determines a first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein the first device group includes multiple first devices that have a device stability correlation with the target device; determines the operating status data of the first device group based on the port performance data of the multiple first devices in the first device group; obtains the first line where the target device is located in the target network, and obtains the first line group corresponding to the first line; wherein the first line group includes multiple second lines that have a line stability correlation with the first line; determines the operating status data of the first line group based on the line operation status of the multiple second lines in the first line group; and determines the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group. Thus, on the one hand, when there are multiple devices and multiple lines in the target network, the anomaly detection efficiency can be improved by detecting whether the target device has an anomaly based on the operating status data of the first device group constructed from the first devices that have a device stability correlation with the target device, and the first line group constructed from the second lines that have a line stability correlation with the first line where the target device is located. On the other hand, from a business perspective, the target device can be monitored for anomalies by group (i.e., equipment group and line group) based on the operating status data of the first equipment group and the first line group corresponding to the target device. This can ensure high efficiency in anomaly detection while avoiding business interruptions caused by equipment or line anomalies, thus ensuring business continuity.
[0101] See Figure 2 As shown in the embodiments of this specification, a specific electronic device is also provided, wherein the electronic device includes a network communication port 201, a processor 202 and a memory 203, and the above structures are connected by internal cables so that the various structures can perform specific data interaction.
[0102] Specifically, the network communication port 201 can be used to receive anomaly detection requests for target devices in the target network.
[0103] The processor 202 can be specifically configured to: determine a first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein the first device group includes multiple first devices that have a device stability correlation with the target device; determine the operating status data of the first device group based on the performance data of the ports of the multiple first devices in the first device group; obtain the first line where the target device is located in the target network, and obtain the first line group corresponding to the first line; wherein the first line group includes multiple second lines that have a line stability correlation with the first line; determine the operating status data of the first line group based on the line operation status of the multiple second lines in the first line group; and determine the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
[0104] The memory 203 can be used to store the corresponding instruction program.
[0105] Based on the above method, the relevant structural performance of electronic devices can be effectively utilized to improve the data processing speed of electronic devices and efficiently realize the method of network device anomaly monitoring.
[0106] In this embodiment, the network communication port 201 can be a virtual port bound to different communication protocols, thereby enabling the sending or receiving of different data. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. Furthermore, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM or CDMA; it can also be a Wi-Fi chip; or it can be a Bluetooth chip.
[0107] In this embodiment, the processor 202 can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification is not limiting.
[0108] In this embodiment, the memory 203 may include multiple layers. In a digital system, anything that can store binary data can be a memory. In an integrated circuit, a circuit with storage function but no physical form is also called a memory, such as RAM, FIFO, etc. In a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.
[0109] This specification also provides a computer-readable storage medium based on the above-described network device anomaly monitoring method. The computer-readable storage medium stores computer program instructions that, when executed, implement the following: receiving an anomaly detection request for a target device in a target network; determining a first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein the first device group includes multiple first devices that have a device stability correlation with the target device; determining the operating status data of the first device group based on the port performance data of the multiple first devices in the first device group; obtaining the first line where the target device is located in the target network, and obtaining a first line group corresponding to the first line; wherein the first line group includes multiple second lines that have a line stability correlation with the first line; determining the operating status data of the first line group based on the line operation status of the multiple second lines in the first line group; and determining an anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
[0110] In this embodiment, the storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions. The network communication unit can be an interface configured according to standards specified in the communication protocol for network connection communication.
[0111] In this embodiment, the specific functions and effects implemented by the program instructions stored in the computer-readable storage medium can be explained in comparison with other embodiments, and will not be repeated here.
[0112] See Figure 3 At the software level, embodiments of this specification also provide a network device anomaly monitoring device, which may specifically include the following structural modules:
[0113] Anomaly detection request module 301 is used to receive anomaly detection requests for target devices in the target network.
[0114] The first group determination module 302 is used to determine the first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein, the first device group includes multiple first devices that have a device stability association relationship with the target device;
[0115] The first state determination module 303 is used to determine the operating state data of the first device group based on the performance data of the ports of multiple first devices in the first device group.
[0116] The second group of determining modules 304 is used to obtain the first line where the target device is located in the target network, and to obtain the first line group corresponding to the first line; wherein, the first line group includes multiple second lines that have a line stability correlation with the first line;
[0117] The second state determination module 305 is used to determine the operating state data of the first line group based on the operating status of multiple second lines in the first line group.
[0118] The anomaly detection module 306 is used to determine the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
[0119] In some embodiments, the first state determination module 303 is specifically implemented to determine the load balancing state of the first device group based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device; and to determine the operating state data of the first device group based on the load balancing state.
[0120] In some embodiments, the first state determination module 303 is specifically implemented to obtain the abnormal detection results of multiple first devices in the first device group within a preset detection period; and to determine the operating state data of the first device group based on the performance data of the ports of the multiple first devices in the first device group and the abnormal detection results.
[0121] In some embodiments, the first state determination module 303 is specifically implemented to aggregate the performance data of the ports of multiple first devices in the first device group, and determine the operating state data of the first device group based on the aggregation processing result.
[0122] In some embodiments, the specific implementation is used to determine whether the target device has an anomaly based on the anomaly detection result; the strategy determination module is used to determine an anomaly handling strategy for the target device based on the operating status data of the first device group and / or the operating status data of the first line group when it is determined that the target device has an anomaly.
[0123] In some embodiments, the strategy determination module is specifically implemented as follows: when it is determined that there is a load imbalance in the first device group based on the operating status data of the first device group, the module determines the exception handling strategy for load balancing the first device group based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device.
[0124] In some embodiments, the strategy determination module is specifically implemented as follows: when determining, based on the operating status data of the first line group, that the number of second lines with signal interruption in the first line group is greater than a preset number threshold, it determines the abnormal handling strategy for lines that have a line stability correlation with the first line.
[0125] It should be noted that the units, devices, or modules described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. For ease of description, the above devices are described by dividing them into various modules according to their functions. Of course, in implementing this specification, the functions of each module can be implemented in one or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection between the devices or units shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0126] As can be seen from the above, a network device anomaly monitoring device provided in the embodiments of this specification receives an anomaly detection request for a target device in a target network; determines a first device group corresponding to the target device based on the topological connection relationship of the target device in the target network; wherein the first device group includes multiple first devices that have a device stability correlation with the target device; determines the operating status data of the first device group based on the performance data of the ports of the multiple first devices in the first device group; obtains the first line where the target device is located in the target network, and obtains a first line group corresponding to the first line; wherein the first line group includes multiple second lines that have a line stability correlation with the first line; determines the operating status data of the first line group based on the line operation status of the multiple second lines in the first line group; and determines the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
[0127] In a specific scenario example, the network device anomaly monitoring method and apparatus provided in this specification can be applied. From a business perspective, based on the operating status data of the first device group and the first line group corresponding to the target device, anomaly monitoring of the target device can be performed. This can ensure high efficiency in anomaly detection while avoiding business interruptions caused by device or line anomalies, thus ensuring business continuity.
[0128] Currently, network monitoring systems are generally capable of monitoring individual devices and lines, but their ability to monitor high-availability groups as a whole is insufficient. In addition to monitoring the operation of individual devices, maintenance personnel often need to prioritize the overall operation of the high-availability group in actual production operations, requiring a comprehensive assessment of the network's overall performance.
[0129] In some embodiments, see Figure 4 As shown, monitoring port traffic balancing for a high-availability device group (i.e., the first device group) can include:
[0130] S1: Input of data related to single-device network monitoring;
[0131] Specifically, the data inputs for single-device network monitoring include obtaining performance data of network device ports and real-time data of each port of a single device as inputs.
[0132] S2: High-availability port group (i.e., first device group) configuration data input;
[0133] Specifically, by analyzing the high availability relationships between network devices and the topological connections between them, certain ports across devices can be calculated to form a high availability group (i.e., the first device group). This high availability group (i.e., the first device group) serves as the configuration basis for subsequent calculations.
[0134] S3: Traffic statistics display for the high-availability port group (i.e., the first device group).
[0135] Specifically, based on the port high availability group configuration relationship obtained above, real-time data of each port of a single device is aggregated and calculated. This allows for the calculation of the maximum, minimum, average, and sample variance of port traffic within the high availability group (i.e., the first device group). This data is used to observe whether the traffic of each port within the high availability unit is normal, evenly distributed, and whether high availability is achieved. Performance curves are also displayed to facilitate troubleshooting by operations and maintenance personnel.
[0136] In some embodiments, see Figure 5 As shown, monitoring the connectivity of a high-availability line group (i.e., the first line group) can include:
[0137] S1: Input of relevant data for single-line monitoring;
[0138] Specifically, the data input for monitoring a single line includes obtaining the monitoring performance data of the single line and the real-time connection / disconnection data of the single line as input.
[0139] S2: High availability line group (i.e., the first line group) configuration data input;
[0140] Specifically, the high availability group of the network line can be identified through the high availability configuration relationship of the network line construction, and this high availability group serves as the configuration basis for subsequent calculations.
[0141] S3: Statistics display of high availability line group (i.e., the first line group).
[0142] Specifically, based on the aforementioned high availability group configuration, the real-time connectivity data of each individual line is aggregated and calculated. The outage status of all lines within the group is assessed; if all lines are down, an additional alarm is triggered, potentially affecting all services handled by that first line group, requiring urgent maintenance intervention.
[0143] Based on the above embodiments, alarms collected and generated by the network monitoring system serve as input for analysis. By summarizing and statistically analyzing device operating performance data (including key performance indicators of network devices such as traffic, error messages, packet loss, number of ports, temperature, and fan status) and alarm data (including major and minor alarms) based on the high availability group information of network lines and devices, the system displays the alarm and performance data of the high availability group. This provides network operation and maintenance colleagues with a method to observe the overall service quality of the network high availability group, enabling faster discovery of network faults and improved fault location capabilities.
[0144] While this specification provides the steps of operation for the methods described in the embodiments or flowcharts, more or fewer steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or client product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.
[0145] Those skilled in the art will also know that, besides implementing the controller using purely computer-readable program code, the same functions can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices within it used to implement various functions can also be considered structures within that hardware component. Alternatively, the devices used to implement various functions can be considered as both software modules implementing the method and structures within a hardware component.
[0146] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of this specification can essentially be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of this specification.
[0147] Although this specification has been described by way of examples, those skilled in the art will recognize that many variations and modifications are possible without departing from the spirit of this specification, and it is intended that the appended claims cover such variations and modifications without departing from the spirit of this specification.
Claims
1. A method for monitoring network device anomalies, characterized in that, include: Receive anomaly detection requests for target devices in the target network; Based on the topological connection relationship of the target device in the target network, obtain multiple devices in the target network that have a cooperative relationship with the target device, and based on preset device configuration information, obtain multiple devices in the target network that have a primary / backup relationship with the target device. The acquired multiple devices are identified as first devices, and a first device group corresponding to the target device is constructed based on the multiple first devices; wherein, the first device group contains multiple first devices that have a high availability relationship with the target device; Based on the performance data of the ports of multiple first devices in the first device group, the performance data of the port of the target device, the preset load weight corresponding to each first device, and the load weight corresponding to the target device, the load balancing status of the first device group is determined, and the load balancing status is used as the operating status data of the first device group. The system acquires a first line in the target network where the target device is located, acquires multiple lines with the same data processing nodes as the first line, and acquires multiple lines in the target network that have a primary / backup relationship and / or bandwidth sharing relationship with the first line according to preset line configuration information; the acquired multiple lines are determined as second lines, and a first line group corresponding to the first line is constructed based on the multiple second lines; wherein, the first line group includes multiple second lines that have a line stability correlation relationship with the first line; Based on the operational status of multiple second lines in the first line group, determine the operational status data of the first line group; Based on the operating status data of the first equipment group and / or the operating status data of the first line group, determine the anomaly detection result for the target equipment.
2. The method according to claim 1, characterized in that, The step of determining the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group includes: Obtain abnormal detection results for multiple first devices in the first device group within a preset detection period; Based on the performance data of the ports of multiple first devices in the first device group and the anomaly detection results, the operating status data of the first device group is determined.
3. The method according to claim 1, characterized in that, The step of determining the operating status data of the first device group based on the performance data of the ports of multiple first devices in the first device group includes: The performance data of the ports of multiple first devices in the first device group are aggregated, and the operating status data of the first device group is determined based on the aggregation results.
4. The method according to claim 1, characterized in that, The method further includes: Based on the anomaly detection results, determine whether the target device has an anomaly; If it is determined that the target device is abnormal, an abnormality handling strategy for the target device is determined based on the operating status data of the first device group and / or the operating status data of the first line group.
5. The method according to claim 4, characterized in that, The step of determining anomaly handling strategies for the target device based on the operating status data of the first device group includes: If, based on the operating status data of the first device group, it is determined that the first device group has an unbalanced load, the anomaly handling strategy for load balancing the first device group is determined based on the performance data of the ports of multiple first devices in the first device group and the performance data of the port of the target device.
6. The method according to claim 5, characterized in that, The step of determining anomaly handling strategies for the target device based on the operating status data of the first line group includes: If, based on the operating status data of the first line group, it is determined that the number of second lines with signal interruption in the first line group is greater than a preset threshold, the abnormal handling strategy for lines that have a line stability correlation with the first line group is determined.
7. A network device anomaly monitoring device, characterized in that, include: The anomaly detection request module is used to receive anomaly detection requests for target devices in the target network; The first group of determining modules is used to obtain multiple devices in the target network that have a cooperative relationship with the target device based on the topological connection relationship of the target device in the target network, and to obtain multiple devices in the target network that have a primary / backup relationship with the target device based on preset device configuration information; The acquired multiple devices are identified as first devices, and a first device group corresponding to the target device is constructed based on the multiple first devices; wherein, the first device group contains multiple first devices that have a high availability relationship with the target device; The first state determination module is used to determine the load balancing state of the first device group based on the performance data of the ports of multiple first devices in the first device group, the maximum and minimum values of the performance data of the port of the target device, the difference between the maximum and minimum values and the preset load weight of each device, and to use the load balancing state as the operating state data of the first device group. The second group of determining modules is used to obtain the first line where the target device is located in the target network, obtain multiple lines with the same data processing nodes as the first line, and obtain multiple lines in the target network that have a primary / backup relationship and / or bandwidth sharing relationship with the first line according to preset line configuration information; determine the multiple lines obtained as second lines, and construct a first line group corresponding to the first line based on the multiple second lines; wherein, the first line group includes multiple second lines that have a line stability correlation relationship with the first line; The second state determination module is used to determine the operating state data of the first line group based on the operating status of multiple second lines in the first line group. The anomaly detection module is used to determine the anomaly detection result for the target device based on the operating status data of the first device group and / or the operating status data of the first line group.
8. An electronic device, characterized in that, It includes a processor and a memory for storing processor-executable instructions, wherein the processor, when executing the instructions, implements the steps of the network device anomaly monitoring method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of the network device anomaly monitoring method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Network processing method, device and equipment
CN119011409A