Method, device, medium and program product for monitoring network performance

By adjusting the monitoring strategy in a targeted manner and dynamically adjusting the monitoring intensity according to the device load and type, the problems of waste of monitoring resources and low efficiency in existing technologies are solved, and the optimization and stable operation of network performance are achieved.

CN120499057BActive Publication Date: 2025-09-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510955102.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-26
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

In the existing technology, network monitoring strategies cannot perform differentiated processing based on the importance and load conditions of different devices, resulting in insufficient monitoring of high-load devices or excessive monitoring of low-load devices, causing resource waste and low monitoring efficiency.

Method used

Adopt targeted monitoring strategies, reduce monitoring intensity for high-load devices, increase monitoring intensity for low-load devices, dynamically adjust monitoring strategies based on load conditions, and optimize monitoring resource configuration based on device type and topology.

Benefits of technology

It realizes differentiated monitoring of different devices, optimizes monitoring resource configuration, improves network performance monitoring efficiency, and ensures stable and secure network operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499057B_ABST
    Figure CN120499057B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, medium, and program product for monitoring network performance, which can be applied to the fields of network management and performance engineering technology. The method includes: monitoring a first device in a target network using a first monitoring policy to determine a load status of the first device; monitoring a second device in the target network that is different from the first device using a second monitoring policy to determine a load status of the second device, wherein the monitoring intensity of the second monitoring policy is lower than the monitoring intensity of the first monitoring policy; updating the first monitoring policy to a monitoring policy with a reduced monitoring intensity based on the load status of the first device indicating an increased load; and updating the second monitoring policy to a monitoring policy with an increased monitoring intensity based on the load status of the second device indicating an increased load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of network management and performance engineering, and in particular to a method, device, medium and program product for monitoring network performance. Background Art

[0002] With the rapid development of network technology, network performance monitoring is crucial to ensuring stable network operation and improving user experience. As network scale continues to expand, network device types are becoming increasingly complex and diverse, and different devices play different roles and have different importance in the network.

[0003] Related technologies often employ a unified monitoring strategy to monitor all devices in the network. This unified management approach fails to differentiate between devices based on their importance and load. Highly loaded devices in the network suffer from insufficient monitoring, making it difficult to obtain timely and accurate critical information when their load fluctuates, making it difficult to ensure the stable operation of core devices. Excessive monitoring of low-load devices in the network wastes monitoring resources, reduces monitoring efficiency, and increases network management costs. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method, device, medium and program product for monitoring network performance.

[0005] According to a first aspect of the present invention, a method for monitoring network performance is provided, comprising: monitoring a first device in a target network using a first monitoring strategy to determine a load condition of the first device; monitoring a second device in the target network that is different from the first device using a second monitoring strategy to determine a load condition of the second device, wherein a monitoring intensity of the second monitoring strategy is lower than a monitoring intensity of the first monitoring strategy; for the first device, updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity based on an indication that the load is increasing by its load condition; and updating the second monitoring strategy to a monitoring strategy with increased monitoring intensity based on an indication that the load is increasing by its load condition.

[0006] The second aspect of the present invention provides an apparatus for monitoring network performance, comprising: a first device monitoring module, for monitoring a first device in a target network using a first monitoring strategy to monitor the first device to determine the load status of the first device; a second device monitoring module, for monitoring a second device in the target network that is different from the first device using a second monitoring strategy to determine the load status of the second device, wherein the monitoring intensity of the second monitoring strategy is lower than the monitoring intensity of the first monitoring strategy; a first policy updating module, for updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity for the first device based on an indication of an increase in load from its load status; and a second policy updating module, for updating the second monitoring strategy to a monitoring strategy with increased monitoring intensity for the second device based on an indication of an increase in load from its load status.

[0007] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0011] Figure 1 An architectural diagram showing a method, device, medium, and program product for monitoring network performance according to an embodiment of the present invention is shown.

[0012] Figure 2 A flow chart of a method for monitoring network performance according to an embodiment of the present invention is shown.

[0013] Figure 3 A timing diagram of a method for monitoring network performance according to an embodiment of the present invention is shown.

[0014] Figure 4 A structural block diagram of an apparatus for monitoring network performance according to an embodiment of the present invention is shown.

[0015] Figure 5A block diagram of an electronic device suitable for implementing a method for monitoring network performance according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0016] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0017] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0019] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0020] The existing monitoring strategy adjustment mechanism is relatively simple, mostly based on a single load indicator for judgment, and lacks comprehensive consideration of multiple factors such as device type and duration of monitoring strategy. As a result, the monitoring strategy adjustment is not scientific and reasonable enough, prone to misjudgment, difficult to adapt to the complex and changing network environment, and difficult to achieve optimal configuration of monitoring resources and effective improvement of network performance monitoring efficiency.

[0021] An embodiment of the present invention provides a method for monitoring network performance, wherein a first monitoring strategy is used to monitor a first device in a target network to determine the load status of the first device; a second monitoring strategy is used to monitor a second device in the target network that is different from the first device to determine the load status of the second device, wherein the monitoring intensity of the second monitoring strategy is lower than the monitoring intensity of the first monitoring strategy; for the first device, based on its load status indicating that the load is increasing, the first monitoring strategy is updated to a monitoring strategy with reduced monitoring intensity; and for the second device, based on its load status indicating that the load is increasing, the second monitoring strategy is updated to a monitoring strategy with increased monitoring intensity.

[0022] Figure 1 An architectural diagram showing a method, device, medium, and program product for monitoring network performance according to an embodiment of the present invention is shown.

[0023] like Figure 1 As shown, the architecture 100 according to this embodiment may include a server 110, a target network 120, and a network 130. The target network 120 may be a hardware network composed of multiple hardware devices. The target network 120 includes at least one first device 121 and at least one second device 122. The target network 120 may also include at least one computing node 123. The first device 121 and the second device 122 may be interconnected in any manner, and the second device 122 and the computing node 123 may be interconnected in any manner. A direct connection relationship may not exist between the first device 121 and the computing node 123. The server 110 and the target network 120 interact via the network 130.

[0024] Server 110 may include a policy calculation module and a policy generation module. The policy calculation module is configured to calculate the load status of multiple devices in target network 120 based on parameters transmitted from target network 120 via network 130. Upon determining that a device in target network 120 requires a change in monitoring policy, the policy generation module sends the modified monitoring policy to the policy generation module. The policy generation module includes multiple basic monitoring protocols, such as NetFlow, sFlow, and Telemetry. Upon receiving the modified monitoring policy from the policy calculation module, the module configures and combines the multiple basic monitoring protocols to obtain a modified monitoring policy, which it then sends to target network 120.

[0025] The target network 120 monitors the at least one first device 121 and the at least one second device 122 respectively using the monitoring policies corresponding to the device types, and reports the parameter results obtained from the monitoring to the server 110. When the target network 120 receives the modified monitoring policy sent by the server 110, it sends the modified monitoring policy to the device corresponding to the monitoring policy and updates the monitoring policy of the device.

[0026] It should be noted that the method for monitoring network performance provided in the embodiment of the present invention can generally be performed by the server 110. Accordingly, the apparatus for monitoring network performance provided in the embodiment of the present invention can generally be set in the server 110. The method for monitoring network performance provided in the embodiment of the present invention can also be performed by a server or server cluster that is different from the server 110 and can communicate with the target network 120, the first device 121, the second device 122 and / or the server 110. Accordingly, the apparatus for monitoring network performance provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 110 and can communicate with the target network 120, the first device 121, the second device 122 and / or the server 110.

[0027] It should be understood that Figure 1 The numbers of servers, target networks, first devices, second devices, computing nodes, and networks are merely illustrative. Any number of servers, target networks, first devices, second devices, computing nodes, and networks may be provided as needed.

[0028] The following will be based on Figure 1 The scene described by Figure 2~Figure 3 The method for monitoring network performance according to an embodiment of the present invention is described in detail.

[0029] Figure 2 A flow chart of a method for monitoring network performance according to an embodiment of the present invention is shown.

[0030] like Figure 2 As shown, the method for monitoring network performance of this embodiment includes operations S210 to S240.

[0031] In operation S210 , a first device in a target network is monitored using a first monitoring strategy to determine a load status of the first device.

[0032] In operation S220 , a second device different from the first device in the target network is monitored using a second monitoring strategy to determine a load status of the second device.

[0033] According to an embodiment of the present invention, the target network may include a first device and a second device. The load of the first device may be higher than the load of the second device. For example, the first device may be a core device in the target network, and the second device may be a non-core device in the target network. However, the present invention is not limited to this. The first device and the second device may be any two different devices in the target network that apply different monitoring strategies. The target network may be a network architecture set up to support the computing tasks of the intelligent computing center. During the process of the intelligent computing center performing training tasks, the loads of multiple devices in the intelligent computing center may change.

[0034] According to an embodiment of the present invention, a first monitoring strategy is applied to monitor the first device, and a second monitoring strategy is applied to monitor the second device, so as to obtain the respective operating parameters of the first device and the second device, and further determine the respective load conditions of the first device and the second device based on the obtained operating parameters.

[0035] According to an embodiment of the present invention, if the load on the first device is higher than the load on the second device, the first device can be monitored more strictly. Therefore, the monitoring intensity of the second monitoring policy can be lower than that of the first monitoring policy. The monitoring policy may include the number of sampling parameters, the sampling frequency, the monitoring protocol used for sampling, etc. When monitoring a device, the greater the number of corresponding sampling parameters and the higher the sampling frequency in the monitoring policy, the higher the monitoring intensity of the monitoring policy. A stronger monitoring policy may require the consumption of more resources.

[0036] In operation S230 , for the first device, based on the load status indicating that the load of the first device is increasing, the first monitoring strategy is updated to a monitoring strategy with reduced monitoring intensity.

[0037] According to the embodiment of the present invention, since monitoring a device requires consuming resources including the device's own resources, the higher the intensity of the monitoring policy, the higher the load brought by monitoring the device.

[0038] According to an embodiment of the present invention, when the load on the first device is relatively high, and when the load condition of the first device indicates that the load on the first device has further increased, the monitoring intensity of the first monitoring policy can be lowered to reduce the load caused by monitoring the first device, thereby minimizing the load on the first device from exceeding its load threshold and ensuring the normal, secure and stable operation of the first device and the target network. According to some embodiments, monitoring losses caused by a decrease in monitoring intensity, such as the loss of certain operating parameters, can be indirectly obtained through other means.

[0039] In operation S240 , for the second device, based on the load status indicating that the load of the second device is increasing, the second monitoring policy is updated to a monitoring policy with increased monitoring intensity.

[0040] According to an embodiment of the present invention, when the load of the second device is relatively low, and when the load condition of the second device indicates that the load of the second device has increased, it means that the second device may have business needs such as resource optimization allocation and equipment failure analysis. Therefore, it is necessary to increase the monitoring intensity of the second monitoring strategy to obtain more detailed and accurate operating parameters and load conditions to meet the data support required for the above-mentioned business needs.

[0041] According to an embodiment of the present invention, by adopting monitoring strategies of different intensities for the first device and the second device in the target network, and dynamically adjusting the monitoring strategies according to the load conditions, differentiated monitoring can be performed based on the importance and load conditions of different devices. More specifically, for the first device with a high load, when its load further increases, monitoring resources are reasonably allocated, and some resources are released by reducing the monitoring intensity so that the first device can bear the increased load, thereby ensuring the stable operation of the first device. In addition, for the second device with a relatively low load, when its load increases, the monitoring intensity is increased so that the collected operating parameters and load conditions can meet the business needs such as fault analysis that may occur after the load of the second device increases, thereby ensuring the stable operation of the second device. By making differentiated adjustments to the monitoring strategies of different devices, it is possible to achieve optimal configuration of monitoring resources, improve the efficiency of overall target network performance monitoring, and ensure the stable and safe operation of the target network.

[0042] According to an embodiment of the present invention, the method for monitoring network performance may further include, for each of the first device and the second device: determining a policy switching evaluation value based on the load condition, the device type of the device and the duration of the monitoring policy corresponding to the device; and when the policy switching evaluation value is higher than the first switching threshold corresponding to the device type, determining that the load condition indicates that the load has increased.

[0043] According to an embodiment of the present invention, the load status of a network device can be determined in the following manner: the load status of the device can be determined based on one or more of the resource utilization rate and congestion index of the device. For example, the resource utilization rate can include at least one of the device's port transmit bandwidth utilization rate, receive bandwidth utilization rate, processor utilization rate, and memory utilization rate, and the congestion index can include at least one of the port queue buffer level and network congestion frame count.

[0044] According to an embodiment of the present invention, the load status L(T) of the network device can be determined by formula (1):

[0045] (1)

[0046] Among them, n represents the number of ports on the device, and m represents the number of port queues on the device. represents the sending bandwidth of the i-th port, represents the receiving bandwidth of the i-th port, Indicates the maximum value of the sum of the sending bandwidth and the receiving bandwidth among n ports. Indicates the buffer water level of the j-th queue, Indicates the maximum value of the buffer water level in m queues, is the processor usage, is the maximum value of processor usage, is the memory usage, is the maximum value of memory usage, are all hyperparameters and satisfy , preferably, can be taken , Indicates the increment of network congestion frames, where network congestion frames include Priority Flow Control (PFC) frames and / or Explicit Congestion Notification (ECN) frames.

[0047] It should be noted that although formula (1) shows an example for calculating the load condition, the present invention is not limited thereto. For example, some parameters may be removed from formula (1), some parameters may be replaced, or other parameters may be added.

[0048] According to an embodiment of the present invention, the increment of network congestion frames can be obtained by subtracting the number of network congestion frames currently collected during device monitoring from the number of network congestion frames used in the last calculation of the load status of the network device.

[0049] According to embodiments of the present invention, load status is determined using one or more of a device's resource utilization and congestion indicators. Furthermore, resource utilization and congestion indicators can encompass multiple parameters, enriching the dimensions of load assessment. Comprehensively assessing device load from multiple perspectives can more comprehensively and accurately reflect the device's actual operating status, avoiding the one-sidedness caused by single-indicator judgments. This provides more accurate baseline data for subsequent monitoring strategy adjustments and ensures the effectiveness of network performance monitoring.

[0050] According to an embodiment of the present invention, since switching of monitoring strategies will also cause consumption of computing resources, if the switching of monitoring strategies is too frequent, it will lead to a large amount of waste of computing resources. Therefore, for each device, the update frequency of the monitoring strategy can be controlled according to the duration of the monitoring strategy corresponding to the device.

[0051] According to an embodiment of the present invention, the calculation of the evaluation policy switching value may also be controlled according to the device type of each device, so as to achieve differentiated management of devices of different types and their corresponding monitoring policies.

[0052] According to an embodiment of the present invention, determining the policy switching evaluation value may include: determining the switching weight corresponding to the device type; determining the time decay item based on the duration; and obtaining the policy switching evaluation value based on the switching weight, the time decay item and the load condition.

[0053] According to an embodiment of the present invention, the strategy switching evaluation value T can be calculated using formula (2):

[0054] (2)

[0055] Where W represents the switching weight, which is different for different device types: for example, for the first device mentioned above, or the core device, W=0.4; for example, for the second device mentioned above, or the non-core device, W=1.25. λ is a hyperparameter used to control the impact of load duration on the strategy switching evaluation value. t is the duration of the current monitoring strategy. is the time decay term.

[0056] According to an embodiment of the present invention, the calculation method for the policy switching evaluation value is further refined by determining the switching weight corresponding to the device type, determining the time decay term based on the duration, and combining the load status to obtain the policy switching evaluation value. Taking into account the differences in the importance of different device types and the impact of time factors on monitoring policies, the calculation of the policy switching evaluation value is made more refined and personalized, thereby more accurately assessing device load changes, making the adjustment of monitoring policies more consistent with the actual needs of the devices, and enhancing the rationality and adaptability of network performance monitoring.

[0057] Although two types of devices (eg, core devices and non-core devices) are exemplified herein, the present invention is not limited thereto. For example, devices may be divided into more categories based on their importance in the target network.

[0058] According to an embodiment of the present invention, when the policy switching evaluation value is higher than the first switching threshold corresponding to the device type, it can be determined that the load condition is sufficiently large for the duration of the current monitoring policy corresponding to the device. Therefore, it can be determined that the load condition indicates that the load is increasing. For example, for different device types, the first switching threshold can be 5 (for example, for the first device or core device described above) or 2.3 (for example, for the second device or non-core device described above). By setting different first switching thresholds for different device types, the differentiated management solution can be further refined, making switching more reasonable and effective.

[0059] According to an embodiment of the present invention, a policy switching evaluation value is determined based on load conditions, device type, and the duration of the monitoring policy. The value is then compared with a first switching threshold to determine whether the load has increased, making monitoring policy adjustments more scientific and reasonable. This approach fully considers multiple factors, including device characteristics, current load conditions, and policy execution time, avoiding misjudgments caused by a single factor. It can more accurately identify changes in device load, providing a reliable basis for timely and accurate monitoring policy adjustments, and improving the accuracy and reliability of network performance monitoring.

[0060] According to an embodiment of the present invention, device types may include core devices and non-core devices, and the method for monitoring network performance may also include: constructing an interconnection topology of devices in the target network including a first device and a second device based on a device topology management protocol, and determining the device type of the device based on the interconnection topology.

[0061] According to an embodiment of the present invention, the device topology management protocol may be a Link Layer Discovery Protocol (LLDP). LLDP is used to discover network devices within a local area network and to build an interconnection topology between the network devices.

[0062] According to an embodiment of the present invention, after the interconnection topology is determined, the device type of each of the multiple network devices included in the interconnection topology can be determined based on the connection status between each of the multiple network devices and other network devices.

[0063] For example, if the interconnection topology reveals that a network device is directly connected to a computing node in the target network, the network device may need to communicate with the edge computing node, making it difficult to focus on the core tasks at the target network level. Therefore, the network device can be determined to be a non-core device. On the other hand, if the interconnection topology reveals that a network device is not directly connected to a computing node in the target network, the network device may not need to communicate with the edge computing node. Therefore, the network device can focus on processing the core tasks at the target network level. Therefore, the network device can be determined to be a core device.

[0064] According to embodiments of the present invention, building a network device interconnection topology and determining device types based on the device topology management protocol provides a solid foundation for implementing differentiated monitoring strategies. By clearly defining the types and location relationships of various devices in the network, different monitoring strategies can be more targeted for core and non-core devices, optimizing network management and monitoring resource allocation, improving the manageability of the network architecture and the effectiveness of monitoring strategies, and ensuring stable network operation.

[0065] According to an embodiment of the present invention, when the first device is a core device, based on its load condition indicating that the load has increased, updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity may include: obtaining the opposite port of the target port with relatively low traffic through the interconnection topology; closing the real-time traffic monitoring of the target port, and using at least one of the sending traffic and receiving traffic of the opposite port to convert at least one of the sending traffic and receiving traffic of the target port.

[0066] According to an embodiment of the present invention, for a first device, target ports with relatively low traffic are determined among at least one port of the first device. The number of target ports can be determined based on the need to reduce monitoring intensity. For example, if reducing monitoring of three ports is necessary to free up sufficient resources, the three ports with the lowest traffic may be determined as target ports.

[0067] According to an embodiment of the present invention, after determining the peer port directly connected to the target port based on the topology, real-time traffic monitoring of the target port is stopped, and the real-time traffic of the peer port is determined based on the monitoring results of the device where the peer port is located. Because the peer port and the target port are directly connected one-to-one, they are typically in a one-to-one configuration, with one port sending and the other port receiving. Therefore, the traffic of the peer port and the target port is generally the same in value, and the send / receive traffic of the target port can be converted based on the send / receive traffic of the peer port.

[0068] According to an embodiment of the present invention, when core devices become heavily loaded, real-time traffic monitoring of the target port is disabled and traffic is converted by acquiring the peer port of the target port with relatively low traffic. This reduces monitoring intensity while ensuring the accuracy and effectiveness of traffic monitoring. This reduces monitoring resource usage when core devices are highly loaded, avoids the impact of excessive monitoring on core device performance, and achieves a balance between monitoring intensity adjustment and monitoring accuracy, ensuring stable operation of core devices and network performance.

[0069] According to an embodiment of the present invention, updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity based on a load condition indicating an increase in load may further include: determining, based on the topology, connection relationships between at least some of the ports on the first device and ports of other devices, defining the ports that are connected to the ports of the other devices as directly connected ports, and determining the peer ports of the directly connected ports. Monitoring of the directly connected ports is disabled, and the send / receive traffic of the directly connected ports is converted using the send / receive traffic of the peer port.

[0070] According to an embodiment of the present invention, by analyzing the topology, ports that are connected to ports of other devices are selected as direct-connected ports, and monitoring of these direct-connected ports is stopped first. This ensures that after stopping monitoring, the send / receive traffic of the opposite port connected to the direct-connected port can be used to convert the send / receive traffic of the direct-connected port, thereby ensuring the feasibility of traffic conversion.

[0071] According to an embodiment of the present invention, when the second device is a non-core device, the second monitoring strategy is used to monitor the parameters of the second device using the first monitoring protocol, and based on its load condition indicating that the load has increased, updating the second monitoring strategy to a monitoring strategy with increased monitoring intensity may include: increasing the sampling frequency of the second monitoring strategy; when it is determined that a parameter mutation occurs in the second device, using the second monitoring protocol to monitor the target parameter that has undergone the parameter mutation, and stopping the first monitoring protocol from monitoring the target parameter.

[0072] According to an embodiment of the present invention, the first monitoring protocol may represent periodic collection of parameters of the second device. Therefore, the sampling frequency of the second monitoring strategy may be increased, and parameter collection may be performed in a shorter period to improve the monitoring intensity of the second monitoring strategy.

[0073] According to an embodiment of the present invention, when the load condition indicates an increase in load, multiple parameters collected using the second monitoring strategy can be analyzed and evaluated to determine whether any of the multiple parameters contain a target parameter that has undergone a sudden change. The range of change corresponding to the sudden change can vary for different parameters.

[0074] For example, for send / receive traffic, if the rate of change in traffic exceeds 30%, a sudden change in that parameter can be determined, and send / receive traffic can be used as the target parameter. For another example, for queue buffer level, if the queue buffer level exceeds 75%, a sudden change in that parameter can be determined, and the queue buffer level can be used as the target parameter. For another example, for network congestion frames, if the increase in network congestion frames exceeds 50%, a sudden change in that parameter can be determined, and the network congestion frame number can be used as the target parameter.

[0075] According to an embodiment of the present invention, when it is determined that a sudden change in the parameters of the second device occurs through the above method, the second monitoring strategy can be adjusted to switch the monitoring protocol for the target parameter from the first monitoring protocol to the second monitoring protocol. The second monitoring protocol can indicate continuous collection of the parameters of the second device.

[0076] According to an embodiment of the present invention, the first monitoring protocol may include a network stream protocol and a sampling stream protocol, which are used to periodically obtain device parameters through active polling to achieve device monitoring. In addition, the second monitoring protocol may include a telemetry protocol, which is used to obtain device parameters through subscription to achieve device monitoring.

[0077] According to an embodiment of the present invention, the device parameters that need to be monitored can be subscribed to through the second monitoring protocol. After the subscribed device parameters change or are updated, the changes will be directly synchronized to the monitoring side to achieve real-time knowledge of the changes in the device parameters that need to be monitored.

[0078] According to embodiments of the present invention, appropriate monitoring protocols can be selected for different devices under different load conditions based on the specific types and operating modes of the first and second monitoring protocols. Combining active polling and subscription methods allows for flexible switching based on device load and demand. While ensuring effective monitoring, this approach leverages the strengths of different protocols, improving monitoring efficiency, optimizing network resource utilization, and enhancing the overall effectiveness of network performance monitoring.

[0079] According to an embodiment of the present invention, for non-core devices, the sampling frequency of the monitoring strategy is increased when the load increases, and the monitoring protocol is switched to the target parameter for parameter mutations, which can promptly capture abnormal changes in the device. Without wasting excessive resources, it can quickly respond to changes in device load, improve the monitoring ability of abnormal conditions in non-core devices, promptly identify potential problems, enhance network stability and reliability, and ensure normal network operation.

[0080] According to an embodiment of the present invention, the method for monitoring network performance may further include: upon determining that the number of parameters collected using the second monitoring protocol on the second device meets a preset condition, monitoring all parameters of the second device using the second monitoring protocol. According to an embodiment, the preset condition may include: the number of parameters reaching a preset percentage of all parameters of the second device. For example, the preset percentage may be 50%, meaning that if more than half of the device parameters on the second device are collected and monitored using the second monitoring protocol, the second monitoring strategy is switched to collecting and monitoring all device parameters of the second device using the second monitoring protocol.

[0081] According to an embodiment of the present invention, the preset condition is set to a predetermined percentage of the total number of parameters of the second device, making the triggering conditions for monitoring policy upgrades clearer and more quantifiable. This facilitates the system's accurate determination of when to fully switch monitoring protocols, ensuring standardized and consistent monitoring policy adjustments, avoiding inappropriate adjustments due to ambiguous conditions, and improving the operability and accuracy of network performance monitoring.

[0082] According to an embodiment of the present invention, when the number of parameters collected by a second device using a second monitoring protocol meets a preset condition, the protocol is used to monitor all parameters, achieving dynamic upgrades to the monitoring strategy. This allows for the gradual expansion of high-precision monitoring scope based on actual device needs, while ensuring effective monitoring while rationally controlling monitoring resource investment and improving monitoring efficiency. This ensures that monitoring of non-core devices can be flexibly and effectively adjusted based on load changes.

[0083] According to an embodiment of the present invention, the method for monitoring network performance may further include: for the first device, after updating the first monitoring policy to a monitoring policy with reduced monitoring intensity, based on its load condition indicating that the load has decreased, restoring the first monitoring policy to the initial first monitoring policy; and for the second device, after updating the second monitoring policy to a monitoring policy with increased monitoring intensity, based on its load condition indicating that the load has decreased, restoring the second monitoring policy to the initial second monitoring policy.

[0084] According to an embodiment of the present invention, when the policy switching evaluation value is lower than the second switching threshold corresponding to the device type, the load condition is determined to indicate a decrease in load. The second switching threshold corresponding to each device type can be lower than the first switching threshold corresponding to that device type. For example, the second switching threshold can be 2.5 (for example, for the first device or core device described above) or 0.3 (for example, for the second device or non-core device described above). By setting different second switching thresholds for different device types, differentiated management solutions can be further refined, making switching more efficient and effective.

[0085] According to the embodiments of the present invention, by comparing the load with the second switching threshold, a clear quantitative standard is provided for restoring the monitoring policy. This avoids the randomness and uncertainty of monitoring policy restoration, ensures that the adjustment and restoration of the monitoring policy have an accurate and clear basis, improves the standardization and reliability of network performance monitoring, and enables the network monitoring system to operate more stably and efficiently.

[0086] According to an embodiment of the present invention, the monitoring policies of the first and second devices are restored to their initial policies when the load decreases, achieving dynamic and reversible adjustment of the monitoring policies. This ensures that the devices return to their optimal monitoring state after the load returns to normal, avoids interference with normal device operation caused by excessive adjustments to the monitoring policies, maintains the stability and continuity of network performance monitoring, and ensures long-term stable network operation.

[0087] Figure 3 A timing diagram of a method for monitoring network performance according to an embodiment of the present invention is shown.

[0088] like Figure 3As shown, the process of monitoring network performance corresponding to the timing diagram includes operations S301 to S319.

[0089] In operation S301 , a link layer discovery protocol is used to discover network devices in a target network, and parameters of the network devices are sent to a server.

[0090] In operation S302, the server uses a link layer discovery protocol to construct an interconnection topology based on network devices, and classifies the network devices according to the interconnection topology to determine a first device and a second device.

[0091] In operation S303 , the second device is monitored using a second monitoring policy, where the second monitoring policy includes the first monitoring protocol.

[0092] In operation S304 , the first device is monitored using a first monitoring policy, where the first monitoring policy includes a second monitoring protocol.

[0093] In operation S305 , the second device returns the acquired data to the server.

[0094] In operation S306 , the server calculates a load status of the second device based on the acquired data of the second device.

[0095] In operation S307 , when the load status of the second device indicates that the load of the second device is increasing, the frequency of polling using the first monitoring protocol is increased.

[0096] In operation S308 , the server analyzes and evaluates parameters of the second device.

[0097] In operation S309 , when the analysis and evaluation determines that a sudden change occurs in the target parameter of the second device, the target parameter is monitored using a second monitoring protocol.

[0098] In operation S310 , the second device returns the number of target parameters currently monitored using the second monitoring protocol to the server.

[0099] In operation S311 , the server judges the target parameter to determine whether it meets a preset condition.

[0100] In operation S312 , when the target parameter meets a preset condition, all parameters of the second device are monitored using a second monitoring protocol.

[0101] In operation S313 , the first device returns the acquired data to the server.

[0102] In operation S314, the server calculates a load status of the first device based on the acquired data of the first device.

[0103] In operation S315 , when the load status of the first device indicates that the load of the first device has increased, monitoring of some ports of the first device is turned off.

[0104] In operation S316 , the server determines the peer port of the closed port through the interconnection topology.

[0105] In operation S317 , the server is used to request parameters of the peer port.

[0106] In operation S318 , the first device and / or the second device returns the parameters of the peer port to the server.

[0107] In operation S319, the parameters of the target port are converted according to the parameters of the peer port.

[0108] In the above monitoring process, the order between operation S303 and operation S304 is not limited, and the two can be performed simultaneously. The order between operation S305 to operation S312 and operation S313 to operation S315 is not limited, and the two can be performed simultaneously.

[0109] Based on the above method for monitoring network performance, the present invention also provides a device for monitoring network performance. Figure 4 The device is described in detail.

[0110] Figure 4 A structural block diagram of an apparatus for monitoring network performance according to an embodiment of the present invention is shown.

[0111] like Figure 4 As shown, the apparatus 400 for monitoring network performance of this embodiment includes a first device monitoring module 410 , a second device monitoring module 420 , a first policy updating module 430 and a second policy updating module 440 .

[0112] The first device monitoring module 410 is used to monitor the first device in the target network using a first monitoring strategy to determine the load status of the first device. In one embodiment, the first device monitoring module 410 can be used to perform the operation S210 described above, which will not be repeated here.

[0113] The second device monitoring module 420 is configured to monitor a second device different from the first device in the target network using a second monitoring policy to determine a load status of the second device, wherein the monitoring intensity of the second monitoring policy is lower than the monitoring intensity of the first monitoring policy. In one embodiment, the second device monitoring module 420 can be configured to perform operation S220 described above, which is not further described here.

[0114] The first policy updating module 430 is configured to update the first monitoring policy to a monitoring policy with reduced monitoring intensity for the first device based on the load condition indicating that the load has increased. In one embodiment, the first policy updating module 430 may be configured to perform the aforementioned operation S230, which will not be described in detail herein.

[0115] The second policy updating module 440 is configured to update the second monitoring policy to a monitoring policy with increased monitoring intensity for the second device based on the load condition indicating that the load has increased. In one embodiment, the second policy updating module 440 may be configured to perform the aforementioned operation S240, which will not be described in detail herein.

[0116] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes an evaluation value determination module and a first load change module.

[0117] The evaluation value determination module is used to determine the strategy switching evaluation value based on the load condition, the device type of the device and the duration of the monitoring strategy corresponding to the device.

[0118] The first load change module is configured to determine that the load condition indicates that the load has increased when the policy switching evaluation value is higher than a first switching threshold corresponding to the device type.

[0119] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes a load status determination module.

[0120] A load status determination module is used to determine the load status of the device based on one or more of the resource utilization rate and congestion index of the device, wherein the resource utilization rate includes at least one of the port sending bandwidth utilization rate, receiving bandwidth utilization rate, processor utilization rate, and memory utilization rate of the device, and the congestion index includes at least one of the port queue buffer water level and network congestion frame count.

[0121] According to an embodiment of the present invention, the evaluation value determination module includes a weight determination submodule, a time item determination submodule and an evaluation value determination submodule.

[0122] The weight determination submodule is used to determine the switching weight corresponding to the device type.

[0123] The time term determination submodule is used to determine the time decay term according to the duration.

[0124] The evaluation value determination submodule is used to obtain the strategy switching evaluation value according to the switching weight, the time decay term and the load condition.

[0125] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes a topology building module.

[0126] The topology construction module is used to construct an interconnection topology of devices including the first device and the second device in the target network based on the device topology management protocol, and determine the device type of the device according to the interconnection topology.

[0127] According to an embodiment of the present invention, the first policy updating module 430 includes a port determination submodule and a traffic conversion submodule.

[0128] The port determination submodule is used to obtain the opposite port of the target port with relatively low traffic through the interconnection topology.

[0129] The traffic conversion submodule is used to close the real-time traffic monitoring of the target port and use at least one of the sending traffic and receiving traffic of the opposite port to convert at least one of the sending traffic and receiving traffic of the target port.

[0130] According to an embodiment of the present invention, the second policy updating module 440 includes a frequency increasing submodule and a protocol switching submodule.

[0131] The frequency increasing submodule is used to increase the sampling frequency of the second monitoring strategy.

[0132] The protocol switching submodule is used to monitor the target parameter with the parameter mutation using the second monitoring protocol when it is determined that the parameter mutation occurs in the second device, and stop monitoring the target parameter with the first monitoring protocol.

[0133] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes a protocol switching module.

[0134] The protocol switching module is used to monitor all parameters of the second device using the second monitoring protocol when it is determined that the number of parameters collected by the second monitoring protocol in the second device meets a preset condition.

[0135] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes a third policy updating module and a fourth policy updating module.

[0136] The third policy updating module is configured to restore the first monitoring policy to the initial first monitoring policy for the first device after updating the first monitoring policy to a monitoring policy with reduced monitoring intensity based on a load condition indicating that the load has decreased.

[0137] The fourth policy updating module is configured to restore the second monitoring policy to the initial second monitoring policy for the second device after updating the second monitoring policy to a monitoring policy with increased monitoring intensity based on a load condition indicating that the load has decreased.

[0138] According to an embodiment of the present invention, the apparatus 400 for monitoring network performance further includes a second load change module.

[0139] The second load change module is configured to determine that the load condition indicates that the load has decreased when the policy switching evaluation value is lower than a second switching threshold corresponding to the device type.

[0140] According to embodiments of the present invention, any multiple modules among the first device monitoring module 410, the second device monitoring module 420, the first policy update module 430, and the second policy update module 440 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the first device monitoring module 410, the second device monitoring module 420, the first policy update module 430, and the second policy update module 440 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the first device monitoring module 410 , the second device monitoring module 420 , the first policy updating module 430 , and the second policy updating module 440 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0141] Figure 5 A block diagram of an electronic device suitable for implementing a method for monitoring network performance according to an embodiment of the present invention is shown.

[0142] like Figure 5 As shown, an electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 502 or programs loaded from a storage unit 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0143] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 executes the programs in the ROM 502 and / or RAM 503 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and RAM 503. The processor 501 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0144] According to an embodiment of the present invention, electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. Electronic device 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or modem. Communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 510 as needed, so that computer programs read from the removable media can be installed into storage section 508 as needed.

[0145] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0146] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above, and / or one or more memories other than ROM 502 and RAM 503.

[0147] The embodiments of the present invention further include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present invention.

[0148] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 501. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0149] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 509, and / or installed from a removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0150] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0151] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0153] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0154] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A method for monitoring network performance, characterized in that: The method comprises: Monitoring a first device in a target network using a first monitoring strategy to determine a load status of the first device; monitoring a second device in the target network that is different from the first device, using a second monitoring policy to determine a load status of the second device, wherein a monitoring intensity of the second monitoring policy is lower than a monitoring intensity of the first monitoring policy, the second device is a non-core device, and the second monitoring policy is used to monitor parameters of the second device using the first monitoring protocol; For the first device, based on its load status indicating that the load has increased, updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity; and For the second device, based on a load condition indicating that the load has increased, updating the second monitoring strategy to a monitoring strategy with increased monitoring intensity; The updating of the second monitoring strategy to a monitoring strategy with increased monitoring intensity based on the load condition indicating that the load has increased includes: Increasing the sampling frequency of the second monitoring strategy; When it is determined that a parameter mutation occurs in the second device, the target parameter with the parameter mutation is monitored using the second monitoring protocol, and the monitoring of the target parameter by the first monitoring protocol is stopped.

2. The method according to claim 1, characterized in that The method further comprises: For each of the first device and the second device: Determining a policy switching evaluation value based on the load condition, the device type of the device, and the duration of the monitoring policy corresponding to the device; and In a case where the policy switching evaluation value is higher than a first switching threshold corresponding to the device type, it is determined that the load condition indicates that the load is increasing.

3. The method according to claim 2, characterized in that The method further comprises: Determine the load status of the device based on one or more of the resource utilization rate and congestion index of the device, wherein the resource utilization rate includes the port sending bandwidth utilization rate, receiving bandwidth utilization rate, processor utilization rate, etc. of the device. The congestion indicator includes at least one of a port queue buffer level and a network congestion frame count.

4. The method according to claim 3, characterized in that Determining the strategy switching evaluation value includes: determining a handover weight corresponding to the device type; Determining a time decay term according to the duration; and The strategy switching evaluation value is obtained according to the switching weight, the time decay term and the load condition.

5. The method according to claim 2, characterized in that The device type includes a core device and a non-core device, and the method further includes: An interconnection topology of devices in the target network, including the first device and the second device, is constructed based on a device topology management protocol, and a device type of the device is determined according to the interconnection topology.

6. The method according to claim 5, characterized in that The first device is a core device, and updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity based on a load status indication that the load has increased comprises: Obtaining a peer port of a target port with relatively low traffic through the interconnection topology; The real-time traffic monitoring of the target port is closed, and at least one of the sending traffic and the receiving traffic of the opposite port is used to convert at least one of the sending traffic and the receiving traffic of the target port.

7. The method according to claim 1, characterized in that The method further comprises: When it is determined that the number of parameters collected by the second monitoring protocol in the second device meets a preset condition, all parameters of the second device are monitored using the second monitoring protocol.

8. The method according to claim 7, characterized in that The preset conditions include: The number of parameters reaches a preset percentage of all parameters of the second device.

9. The method according to claim 7 or 8, characterized in that The first monitoring protocol includes a network flow protocol and a sampling flow protocol, which are used to periodically obtain the parameters of the device through active polling to achieve monitoring of the device. The second monitoring protocol includes a telemetry protocol, which is used to obtain the parameters of the device through subscription to achieve monitoring of the device.

10. The method according to claim 1, characterized in that The method further comprises: For the first device, after updating the first monitoring strategy to a monitoring strategy with reduced monitoring intensity, based on its load status indicating that the load has become smaller, restoring the first monitoring strategy to the initial first monitoring strategy; and For the second device, after the second monitoring strategy is updated to a monitoring strategy with increased monitoring intensity, based on its load status indicating that the load becomes smaller, the second monitoring strategy is restored to the initial second monitoring strategy.

11. The method according to claim 2, characterized in that The method further comprises: In a case where the policy switching evaluation value is lower than a second switching threshold corresponding to the device type, it is determined that the load condition indicates that the load has become smaller.

12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 11.

13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • VoLTE system monitoring method and device, storage medium and electronic equipment

    CN116599873A