A high-performance intelligent management method and system for network congestion self-optimization
Patent Information
- Application Number
- CN202610457921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-04-09
AI Technical Summary
该方法存在技术局限性:一方面人工操作链路长、决策周期久,响应时效性不足,难以应对突发流量引发的瞬时拥堵,易造成数据传输时延增加、丢包率上升,另一方面缺乏终端设备精准识别与分类管理机制,无法针对工业控制终端、办公终端等不同业务属性设备制定差异化调度方案,导致资源分配失衡
[0015]本发明通过接收高性能智能管理指令,根据高性能智能管理指令获取网络拓扑图,其中,网络拓扑图包括多个节点,本发明通过网络拓扑图完整掌握全网节点分布及连接关系,为后续关键节点识别、拥堵定位提供可视化的基础数据支撑,从网络拓扑图识别出多个关键路径节点,对多个关键路径节点中的每个关键路径节点均执行如下操作:对关键路径节点进行采集模块配置,得到已配置节点,基于已配置节点获取节点网络参数及节点关联影响特征值,本发明通过采集模块配置使节点具备数据采集能力,精准获取节点网络参数与关联影响特征值,为拥堵根因判定提供量化、精准的指标依据,解决传统管理中因数据缺失导致的拥堵定位不准问题,根据节点网络参数及节点关联影响特征值确认出根因拥堵节点、已恢复节点或正常节点,基于量化指标实现节点状态的精准分类,能够快速区分需要优化的拥堵节点、已自愈的恢复节点及无需干预的正常节点,提升拥堵管理的针对性,避免对正常节点的无效操作,节省管理资源,根据根因拥堵节点确认出根因拥堵设备类型,其中,根因拥堵设备类型为核心路由器或网络交换设备,若根因拥堵设备类型为核心路由器,则将根因拥堵节点作为路径优化节点,对路径优化节点进行路由路径切换优化操作,得到已优化节点,若根因拥堵设备类型为网络交换设备,则将根因拥堵节点作为待负载均衡节点,对待负载均衡节点进行负载均衡优化操作,得到已优化节点,本发明对核心路由器采取路径切换优化,可快速规避拥堵链路,保障数据传输的连续性与低时延,对网络交换设备采取负载均衡优化,能均衡设备负载分布,提升资源利用率,通过两种差异化策略实现拥堵问题的精准消解,提升优化效率与效果,分别汇总已优化节点、已恢复节点及正常节点,得到已优化节点集、已恢复节点集及正常节点集,基于已优化节点集、已恢复节点集及正常节点集完成网络拥堵自优化的高性能智能管理。因此,本发明可全面提升网络运行的稳定性与资源利用效率。
Smart Images

Figure CN122001804B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network self-optimization management technology, and in particular to a high-performance intelligent management method and system for network congestion self-optimization. Background Technology
[0002] Network congestion self-optimization is an automated network management process that automatically detects network congestion and uses a series of intelligent algorithms and strategies to automatically adjust the allocation of network resources and path selection to alleviate congestion and improve network transmission efficiency and performance. High-performance intelligent management is an advanced network management approach that utilizes high-performance computing power and intelligent algorithm models to perform comprehensive and automated network management.
[0003] Traditional network congestion management typically employs a combination of manual monitoring and intervention. Administrators rely on monitoring tools to manually collect network-wide traffic data, analyze traffic characteristics to locate congested nodes, and then manually adjust core configurations such as routing policies and bandwidth allocation rules. This method has technical limitations: firstly, manual operation involves long processing times and decision-making cycles, resulting in insufficient response time and difficulty in handling instantaneous congestion caused by sudden traffic surges, easily leading to increased data transmission latency and packet loss rates; secondly, it lacks a precise identification and classification management mechanism for terminal devices, making it impossible to develop differentiated scheduling schemes for devices with different business attributes, such as industrial control terminals and office terminals, leading to resource imbalances. Therefore, how to comprehensively improve network operation stability and resource utilization efficiency is an urgent technical problem to be solved. Summary of the Invention
[0004] This invention provides a high-performance intelligent management method and system for network congestion self-optimization, the main purpose of which is to comprehensively improve the stability of network operation and the efficiency of resource utilization.
[0005] To achieve the above objectives, the present invention provides a high-performance intelligent management method for network congestion self-optimization, comprising: Receive high-performance intelligent management instructions and obtain a network topology map based on the instructions. The network topology map includes multiple nodes. Multiple critical path nodes were identified from the network topology graph. For each critical path node, the following operations were performed: The acquisition module is configured for critical path nodes to obtain the configured nodes. Based on the configured nodes, the node network parameters and node association impact feature values are obtained. Based on the node network parameters and the characteristic values of node associations, the root cause congestion nodes, recovered nodes, or normal nodes can be identified. The type of device causing the congestion is determined based on the root cause congestion node. The root cause congestion device type is either a core router or a network switching device. If the root cause congestion device type is a core router, then the root cause congestion node is used as a path optimization node, and a route path switching optimization operation is performed on the path optimization node to obtain the optimized node. If the root cause congestion device type is a network switching device, then the root cause congestion node is regarded as a node to be load balanced, and load balancing optimization operation is performed on the node to be load balanced to obtain the optimized node. The optimized nodes, recovered nodes, and normal nodes are summarized separately to obtain the optimized node set, the recovered node set, and the normal node set; High-performance intelligent management that achieves self-optimization of network congestion based on optimized node sets, recovered node sets, and normal node sets.
[0006] Optionally, the step of obtaining node network parameters and node association influence feature values based on configured nodes includes: Obtain historical byte count and maximum interface bandwidth, collect data from the configured nodes using a preset collection interval, obtain the current byte count, and calculate the node bandwidth utilization rate based on the historical byte count, current byte count, and maximum interface bandwidth. Once the target and probe packet are identified, multiple data packets are sent to the target using the configured nodes and probe packets to obtain the number of transmissions, the number of receptions, and the average node latency. The node packet loss rate is calculated based on the number of transmissions and receptions. The node association impact characteristic value is calculated based on the configured nodes. The node network parameters are determined based on the node bandwidth utilization, the node average latency, and the node packet loss rate.
[0007] Optionally, the formula for calculating the node bandwidth utilization is as follows: in, Indicates node bandwidth utilization. Indicates the current number of bytes. Indicates the number of historical bytes. Indicates the data collection interval. This indicates the maximum bandwidth of the interface.
[0008] Optionally, the step of calculating the node association influence feature value based on the configured nodes includes: Based on the configured nodes in the network topology diagram, the associated target node set is identified, and the associated target node set is filtered to obtain the downstream directly connected node set; The downstream directly connected device set is determined based on the downstream directly connected node set. The downstream directly connected devices are extracted sequentially from the downstream directly connected device set. Network access parameters and device operating parameters are obtained based on the extracted downstream directly connected devices. Based on network access parameters and device operating parameters, anomalies are judged on the extracted downstream directly connected devices to obtain performance judgment results, which are either abnormal or normal. If the performance assessment result is abnormal, then the extracted downstream directly connected devices will be regarded as devices with abnormal performance. Sum the devices with abnormal performance to obtain the set of devices with abnormal performance. Then sum the set of devices with abnormal performance and the set of downstream directly connected devices to obtain the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is calculated based on the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is the value obtained by dividing the number of devices with abnormal performance by the number of downstream directly connected devices.
[0009] Optionally, the step of identifying the root cause congestion node, recovered node, or normal node based on node network parameters and node association impact characteristic values includes: If the node network parameters do not meet the preset standard node network parameter range conditions, the configured node is regarded as a suspected congested node, a first proportion threshold is obtained, and it is determined whether the node association influence feature value of the suspected congested node is greater than the first proportion threshold. If the node association impact characteristic value of a suspected congested node is greater than the first proportional threshold, then the suspected congested node is regarded as the root cause congestion node. If the node association impact feature value of a suspected congested node is not greater than the first proportion threshold, then the suspected congested node is regarded as an affected node, and a self-optimization operation is performed on the affected node to obtain the recovered node. If the node network parameters meet the standard node network parameter range conditions, then the configured node is considered a normal node.
[0010] Optionally, obtaining the first ratio threshold includes: Obtain the historical alarm root cause device set, which includes multiple historical alarm root cause devices, and each historical alarm root cause device corresponds to an alarm trigger timestamp. Historical alarm root cause devices are extracted sequentially from the historical alarm root cause device set. A time window is obtained based on the alarm trigger timestamp corresponding to the extracted historical alarm root cause device. The set of directly connected devices is identified based on the time window. The proportion of abnormal downstream devices in history is calculated based on the set of directly connected devices. The historical downstream equipment abnormality ratios are summarized to obtain a historical downstream equipment abnormality ratio set. The median of the historical downstream equipment abnormality ratio set is extracted to obtain the median abnormality ratio, which is then used as the first ratio threshold.
[0011] Optionally, the step of performing route path switching optimization operation on the path optimization node to obtain the optimized node includes: The alternative route path set is retrieved from the pre-built routing information database based on the path optimization node, and alternative route paths are extracted from the alternative route path set in turn. The path node set is confirmed based on the extracted alternative route paths, wherein the path node set includes one or more path nodes. Calculate the path node bandwidth utilization and path node latency value for each path node in the path node set to obtain the path node bandwidth utilization set and path node latency value set. The maximum value is extracted from the set of bandwidth utilization rates of path nodes to obtain the maximum bandwidth utilization rate of the path node. The total path delay value is obtained by summing the set of delay values of path nodes. The evaluation value of the backup path is calculated based on the maximum path node bandwidth utilization and the total path delay value. Summarize the evaluation values of the alternative paths to obtain the set of alternative path evaluation values, and take the alternative route path corresponding to the alternative path with the largest evaluation value in the set of alternative path evaluation values as the optimal alternative path. Based on the optimal backup path, the routing path of the path optimization node is switched to obtain the optimized node.
[0012] Optionally, the step of performing load balancing optimization on the node to be load balanced to obtain an optimized node includes: Based on the nodes to be load balanced, identify the high-load source devices, obtain multiple terminal devices and multiple similar devices based on the high-load source devices, perform minimum load rate search on the multiple similar devices, and obtain candidate target devices; Identify a set of non-critical terminal devices from multiple terminal devices and count the number of non-critical terminal devices in the set; Based on the high load source devices and candidate target devices, obtain the high load rate and target device load rate, and calculate the theoretical number of devices to be migrated based on the number of non-critical terminal devices, the high load rate and the target device load rate. Obtain the maximum physical number of connections and the current number of connections for the candidate target device, and calculate the remaining number of connections based on the maximum physical number of connections and the current number of connections; The theoretical number of devices to be migrated is adjusted based on the remaining number of connections and the preset migration security ratio to obtain the final number of devices to be migrated. Based on the final number of devices to be migrated, a set of devices to be migrated is matched from the set of non-critical terminal devices. The set of devices to be migrated is then migrated to the candidate target devices to obtain optimized nodes.
[0013] Optionally, the formula for calculating the theoretical number of relocated devices is as follows: in, Indicates the theoretical number of devices to be migrated. Indicates a high load rate. Indicates the target device load rate. This indicates the number of non-critical terminal devices. The symbol indicates rounding up.
[0014] To achieve the above objectives, the present invention also provides a high-performance intelligent management system for network congestion self-optimization, comprising: The critical path node identification module is used to receive high-performance intelligent management instructions, obtain a network topology map based on the instructions, wherein the network topology map includes multiple nodes, and identify multiple critical path nodes from the network topology map; The node status confirmation module is used to perform the following operations on each critical path node among multiple critical path nodes: configure the acquisition module of the critical path node to obtain the configured node, obtain the node network parameters and node association impact feature values based on the configured node, and confirm the root cause congestion node, recovered node or normal node based on the node network parameters and node association impact feature values. The root cause node optimization module is used to identify the root cause congestion device type based on the root cause congestion node. The root cause congestion device type includes core routers or network switching devices. If the root cause congestion device type is a core router, the root cause congestion node is used as a path optimization node, and a routing path switching optimization operation is performed on the path optimization node to obtain an optimized node. If the root cause congestion device type is a network switching device, the root cause congestion node is used as a load balancing node, and a load balancing optimization operation is performed on the load balancing node to obtain an optimized node. The network self-optimization module is used to summarize optimized nodes, recovered nodes, and normal nodes respectively, to obtain optimized node sets, recovered node sets, and normal node sets. Based on the optimized node sets, recovered node sets, and normal node sets, it completes high-performance intelligent management of network congestion self-optimization.
[0015] This invention receives high-performance intelligent management commands and obtains a network topology map based on these commands. The network topology map includes multiple nodes. This invention provides a complete understanding of the distribution and connections of all nodes in the network, offering visualized data support for subsequent key node identification and congestion localization. Multiple key path nodes are identified from the network topology map. For each key path node, the following operations are performed: A data acquisition module is configured to obtain the configured node; based on the configured node, node network parameters and node-related influence characteristic values are obtained. This invention enables nodes to acquire data through the data acquisition module configuration, accurately obtaining node network parameters and related influence characteristic values. This provides quantitative and accurate indicators for determining the root cause of congestion, solving the problem of inaccurate congestion localization caused by data gaps in traditional management. Based on the node network parameters and node-related influence characteristic values, the root cause congested nodes, recovered nodes, or normal nodes are identified. Based on quantitative indicators, precise classification of node status is achieved, quickly distinguishing between congested nodes requiring optimization, self-healed recovered nodes, and normal nodes requiring no intervention, thus improving congestion management. This targeted approach avoids ineffective operations on normal nodes, saving management resources. Based on the root cause congestion node, the type of the root cause congestion device is identified. If the root cause congestion device type is a core router or a network switch, and it is a core router, then the root cause congestion node is designated as a path optimization node, and routing path switching optimization is performed on it to obtain an optimized node. If the root cause congestion device type is a network switch, then it is designated as a node to be load balanced, and load balancing optimization is performed on it to obtain an optimized node. This invention optimizes path switching for core routers, quickly avoiding congested links and ensuring continuous and low-latency data transmission. It also optimizes load balancing for network switching equipment, balancing device load distribution and improving resource utilization. Through these two differentiated strategies, congestion problems are precisely resolved, improving optimization efficiency and effectiveness. Optimized nodes, recovered nodes, and normal nodes are aggregated to obtain optimized node sets, recovered node sets, and normal node sets, respectively. Based on these sets, high-performance intelligent management of network congestion self-optimization is achieved. Therefore, this invention can comprehensively improve network operational stability and resource utilization efficiency. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a high-performance intelligent management method for network congestion self-optimization provided in an embodiment of the present invention. Figure 2 This is a functional block diagram of a high-performance intelligent management system for network congestion self-optimization provided in an embodiment of the present invention.
[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a high-performance intelligent management method for self-optimization of network congestion. The executing entity of this high-performance intelligent management method for self-optimization of network congestion includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the high-performance intelligent management method for self-optimization of network congestion can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a high-performance intelligent management method for network congestion self-optimization according to an embodiment of the present invention. In this embodiment, the high-performance intelligent management method for network congestion self-optimization includes: S1. Receive high-performance intelligent management instructions and obtain a network topology map based on the instructions. The network topology map includes multiple nodes.
[0021] It should be explained that high-performance intelligent management commands are commands that trigger the entire network intelligent management process. A network topology diagram is a topology diagram describing all devices in the current network and the connections between them. Nodes are the constituent units of the network topology diagram. Nodes correspond to various hardware devices or logical functional units in the network (such as router nodes, switch nodes, industrial control terminal nodes, etc.), and each node has independent network parameters (such as bandwidth utilization, transmission latency).
[0022] S2. Identify multiple critical path nodes from the network topology diagram, and perform the following operations on each of the multiple critical path nodes.
[0023] It should be explained that critical path nodes are nodes identified from the network topology diagram that play a decisive role in the transmission of core network services. For example, critical path nodes are core router nodes responsible for command transmission and server nodes carrying critical data interactions in an industrial control network. This invention clearly distinguishes between critical path nodes and non-critical path nodes in the network topology. Their impact on network services differs: the operational status of critical path nodes directly affects the continuity, real-time performance, and reliability of core service traffic transmission. Once a fault occurs (such as congestion or downtime), it will lead to the interruption or severe performance degradation of core services (such as the issuance of industrial control commands and critical data interactions). Non-critical path nodes, on the other hand, only carry non-core services (such as ordinary office data and non-real-time monitoring information). Even if a non-critical path node fails, it only affects localized secondary services and will not have a decisive impact on the overall core business process.
[0024] S3. Configure the acquisition module for critical path nodes to obtain the configured nodes, and obtain the node network parameters and node association influence feature values based on the configured nodes.
[0025] It should be explained that configuring the data acquisition module for critical path nodes involves enabling dedicated data acquisition plugins (such as core server acquisition plugins or industrial PLC control node acquisition plugins) for critical path nodes and configuring the acquisition parameters (such as acquisition cycle and acquisition indicator type) of the data acquisition plugins. The configured node is the critical path node after the data acquisition module configuration is completed.
[0026] Specifically, the step of obtaining node network parameters and node association influence feature values based on configured nodes includes: Obtain historical byte count and maximum interface bandwidth, collect data from the configured nodes using a preset collection interval, obtain the current byte count, and calculate the node bandwidth utilization rate based on the historical byte count, current byte count, and maximum interface bandwidth. Once the target and probe packet are identified, multiple data packets are sent to the target using the configured nodes and probe packets to obtain the number of transmissions, the number of receptions, and the average node latency. The node packet loss rate is calculated based on the number of transmissions and receptions. The node association impact characteristic value is calculated based on the configured nodes. The node network parameters are determined based on the node bandwidth utilization, the node average latency, and the node packet loss rate.
[0027] It should be explained that the historical byte count is the cumulative traffic value collected from the network device interface at the last collection moment. The maximum interface bandwidth is the theoretical maximum data transmission rate of the network interface. Optionally, the maximum interface bandwidth is 100Mbps, 1000Mbps, or 10Gbps. The collection interval is the time interval between two consecutive data collection operations. For example, the collection interval is 5 seconds. The current byte count is the cumulative traffic value read in real-time from the network device interface within the collection interval, since the device corresponding to the configured node started. Node bandwidth utilization is the percentage of the average bandwidth actually used by the network device interface relative to the interface's maximum bandwidth within the collection interval. The sending destination is the destination network address to which the probe packet is sent. For example, the sending destination is a public DNS server or default gateway. A probe packet is a small data packet actively sent to measure network performance. For example, a probe packet is an ICMP Echo Request packet or a UDP packet.
[0028] Understandably, the step of sending multiple data packets to the target using configured nodes and probe packets involves: using configured nodes to send probe packets to the target at fixed time intervals (e.g., one every 2 seconds), waiting for reply packets from the target, and repeating this process multiple times until the number of transmissions equals a pre-set maximum (e.g., 10 times). The number of transmissions is the total number of probe packets sent by the configured nodes to the target within a statistical period. The number of receptions is the total number of reply packets successfully received by the configured nodes for the corresponding probe packets within the same statistical period. The average node latency is the average round-trip time of all successfully received probe packets within a statistical period. For example, if 30 probe packets are sent within 1 minute and 28 reply packets are successfully received, the latency values of these 28 reply packets are added together and divided by 28 to obtain the average node latency. The node data packet loss rate is the percentage of lost probe packets out of the total number of transmissions within the statistical period. Detailed steps for calculating the node association influence characteristic value based on the configured nodes will be given later. The node network parameter is a set of parameters that includes node bandwidth utilization, average node latency, and node packet loss rate. The formula for calculating the node packet loss rate is as follows: in, Indicates the node packet loss rate. Indicates the number of times it was sent. Indicates the number of times the message was received.
[0029] In detail, the formula for calculating the node bandwidth utilization is as follows: in, Indicates node bandwidth utilization. Indicates the current number of bytes. Indicates the number of historical bytes. Indicates the data collection interval. This indicates the maximum bandwidth of the interface.
[0030] It should be explained that in the calculation of node bandwidth utilization, firstly based on The actual data transmission increment within the acquisition interval is calculated. Since the standard unit of measurement for network bandwidth is bits, while node data transmission is usually measured in bytes, therefore, the actual data transmission increment within the acquisition interval is calculated using... After completing the unit conversion, the byte increment is converted into a bit increment, which is then divided by the collection interval to obtain the average transmission rate within the collection interval. Finally, the average transmission rate is divided by the maximum bandwidth of the node interface and multiplied by 100% to transform the abstract bandwidth usage into a directly quantifiable node bandwidth utilization rate expressed as a percentage. The purpose of introducing the above formula in this invention is to accurately quantify the bandwidth resource occupancy of configured nodes (such as router and switch ports) within the collection interval, providing data support for subsequent network congestion judgment, resource allocation optimization, and device load monitoring.
[0031] Specifically, the calculation of node association influence feature values based on configured nodes includes: Based on the configured nodes in the network topology diagram, the associated target node set is identified, and the associated target node set is filtered to obtain the downstream directly connected node set; The downstream directly connected device set is determined based on the downstream directly connected node set. The downstream directly connected devices are extracted sequentially from the downstream directly connected device set. Network access parameters and device operating parameters are obtained based on the extracted downstream directly connected devices. Based on network access parameters and device operating parameters, anomalies are judged on the extracted downstream directly connected devices to obtain performance judgment results, which are either abnormal or normal. If the performance assessment result is abnormal, then the extracted downstream directly connected devices will be regarded as devices with abnormal performance. Sum the devices with abnormal performance to obtain the set of devices with abnormal performance. Then sum the set of devices with abnormal performance and the set of downstream directly connected devices to obtain the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is calculated based on the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is the value obtained by dividing the number of devices with abnormal performance by the number of downstream directly connected devices.
[0032] It should be explained that the associated target node set is a collection of all associated target nodes. Associated target nodes are the set of all nodes in the network topology that have a connection relationship with the currently analyzed node (the configured node). For example, associated target nodes include sibling nodes (other devices under the same switch) and downstream nodes (such as connected sub-switches and terminal devices). Filtering the associated target node set involves retaining all downstream nodes in the associated target node set that have a direct physical or logical link connection with the current node. The downstream directly connected node set is a collection of all downstream directly connected nodes. Downstream directly connected nodes are downstream nodes that have a direct physical or logical link connection with the current node. The downstream directly connected device set is a collection of all downstream directly connected devices. Downstream directly connected devices are the devices corresponding to downstream directly connected nodes. Identification information is information used to uniquely identify and locate a network device. For example, identification information includes IP address, device name, MAC address, etc. The identification information set is a collection of all identification information.
[0033] Understandably, network access parameters include: device real-time bandwidth utilization, device transmission latency, and packet loss rate. Device operating parameters include: device CPU utilization, memory utilization, and number of device disconnections. Device real-time bandwidth utilization is the real-time bandwidth utilization of the downstream directly connected device's own network interface. Device transmission latency and packet loss rate are the network round-trip latency from the monitoring system to the downstream directly connected device and the proportion of packet loss during communication between the monitoring system and the downstream directly connected device, respectively. Device CPU utilization is the CPU utilization of the downstream directly connected device. Memory utilization is the memory (RAM) utilization of the downstream directly connected device. Number of device disconnections is the number of times the device loses connection with the monitoring system within a certain time window. The step of determining the performance of extracted downstream directly connected devices based on network access parameters and device operating parameters is as follows: If the device's real-time bandwidth utilization (e.g., 70%), transmission latency (e.g., 50ms), and packet loss rate (e.g., 1%) are greater than the preset standard, or if the device's CPU utilization (e.g., 80%), memory utilization (e.g., 85%), or number of disconnections is greater than the preset standard (e.g., 3 times / hour), then the downstream directly connected device is determined to have abnormal performance. Conversely, if the conditions are not met, the downstream directly connected device is determined to have normal performance. The performance abnormality or normality is then confirmed as the performance judgment result. An abnormal performance device is a downstream directly connected device whose performance judgment result is abnormal. If the performance judgment result is abnormal, then the extracted downstream directly connected devices are considered abnormal performance devices. Normal performance devices are downstream directly connected devices whose performance judgment result is normal performance. The abnormal performance device set is a collection of all abnormal performance devices. The number of devices with performance issues and the number of downstream directly connected devices are respectively the number of devices with performance issues in the performance issue cluster and the number of downstream directly connected devices in the downstream directly connected device cluster.
[0034] Importantly, in traditional network congestion self-optimization management, relying solely on the network parameters of a node itself (such as bandwidth utilization and transmission latency) to determine the root cause of congestion has significant limitations. Such single parameters can only reflect the node's own operating status and cannot reflect the chain effect of node anomalies on downstream devices. It is easy to mistakenly identify nodes with abnormal parameters but no downstream impact as root cause nodes, or to overlook the problem that their own parameters are within limits but cause a large number of downstream devices to malfunction. Therefore, this invention introduces node correlation impact feature values to solve this problem.
[0035] S4. Identify the root cause congestion nodes, recovered nodes, or normal nodes based on the node network parameters and the characteristic values of node association impact.
[0036] Specifically, the step of identifying the root cause congestion node, recovered node, or normal node based on node network parameters and node association impact characteristic values includes: If the node network parameters do not meet the preset standard node network parameter range conditions, the configured node is regarded as a suspected congested node, a first proportion threshold is obtained, and it is determined whether the node association influence feature value of the suspected congested node is greater than the first proportion threshold. If the node association impact characteristic value of a suspected congested node is greater than the first proportional threshold, then the suspected congested node is regarded as the root cause congestion node. If the node association impact feature value of a suspected congested node is not greater than the first proportion threshold, then the suspected congested node is regarded as an affected node, and a self-optimization operation is performed on the affected node to obtain the recovered node. If the node network parameters meet the standard node network parameter range conditions, then the configured node is considered a normal node.
[0037] It should be explained that the standard node network parameter range conditions include the standard bandwidth utilization range, the standard transmission delay range, and the standard packet loss rate range, and the standard node network parameter range corresponds one-to-one with the parameters in the node network parameters. The standard bandwidth utilization range, the standard transmission delay range, and the standard packet loss rate range are all pre-set normal ranges used to determine whether a node's network status is healthy. In this invention, a node's network parameters not meeting the preset standard node network parameter range conditions means that the node's bandwidth utilization is not within the standard bandwidth utilization range, the node's average latency value is not within the standard transmission delay range, and the node's packet loss rate is not within the standard packet loss rate range. A suspected congested node is a configured node whose network parameters do not meet the preset standard node network parameter range conditions. If the node association impact characteristic value of a suspected congested node is greater than the first proportional threshold, it indicates that the performance problem of the suspected congested node has affected multiple downstream networks of the suspected congested node. A root cause congested node is a network device (suspected congested node) whose own performance is abnormal, and whose abnormality is the root cause of the performance degradation of a network area. In this invention, node network parameters meeting the standard node network parameter range conditions are defined as follows: node bandwidth utilization is within the standard bandwidth utilization range, node average latency is within the standard transmission latency range, and node packet loss rate is within the standard packet loss rate range. A normal node is a configured node whose network parameters meet the standard node network parameter range conditions. The self-optimization operation for affected nodes prioritizes ensuring the bandwidth requirements of critical business data (such as file transfers and video conferencing in enterprise offices, and online classes and video calls on home networks) for affected nodes, while temporarily reducing the bandwidth share of non-critical businesses (such as downloads and game updates). For example, when congestion is detected affecting video conferencing data transmission in an enterprise network, the bandwidth share of that meeting is automatically increased from 20% to 40%. A recovered node is a node obtained after performing self-optimization operations on the affected node.
[0038] Furthermore, this invention employs a two-layer determination mechanism to pinpoint the root cause of network congestion: First, a preliminary screening is performed based on a standard node network parameter range, classifying configured nodes that do not meet the conditions of this range as suspected congested nodes, while nodes that meet the conditions are directly identified as normal nodes. Further, a first proportional threshold is introduced for suspected congested nodes. By verifying whether the node association influence characteristic value of a suspected congested node exceeds this threshold, the final determination is made as to whether the suspected congested node is the root cause of congestion. This two-layer determination mechanism effectively avoids the problems of misclassifying normal nodes as congested nodes or missing true root cause congested nodes that are easily caused by relying solely on a single parameter, thereby improving the accuracy of congestion node identification and providing a precise basis for subsequent targeted network congestion optimization strategies, ensuring network management efficiency and stability.
[0039] Specifically, obtaining the first ratio threshold includes: Obtain the historical alarm root cause device set, which includes multiple historical alarm root cause devices, and each historical alarm root cause device corresponds to an alarm trigger timestamp. Historical alarm root cause devices are extracted sequentially from the historical alarm root cause device set. A time window is obtained based on the alarm trigger timestamp corresponding to the extracted historical alarm root cause device. The set of directly connected devices is identified based on the time window. The proportion of abnormal downstream devices in history is calculated based on the set of directly connected devices. The historical downstream equipment anomaly ratios are summarized to obtain a historical downstream equipment anomaly ratio set. The median of the historical downstream equipment anomaly ratio set is extracted to obtain the median anomaly ratio, which is then used as the first ratio threshold.
[0040] It should be explained that the historical alarm root cause device is the network device that was ultimately identified as the root cause of the problem in a congestion event that has been successfully verified by the system in the past. The alarm trigger timestamp is the time when the historical alarm root cause device was ultimately identified as the root cause of the problem. The step of obtaining the time window based on the alarm trigger timestamp corresponding to the extracted historical alarm root cause device is as follows: according to a preset backtracking time window (e.g., 2 minutes), the interval formed by subtracting the backtracking time window from the alarm trigger timestamp is used as the time window. The data within the backtracking time window best reflects the direct state before the problem occurred. Therefore, this invention identifies the set of directly connected devices based on the time window, avoiding the interference of chaotic data analysis after the problem occurs. The step of identifying the set of directly connected devices based on the time window involves obtaining all downstream devices of the extracted historical alarm root cause device within the time window and taking all downstream devices as the set of directly connected devices. The method of calculating the abnormal proportion of historical downstream devices based on the set of directly connected devices is the same as the method of calculating the node association impact feature value based on the configured nodes, and will not be repeated here. The historical downstream device anomaly ratio is the percentage of devices with performance abnormalities among all directly downstream devices corresponding to the root cause device of a historical alarm. The historical downstream device anomaly ratio set is a collection of all historical downstream device anomaly ratios. The median anomaly ratio is the median of the historical downstream device anomaly ratio set. The method for extracting the median is existing technology and will not be elaborated here.
[0041] It should be noted that this invention uses historical alarm root cause device set—a set of real historical congestion data—as a benchmark to generate the first proportional threshold. This avoids the subjectivity of manually setting thresholds and leverages the median's characteristic of being less susceptible to interference from extreme abnormal proportional values, ensuring that the first proportional threshold accurately reflects the actual impact of root cause devices on downstream devices in most congestion scenarios. Compared to the shortcomings of the mean, which is easily skewed by extreme values, and the mode, which fails to reflect general trends, this invention uses the median as the first proportional threshold. This effectively eliminates the threshold deviation caused by individual extreme cases, thus providing a stable and quantitative basis that fits the actual network operation scenario for the accurate determination of current root cause congestion nodes.
[0042] S5. Identify the root cause congestion device type based on the root cause congestion node. The root cause congestion device type is a core router or network switching device.
[0043] It should be explained that core routers are critical network devices located in the network backbone layer. They primarily handle data forwarding and routing decisions between different network segments, subnets, or wide area networks. Their deployment location is typically at the core hub of the network architecture, responsible for processing high-speed data transmission and routing information exchange across the entire network. Network switching devices are network devices used to connect multiple terminal devices and nodes within a local area network.
[0044] S6. If the root cause congestion device type is a core router, then the root cause congestion node is used as a path optimization node. The path optimization node is subjected to route path switching optimization operation to obtain the optimized node.
[0045] In detail, the step of performing route path switching optimization operations on the path optimization node to obtain the optimized node includes: The alternative route path set is retrieved from the pre-built routing information database based on the path optimization node, and alternative route paths are extracted from the alternative route path set in turn. The path node set is confirmed based on the extracted alternative route paths, wherein the path node set includes one or more path nodes. Calculate the path node bandwidth utilization and path node latency value for each path node in the path node set to obtain the path node bandwidth utilization set and path node latency value set. The maximum value is extracted from the set of bandwidth utilization rates of path nodes to obtain the maximum bandwidth utilization rate of the path node. The total path delay value is obtained by summing the set of delay values of path nodes. The evaluation value of the backup path is calculated based on the maximum path node bandwidth utilization and the total path delay value. Summarize the evaluation values of the alternative paths to obtain the set of alternative path evaluation values, and take the alternative route path corresponding to the alternative path with the largest evaluation value in the set of alternative path evaluation values as the optimal alternative path. Based on the optimal backup path, routing path switching is performed on the optimized nodes to obtain the optimized nodes. It should be explained that in the process of calculating the backup path evaluation value: first, the bandwidth utilization of the maximum path node and the total path delay value are normalized, and then the backup path evaluation value is calculated using the following formula: [Formula omitted for brevity] in, This represents the evaluation value of the alternative path. This indicates the preset bandwidth utilization weight. This indicates the maximum bandwidth utilization of the path node. This indicates the preset path delay weight. This represents the total path delay value.
[0046] Importantly, the purpose of the calculation formula for the backup path evaluation value in this invention is to quantitatively evaluate the overall performance of backup routing paths in order to select the optimal backup path. To eliminate the dimensional difference between the maximum path node bandwidth utilization and the total path delay value, and to ensure that the weight allocation of the two evaluation dimensions is more practically meaningful, this invention uses a minimum-maximum normalization method to normalize the maximum path node bandwidth utilization and the total path delay value before calculating the backup path evaluation value, mapping the values of both to the [0, 1] interval. The minimum-maximum normalization method is existing technology and will not be elaborated upon here. The calculation formula uses... Because The maximum path node bandwidth utilization rate in the path node set represents the bandwidth bottleneck of the backup routing path. Because the maximum path node bandwidth utilization rate is negatively correlated with path performance, therefore... The bandwidth utilization rate of the maximum path node can be converted into a positively correlated bandwidth availability metric. A higher metric indicates more remaining available bandwidth and better bandwidth performance for the node corresponding to the maximum path node utilization rate. Furthermore, the reciprocal of the total path delay value is used because total path delay is negatively correlated with path performance; taking the reciprocal yields... It can represent the path transmission efficiency of the alternative routing path, and A higher value indicates a lower total path delay and higher transmission efficiency. Furthermore, this reciprocal approach avoids excessive influence on the backup path evaluation value due to large differences in total path delay, ensuring a balanced evaluation weight for both bandwidth and delay dimensions. Ultimately, a higher backup path evaluation value indicates better overall performance in terms of bandwidth and delay for the corresponding backup route. Therefore, this invention achieves a scientific quantitative evaluation of backup routes, improving the efficiency and accuracy of route switching.
[0047] It should be explained that the routing information database is a pre-built database that stores information on all known alternative routing paths in the network that can reach different destinations. This alternative routing path information includes, but is not limited to, all nodes on the path. The backup routing path set is a collection of all backup routing paths. A backup routing path is a path retrieved from the routing information database that bypasses a currently congested path optimization node and the affected destination. The path node set is a collection of all path nodes. A path node is a node of a network device within a backup routing path. The method for calculating the path node bandwidth utilization and path node latency value for each path node in the path node set is the same as the method for obtaining node network parameters based on configured nodes, and will not be repeated here. The path node bandwidth utilization is the real-time bandwidth utilization of the critical forwarding interface of a path node on a backup path.
[0048] For example, the original congested path is: office area user → edge switch → core router A (path optimization node) → data center server. The alternative route is: office area user → edge switch → core router B → core router C → data center server. On the alternative route, the path node delay value of the core router B node is the delay from the edge switch to the core router B, and the path node delay value of the core router C node is the delay from the edge switch to the core router C.
[0049] It should also be explained that the path node bandwidth utilization set and the path node delay value set are sets composed of path node bandwidth utilization and path node delay values, respectively. The maximum path node bandwidth utilization is the path node bandwidth utilization with the highest value in the path node bandwidth utilization set. The total path delay value is the sum of the path node delay values. The backup path evaluation value is a numerical value used to comprehensively compare the advantages and disadvantages of different backup routing paths. The bandwidth utilization weight and path delay weight are preset values used to adjust the relative importance of bandwidth utilization and delay in the backup path evaluation value, respectively. It should be noted that the sum of the bandwidth utilization weight and the path delay weight is 1. The backup path evaluation value set is a set composed of all backup path evaluation values. The optimal backup path is the backup routing path corresponding to the backup path with the highest evaluation value in the backup path evaluation value set. The routing path switching based on the optimal backup path is: adding a static route configuration for service traffic that cannot be smoothly transmitted to the final destination network due to congestion at the path optimization node, and explicitly specifying the next hop of the static route as the entry point of the optimal backup path. The service traffic shown is various types of service data transmitted through the network. Examples include server-client requests, video streams, or file transfer data packets. Optimized nodes are those that have successfully undergone routing path switching optimization or load balancing optimization operations.
[0050] It should be noted that this invention selects the maximum bandwidth utilization from the set of path node bandwidth utilization rates because the worst-performing path node in the backup routing path becomes the bottleneck of the entire backup routing path. Summing the path delay values directly reflects the overall transmission time of the backup routing path. Combining the maximum bandwidth utilization rate and the total path delay value allows for a comprehensive evaluation of the backup routing path performance. Secondly, this invention introduces bandwidth utilization rate weights and path delay weights into the calculation formula for backup path evaluation values. This allows for flexible adjustment of the priority of bandwidth and delay in the path evaluation system according to actual network service requirements, making the backup path evaluation values more aligned with specific application scenarios. Finally, summarizing the backup path evaluation values and selecting the backup routing path corresponding to the maximum value as the optimal backup path for switching avoids the subjectivity and randomness of manual decision-making, ensuring that the selected path achieves the best overall performance in terms of bandwidth utilization and transmission delay. This achieves precise optimization of routing paths, improving network transmission efficiency and stability.
[0051] S7. If the root cause congestion device type is a network switching device, then the root cause congestion node is taken as the node to be load balanced, and load balancing optimization operation is performed on the node to be load balanced to obtain the optimized node.
[0052] In detail, the load balancing optimization operation performed on the node to be load balanced to obtain the optimized node includes: Based on the nodes to be load balanced, identify the high-load source devices, obtain multiple terminal devices and multiple similar devices based on the high-load source devices, perform minimum load rate search on the multiple similar devices, and obtain candidate target devices; Identify a set of non-critical terminal devices from multiple terminal devices and count the number of non-critical terminal devices in the set; Based on the high load source devices and candidate target devices, obtain the high load rate and target device load rate, and calculate the theoretical number of devices to be migrated based on the number of non-critical terminal devices, the high load rate and the target device load rate. Obtain the maximum physical number of connections and the current number of connections for the candidate target device, and calculate the remaining number of connections based on the maximum physical number of connections and the current number of connections; The theoretical number of devices to be migrated is adjusted based on the remaining number of connections and the preset migration security ratio to obtain the final number of devices to be migrated. Based on the final number of devices to be migrated, a set of devices to be migrated is matched from the set of non-critical terminal devices. The set of devices to be migrated is then migrated to the candidate target devices to obtain optimized nodes.
[0053] It is important to understand that the load balancing node is the root cause congestion node whose device type is a network switch. The high-load source device is the device identified as the root cause congestion node and whose device type is a network switch. For example, the network switch is an aggregation switch. Terminal devices are end-user devices directly connected to the high-load source device. For example, end-user devices are personal computers, printers, servers, etc. Similar devices are devices that are at the same layer in the network as the high-load source device, perform the same function, and are configured with the same service policies (such as the same VLAN). The step of searching for the minimum load rate among multiple similar devices to obtain candidate target devices is as follows: query the current load rate of each similar device among multiple similar devices, find the minimum load rate from all current load rates, and select the similar device corresponding to the minimum load rate as the candidate target device. The non-critical terminal device set is a collection of all non-critical terminal devices. Non-critical terminal devices include, but are not limited to: servers, network printers.
[0054] Understandably, the maximum physical number of connections and the current number of connections represent the total number of hardware ports on the candidate target device and the number of terminal devices currently connected to the candidate target device, respectively. The remaining number of connections is the value obtained by subtracting the current number of connections from the maximum physical number of connections. The step of correcting the theoretical number of migrated devices based on the remaining number of connections and a preset migration security ratio to obtain the final number of migrated devices is as follows: multiply the migration security ratio by the number of non-critical terminal devices, round up to obtain the security constraint value, and take the minimum value among the security constraint value, the theoretical number of migrated devices, and the remaining number of connections. The step of matching the set of devices to be migrated from the set of non-critical terminal devices based on the final number of migrated devices is as follows: sort the current network traffic corresponding to each non-critical terminal device in the set of non-critical terminal devices from smallest to largest, and extract the corresponding devices from the sorted non-critical terminal devices based on the final number of migrated devices as the set of devices to be migrated.
[0055] Importantly, the above steps of this invention first obtain multiple terminal devices and multiple similar devices based on the high-load source device, and retrieve the minimum load rate of similar devices to obtain candidate target devices. This step is to clarify the carrier for load migration. Selecting the similar device with the minimum load rate as the candidate target device can ensure that the candidate target device has sufficient load-bearing capacity and avoid new performance bottlenecks caused by excessive load on the target device. Then, non-critical terminal devices are identified from multiple terminal devices and their numbers are counted. Focusing on non-critical terminal devices for migration is to avoid affecting the normal operation of critical business terminals while achieving load balancing, and to ensure the stability of core business. Subsequently, the theoretical number of devices to be migrated is calculated by combining the high load rate of the high-load source device, the load rate of the candidate target device, and the number of non-critical terminal devices. This can preliminarily determine a reasonable migration scale and provide a theoretical basis for load balancing. Then, the physical maximum number of connections and the current number of connections of the candidate target device are obtained, and the remaining number of connections is calculated. Since the physical maximum number of connections of the device is the upper limit of the hardware of the candidate target device, the remaining number of connections directly determines the scale of terminal connections that the candidate target device can add, and a more accurate assessment of the actual load-bearing capacity of the candidate target device is made. The migration safety ratio is then multiplied by the number of non-critical terminal devices, and the resulting value is rounded up to obtain the safety constraint value. The minimum value among the safety constraint value, the theoretical number of migrated devices, and the number of remaining connections is taken as the final number of migrated devices. The migration safety ratio is introduced in this invention to reserve a certain amount of load redundancy space to prevent the candidate target device from experiencing a sudden increase in load due to the instantaneous migration of too many terminal devices, thereby improving the safety and stability of load migration. Finally, the set of devices to be migrated is matched according to the final number of migrated devices and migrated to the candidate target device to obtain optimized nodes. This load migration method based on quantitative analysis and safety correction can reduce the load of high-load source devices, effectively balance the load distribution among similar devices, and improve the operating efficiency of the entire device cluster.
[0056] In detail, the formula for calculating the theoretical number of relocated devices is as follows: in, Indicates the theoretical number of devices to be migrated. Indicates a high load rate. Indicates the target device load rate. This indicates the number of non-critical terminal devices. The symbol indicates rounding up.
[0057] It should be explained that "high load rate" and "target device load rate" refer to the current load rate of the high-load source device and the current load rate of the candidate target device, respectively. The number of non-critical terminal devices is the number of non-critical terminal devices in the non-critical terminal device set. In the formula for calculating the theoretical number of migration devices in this invention, It is used to calculate the load rate difference between high-load source devices and candidate target devices, and then... Dividing by 2 is to achieve an approximate balanced load distribution, and then... The rounding is performed because the number of devices to be migrated must be a positive integer, and then multiplied by the nearest integer. This method can convert the load rate difference into the actual number of devices to be migrated, while also eliminating the influence of percentage units.
[0058] S8. Summarize the optimized nodes, recovered nodes, and normal nodes respectively to obtain the optimized node set, recovered node set, and normal node set.
[0059] It should be explained that the optimized node set is the set of all optimized nodes. The recovered node set is the set of all recovered nodes. The normal node set is the set of all normal nodes.
[0060] S9. High-performance intelligent management that performs self-optimization of network congestion based on optimized node sets, recovered node sets, and normal node sets.
[0061] It should be explained that this invention achieves automated and intelligent management of network congestion, reduces manual intervention, and improves the stability and reliability of network operation.
[0062] This invention receives high-performance intelligent management commands and obtains a network topology map based on these commands. The network topology map includes multiple nodes. This invention provides a complete understanding of the distribution and connections of all nodes in the network, offering visualized data support for subsequent key node identification and congestion localization. Multiple key path nodes are identified from the network topology map. For each key path node, the following operations are performed: A data acquisition module is configured to obtain the configured node; based on the configured node, node network parameters and node-related influence characteristic values are obtained. This invention enables nodes to acquire data through the data acquisition module configuration, accurately obtaining node network parameters and related influence characteristic values. This provides quantitative and accurate indicators for determining the root cause of congestion, solving the problem of inaccurate congestion localization caused by data gaps in traditional management. Based on the node network parameters and node-related influence characteristic values, the root cause congested nodes, recovered nodes, or normal nodes are identified. Based on quantitative indicators, precise classification of node status is achieved, quickly distinguishing between congested nodes requiring optimization, self-healed recovered nodes, and normal nodes requiring no intervention, thus improving congestion management. This targeted approach avoids ineffective operations on normal nodes, saving management resources. Based on the root cause congestion node, the type of the root cause congestion device is identified. If the root cause congestion device type is a core router or a network switch, and it is a core router, then the root cause congestion node is designated as a path optimization node, and routing path switching optimization is performed on it to obtain an optimized node. If the root cause congestion device type is a network switch, then it is designated as a node to be load balanced, and load balancing optimization is performed on it to obtain an optimized node. This invention optimizes path switching for core routers, quickly avoiding congested links and ensuring continuous and low-latency data transmission. It also optimizes load balancing for network switching equipment, balancing device load distribution and improving resource utilization. Through these two differentiated strategies, congestion problems are precisely resolved, improving optimization efficiency and effectiveness. Optimized nodes, recovered nodes, and normal nodes are aggregated to obtain optimized node sets, recovered node sets, and normal node sets, respectively. Based on these sets, high-performance intelligent management of network congestion self-optimization is achieved. Therefore, this invention can comprehensively improve network operational stability and resource utilization efficiency.
[0063] like Figure 2 The diagram shown is a functional block diagram of a high-performance intelligent management system for network congestion self-optimization provided in an embodiment of the present invention.
[0064] The high-performance intelligent management system 100 for network congestion self-optimization described in this invention can be installed in an electronic device. Depending on the functions implemented, the high-performance intelligent management system 100 for network congestion self-optimization may include a critical path node identification module 101, a node status confirmation module 102, a root cause node optimization module 103, and a network self-optimization completion module 104. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device. The critical path node identification module 101 is used to receive high-performance intelligent management instructions, obtain a network topology map according to the high-performance intelligent management instructions, wherein the network topology map includes multiple nodes, and identify multiple critical path nodes from the network topology map; The node status confirmation module 102 is used to perform the following operations on each of the multiple critical path nodes: configure the acquisition module of the critical path node to obtain the configured node, obtain the node network parameters and node association impact feature values based on the configured node, and confirm the root cause congestion node, recovered node or normal node according to the node network parameters and node association impact feature values. The root cause node optimization module 103 is used to determine the root cause congestion device type based on the root cause congestion node. The root cause congestion device type includes core routers or network switching devices. If the root cause congestion device type is a core router, the root cause congestion node is used as a path optimization node, and a routing path switching optimization operation is performed on the path optimization node to obtain an optimized node. If the root cause congestion device type is a network switching device, the root cause congestion node is used as a load balancing node, and a load balancing optimization operation is performed on the load balancing node to obtain an optimized node. The network self-optimization module 104 is used to summarize the optimized nodes, recovered nodes and normal nodes respectively to obtain the optimized node set, the recovered node set and the normal node set, and to complete the high-performance intelligent management of network congestion self-optimization based on the optimized node set, the recovered node set and the normal node set.
[0065] In detail, the modules in the high-performance intelligent management system 100 for network congestion self-optimization described in this embodiment of the invention employ the same methods as described above. Figure 1 The method uses the same high-performance intelligent management approach for network congestion self-optimization as described in the article, and can produce the same technical effect, so it will not be elaborated here.
[0066] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.
[0067] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0068] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0069] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A high-performance intelligent management method for self-optimization of network congestion, characterized in that, The method includes: Receive high-performance intelligent management instructions and obtain a network topology map based on the instructions. The network topology map includes multiple nodes. Multiple critical path nodes were identified from the network topology graph. For each critical path node, the following operations were performed: The acquisition module is configured for critical path nodes to obtain the configured nodes. Based on the configured nodes, the node network parameters and node association impact feature values are obtained. The step of obtaining node network parameters and node association influence feature values based on configured nodes includes: Obtain historical byte count and maximum interface bandwidth, collect data from the configured nodes using a preset collection interval, obtain the current byte count, and calculate the node bandwidth utilization rate based on the historical byte count, current byte count, and maximum interface bandwidth. Once the target and probe packet are identified, multiple data packets are sent to the target using the configured nodes and probe packets to obtain the number of transmissions, the number of receptions, and the average node latency. The node packet loss rate is calculated based on the number of transmissions and receptions. The node association impact characteristic value is calculated based on the configured nodes. The node network parameters are determined based on the node bandwidth utilization, the node average latency, and the node packet loss rate. The step of calculating the node association influence feature value based on the configured nodes includes: Based on the configured nodes in the network topology diagram, the associated target node set is identified, and the associated target node set is filtered to obtain the downstream directly connected node set; The downstream directly connected device set is determined based on the downstream directly connected node set. The downstream directly connected devices are extracted sequentially from the downstream directly connected device set. Network access parameters and device operating parameters are obtained based on the extracted downstream directly connected devices. Based on network access parameters and device operating parameters, anomalies are judged on the extracted downstream directly connected devices to obtain performance judgment results, which are either abnormal or normal. If the performance assessment result is abnormal, then the extracted downstream directly connected devices will be regarded as devices with abnormal performance. Sum the devices with abnormal performance to obtain the set of devices with abnormal performance. Then sum the set of devices with abnormal performance and the set of downstream directly connected devices to obtain the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is calculated based on the number of devices with abnormal performance and the number of downstream directly connected devices. The node association impact characteristic value is the value obtained by dividing the number of devices with abnormal performance by the number of downstream directly connected devices. Based on the node network parameters and the characteristic values of node associations, the root cause congestion nodes, recovered nodes, or normal nodes can be identified. The type of device causing the congestion is determined based on the root cause congestion node. The root cause congestion device type is either a core router or a network switching device. If the root cause congestion device type is a core router, then the root cause congestion node is used as a path optimization node, and a route path switching optimization operation is performed on the path optimization node to obtain the optimized node. If the root cause congestion device type is a network switching device, then the root cause congestion node is regarded as a node to be load balanced, and load balancing optimization operation is performed on the node to be load balanced to obtain the optimized node. The optimized nodes, recovered nodes, and normal nodes are summarized separately to obtain the optimized node set, the recovered node set, and the normal node set; High-performance intelligent management that achieves self-optimization of network congestion based on optimized node sets, recovered node sets, and normal node sets.
2. The high-performance intelligent management method for network congestion self-optimization as described in claim 1, characterized in that, The formula for calculating the node bandwidth utilization is as follows: in, Indicates node bandwidth utilization. Indicates the current number of bytes. Indicates the number of historical bytes. Indicates the data collection interval. This indicates the maximum bandwidth of the interface.
3. The high-performance intelligent management method for network congestion self-optimization as described in claim 2, characterized in that, The process of identifying root cause congested nodes, recovered nodes, or normal nodes based on node network parameters and node association impact characteristic values includes: If the node network parameters do not meet the preset standard node network parameter range conditions, the configured node is regarded as a suspected congested node, a first proportion threshold is obtained, and it is determined whether the node association influence feature value of the suspected congested node is greater than the first proportion threshold. If the node association impact characteristic value of a suspected congested node is greater than the first proportional threshold, then the suspected congested node is regarded as the root cause congestion node. If the node association impact feature value of a suspected congested node is not greater than the first proportion threshold, then the suspected congested node is regarded as an affected node, and a self-optimization operation is performed on the affected node to obtain the recovered node. If the node network parameters meet the standard node network parameter range conditions, then the configured node is considered a normal node.
4. The high-performance intelligent management method for network congestion self-optimization as described in claim 3, characterized in that, The process of obtaining the first ratio threshold includes: Obtain the historical alarm root cause device set, which includes multiple historical alarm root cause devices, and each historical alarm root cause device corresponds to an alarm trigger timestamp. Historical alarm root cause devices are extracted sequentially from the historical alarm root cause device set. A time window is obtained based on the alarm trigger timestamp corresponding to the extracted historical alarm root cause device. The set of directly connected devices is identified based on the time window. The proportion of abnormal downstream devices in history is calculated based on the set of directly connected devices. The historical downstream equipment abnormality ratios are summarized to obtain a historical downstream equipment abnormality ratio set. The median of the historical downstream equipment abnormality ratio set is extracted to obtain the median abnormality ratio, which is then used as the first ratio threshold.
5. The high-performance intelligent management method for network congestion self-optimization as described in claim 4, characterized in that, The step of performing route path switching optimization operations on the path optimization nodes to obtain optimized nodes includes: The alternative route path set is retrieved from the pre-built routing information database based on the path optimization node, and alternative route paths are extracted from the alternative route path set in turn. The path node set is confirmed based on the extracted alternative route paths, wherein the path node set includes one or more path nodes. Calculate the path node bandwidth utilization and path node latency value for each path node in the path node set to obtain the path node bandwidth utilization set and path node latency value set. The maximum value is extracted from the set of bandwidth utilization rates of path nodes to obtain the maximum bandwidth utilization rate of the path node. The total path delay value is obtained by summing the set of delay values of path nodes. The evaluation value of the backup path is calculated based on the maximum path node bandwidth utilization and the total path delay value. Summarize the evaluation values of the alternative paths to obtain the set of alternative path evaluation values, and take the alternative route path corresponding to the alternative path with the largest evaluation value in the set of alternative path evaluation values as the optimal alternative path. Based on the optimal backup path, the routing path of the path optimization node is switched to obtain the optimized node.
6. The high-performance intelligent management method for network congestion self-optimization as described in claim 5, characterized in that, The process of performing load balancing optimization on the node to be load balanced, resulting in an optimized node, includes: Based on the nodes to be load balanced, identify the high-load source devices, obtain multiple terminal devices and multiple similar devices based on the high-load source devices, perform minimum load rate search on the multiple similar devices, and obtain candidate target devices; Identify a set of non-critical terminal devices from multiple terminal devices and count the number of non-critical terminal devices in the set; Based on the high load source devices and candidate target devices, obtain the high load rate and target device load rate, and calculate the theoretical number of devices to be migrated based on the number of non-critical terminal devices, the high load rate and the target device load rate. Obtain the maximum physical number of connections and the current number of connections for the candidate target device, and calculate the remaining number of connections based on the maximum physical number of connections and the current number of connections; The theoretical number of devices to be migrated is adjusted based on the remaining number of connections and the preset migration security ratio to obtain the final number of devices to be migrated. Based on the final number of devices to be migrated, a set of devices to be migrated is matched from the set of non-critical terminal devices. The set of devices to be migrated is then migrated to the candidate target devices to obtain optimized nodes.
7. The high-performance intelligent management method for network congestion self-optimization as described in claim 6, characterized in that, The formula for calculating the theoretical number of relocated devices is as follows: in, Indicates the theoretical number of devices to be migrated. Indicates a high load rate. Indicates the target device load rate. This indicates the number of non-critical terminal devices. The symbol indicates rounding up.
8. A system applied to the high-performance intelligent management method for network congestion self-optimization as described in claim 1, characterized in that, The system includes: The critical path node identification module is used to receive high-performance intelligent management instructions, obtain a network topology map based on the instructions, wherein the network topology map includes multiple nodes, and identify multiple critical path nodes from the network topology map; The node status confirmation module is used to perform the following operations on each critical path node among multiple critical path nodes: configure the acquisition module of the critical path node to obtain the configured node, obtain the node network parameters and node association impact feature values based on the configured node, and confirm the root cause congestion node, recovered node or normal node based on the node network parameters and node association impact feature values. The root cause node optimization module is used to identify the root cause congestion device type based on the root cause congestion node. The root cause congestion device type includes core routers or network switching devices. If the root cause congestion device type is a core router, the root cause congestion node is used as a path optimization node, and a routing path switching optimization operation is performed on the path optimization node to obtain an optimized node. If the root cause congestion device type is a network switching device, the root cause congestion node is used as a load balancing node, and a load balancing optimization operation is performed on the load balancing node to obtain an optimized node. The network self-optimization module is used to summarize optimized nodes, recovered nodes, and normal nodes respectively, to obtain optimized node sets, recovered node sets, and normal node sets. Based on the optimized node sets, recovered node sets, and normal node sets, it completes high-performance intelligent management of network congestion self-optimization.
Citation Information
Patent Citations
Multi-platform collaborative intelligent comprehensive operation management system
CN118612086A
Multi-purpose routing optimization method and optimization system based on hierarchical traversal, and router
CN118714062A