Method for locating reason of network dial testing quality index abnormity

Through multi-source data acquisition and network topology modeling, combined with graph theory, information theory, game theory and deep reinforcement learning technology, network paths and loads are dynamically adjusted, and the static and singular problems of network performance monitoring and traffic optimization in the existing technology are solved, real-time monitoring and efficient optimization of network performance are achieved.

CN119996196APending Publication Date: 2025-05-13TOP XINGDA
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510155703.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has problems such as insufficient static configuration, single-objective optimization and real-time performance in network performance monitoring and traffic optimization, resulting in network bottlenecks not being discovered and resolved in a timely manner, and the accuracy of abnormal detection is insufficient.

Method used

Through multi-source data acquisition and network topology modeling, based on graph theory, information theory, game theory and deep reinforcement learning technology, dynamic path adjustment, load optimization and multi-objective optimization are achieved, and the causes of network abnormalities are located and optimization suggestions are provided.

Benefits of technology

Real-time monitoring and dynamic optimization of network performance are realized, which avoids inefficiency and misjudgment risks in traditional methods, improves the robustness and response speed of the network, and ensures efficient abnormal positioning and optimization adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996196A_ABST
    Figure CN119996196A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer networks, and discloses a method for locating reasons for network dial test quality index abnormity, which comprises the following steps of: multi-source data acquisition and network topology modeling: generating state information of network equipment and links by acquiring performance data of each equipment and flow data of the links in a network, and establishing a network topology model; establishing a network topological graph model; shortest path calculation and abnormal path identification based on the graph theory: calculating a shortest path from source equipment to target equipment by using a shortest path algorithm in the graph theory, and identifying abnormal links and equipment in the path; and load evaluation and bottleneck analysis based on an information theory. Through combination of multi-source data acquisition, graph theory optimization, information theory evaluation, game theory resource allocation, deep reinforcement learning and multi-target optimization, network paths and loads are dynamically adjusted, abnormal sources are accurately positioned, and network performance, stability and adaptive capacity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of computer networks, and in particular to a method for locating the cause of abnormal network dialing quality indicators. Background Art

[0002] In modern society, data communication networks have become the cornerstone of people's lives and work. From daily Internet browsing to enterprise-level cloud computing and big data processing, the quality of the network directly affects the efficiency and stability of information transmission. However, performance issues such as latency, packet loss rate, and bandwidth utilization in the network often lead to service interruptions, poor user experience, and even business failures. In a complex network environment, how to monitor, accurately locate, and optimize network anomalies in real time has become the key to improving network performance and ensuring business continuity.

[0003] At present, existing technologies mainly rely on network management protocols (such as SNMP) and traffic monitoring protocols (such as NetFlow) for data collection, and optimize network paths and resource allocation based on static configurations or rules. These methods can perform basic monitoring and management of the network, and are mainly used to evaluate link bandwidth utilization, latency, packet loss rate and other indicators.

[0004] Although existing technologies have provided some help in network performance monitoring and traffic optimization, their limitations are still obvious. First, most existing methods rely on static configurations and rules, and are slow to respond to frequently changing loads and fault conditions in the network. Factors such as latency and packet loss rate in the network need to be evaluated in real time, but traditional methods cannot dynamically adjust path selection and load distribution in a timely manner, resulting in network bottlenecks not being discovered and resolved in a timely manner. Second, the optimization methods of existing technologies often only focus on a single goal and ignore the balance between multiple network quality indicators. In a complex network environment, optimizing latency may affect bandwidth utilization, and reducing packet loss rate may increase latency. The contradictions between these goals need to be effectively considered comprehensively. However, the multi-objective optimization in existing technologies is mostly simplified and cannot accurately find the optimal balance between multiple goals. Finally, the existing anomaly detection and network optimization methods are not accurate enough. They usually rely on historical performance data or empirical rules to predict and identify network failures, lack sufficient real-time performance and accuracy, and traditional methods often have the risk of misjudgment and missed judgment. They cannot quickly respond to sudden changes in network status and are difficult to achieve efficient anomaly location and optimization adjustments. Summary of the invention

[0005] In view of the deficiencies of the prior art, the present invention provides a method for locating the cause of abnormal network dialing quality indicators, which solves the problems of static, single and slow response of network optimization methods in the prior art.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for locating the cause of abnormal network dialing quality indicators, the method comprising the following steps: Multi-source data collection and network topology modeling: By collecting the performance data of each device in the network and the traffic data of the link, the status information of the network device and link is generated, and the network topology model is established; Shortest path calculation and abnormal path identification based on graph theory: Use the shortest path algorithm in graph theory to calculate the shortest path from the source device to the target device and identify abnormal links and devices in the path; Load evaluation and bottleneck analysis based on information theory: Information entropy is calculated based on the performance status of devices and links to evaluate the load of devices or links and identify performance bottlenecks in the network; Resource allocation and optimization based on game theory: Use game theory models to optimize the resources of devices and links in the network, avoid resource conflicts in the network, and improve network performance; Dynamic path adjustment and load optimization based on deep reinforcement learning: Through deep reinforcement learning training, network path selection and load distribution strategies are dynamically adjusted to optimize network performance and reduce quality anomalies; Multi-objective optimization and comprehensive anomaly location: A multi-objective optimization algorithm is used to balance and optimize multiple quality indicators, locate the causes of anomalies in the network, and provide network configuration optimization suggestions.

[0007] Preferably, the multi-source data collection and network topology modeling further includes: Collect real-time performance data of each device through network management protocols, including CPU utilization, memory usage and motherboard temperature; Collect link bandwidth utilization, packet loss rate, and latency data through network traffic monitoring protocols; Based on the collected device and link data, a network topology diagram is constructed, the devices and links are abstracted as nodes and edges in the diagram, and each link is assigned a corresponding quality indicator weight.

[0008] Preferably, the shortest path calculation and abnormal path identification based on graph theory further includes: For each link, its weight is calculated based on its latency, packet loss rate, and bandwidth performance indicators. The link weight is obtained through a weighted calculation formula. The Dijkstra algorithm is used to calculate the shortest path from the source device to the target device, identify potential abnormal paths, and make abnormal judgments based on the weights of the links in the path.

[0009] Preferably, the load assessment and bottleneck analysis based on information theory further includes: By calculating the load status information entropy of the device or link, the current load level is evaluated; By analyzing the redundancy of devices or links, possible performance bottlenecks are identified. If the redundancy of a device or link is lower than a preset threshold, it is considered a bottleneck.

[0010] Preferably, the resource allocation and optimization based on game theory further includes: A resource allocation model between devices is constructed through the framework of game theory. The device strategies include adjusting bandwidth allocation and load balancing. Nash equilibrium is solved to determine the optimal resource allocation strategy for each device under the condition that the strategies of other devices are fixed, thereby optimizing the use of network resources.

[0011] Preferably, the dynamic path adjustment and load optimization based on deep reinforcement learning further includes: The state space is defined to include the performance status of the device and the traffic status of the link, and the action space includes path selection and load adjustment. The reward function is designed to evaluate the reward for each action based on the network's latency and packet loss rate, and encourage the selection of actions that reduce network quality anomalies. The Q-learning algorithm is used to train the intelligent agent, and the path selection and load distribution strategies are continuously optimized based on historical network status and action feedback.

[0012] Preferably, the multi-objective optimization and comprehensive anomaly positioning further includes: Construct a multi-objective optimization function, comprehensively consider multiple network quality indicators, and aim to minimize the overall performance loss of the network; calculate the Pareto optimal solution through a multi-objective optimization algorithm, and select an optimal solution that does not compromise other objectives.

[0013] Preferably, the shortest path calculation further comprises: When calculating the shortest path, the link load and real-time performance changes are taken into account and the path selection is dynamically adjusted.

[0014] Preferably, the game model establishment further includes: Define the strategies of devices in the game, including bandwidth allocation, data forwarding strategy, and link selection, to reduce network resource conflicts.

[0015] Preferably, the deep reinforcement learning training further comprises: The training process of deep reinforcement learning is optimized, and a more efficient learning rate adjustment method and reward mechanism are adopted to ensure that the intelligent agent can quickly adapt to changes in the network environment and optimize path selection and load distribution.

[0016] The present invention provides a method for locating the cause of abnormal network dialing quality indicators. It has the following beneficial effects: 1. The resource allocation and deep reinforcement learning technology based on game theory can dynamically adjust network paths and load distribution and respond to network load changes in real time. This intelligent resource optimization method avoids the inefficiency caused by traditional static path selection and manual adjustment, and ensures load balancing and stability in complex environments.

[0017] 2. Through multi-objective optimization and information entropy analysis, the present invention can simultaneously consider multiple performance indicators such as latency, bandwidth utilization and packet loss rate, accurately identify and locate bottlenecks and abnormal sources in the network. Compared with the prior art, this provides a more comprehensive and accurate abnormality diagnosis method, which helps to optimize the network and allocate resources more efficiently.

[0018] 3. The present invention combines graph theory, information theory and deep reinforcement learning technology to dynamically adjust the optimization strategy according to the real-time network status. This adaptive capability enables the network to automatically adjust resource allocation when traffic fluctuates and device load changes, avoiding the limitations of static configuration in traditional methods and improving the robustness and response speed of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] Please see attached Figure 1 The embodiment of the present invention provides a method for locating the cause of abnormal network dialing quality indicators, the method comprising the following steps: S1. Multi-source data collection and network topology modeling: By collecting the performance data of each device in the network and the traffic data of the link, the status information of the network device and link is generated, and the network topology model is established; By comprehensively collecting the performance data of devices in the network and the traffic data of links, a complete network status view can be constructed to help accurately locate the specific location and cause of the anomaly. This step not only includes the real-time status data collection of devices and links, but also involves how to organize this data into a network topology diagram to provide a reliable data foundation for subsequent algorithm analysis and optimization.

[0022] Specifically, in this embodiment, different types of data are collected through multiple protocols, and an efficient and accurate network monitoring and analysis framework is constructed by using data fusion and network topology modeling. This framework provides a solid foundation for subsequent graph-based shortest path calculation, information theory load assessment, game theory optimization, and deep reinforcement learning path adjustment.

[0023] Multi-source data collection In this embodiment, the performance data of each device in the network is first collected through a network management protocol (such as SNMP). These data include the CPU utilization, memory occupancy and motherboard temperature of the device, which can reflect the operating status and load of the device. In actual applications, the SNMP protocol requests the device status at regular intervals and returns corresponding statistical data. The device performance indicators are usually collected every five minutes and continuously updated.

[0024] At the same time, the bandwidth utilization, packet loss rate and delay data of the network link are collected through network traffic monitoring protocols (such as NetFlow protocol). NetFlow protocol tracks the changes of traffic in real time by sampling link traffic. These data help reflect the performance status of the links in the network and whether there are problems such as network bottlenecks and link overloads.

[0025] In the process of collecting data, the integration of device performance and link traffic data is very important. The data obtained through these two protocols can fully present the status of each device and link in the network.

[0026] After data collection, the next step is to build a network topology diagram to describe the connection relationship between devices. The network topology diagram is represented by a graph theory model, with device nodes as nodes in the graph and links between devices as edges in the graph. The weight of each link is determined by the quality indicators of the link (such as latency, packet loss rate, and bandwidth). The construction of network topology is not only about modeling the physical location or structure of network devices, but more importantly, how to describe the relationship between devices in the network through the performance indicators of the links.

[0027] In a possible implementation, the construction of the network topology map can be completed by the following steps: Node definition: Each network device is uniquely identified by a device ID. The performance status of the device node is represented by the performance data vector of the device. The device performance vector includes indicators such as CPU utilization, memory usage, and motherboard temperature, which are used to describe the health status of the device.

[0028] Edge definition: The connection between devices is represented by the link quality index. In the network topology, the edge weight can be composed of link delay, packet loss rate and bandwidth utilization. The link weight is calculated by the weighted calculation formula: w ij =α·Delayij +β·PacketLoss ij +γ·(1-BandwidthUtilization ij ) Where: w ij is the weight of the link between device i and device j, indicating the quality of the link; Delay ij is the link delay, in milliseconds (ms), reflecting the delay of link transmission; PacketLoss ij The link packet loss rate is expressed in percentage (%), reflecting the ratio of data packets lost on the link; BandwidthUtilization ij is the link bandwidth utilization, in percentage (%), indicating the usage of the link bandwidth; α, β, γ are weight coefficients, which are used to adjust the influence of link delay, packet loss rate and bandwidth utilization on link quality, and are usually adjusted based on experience or network requirements.

[0029] As an option, these weight coefficients α, β, γ can be dynamically adjusted according to the actual situation of the network to give priority to optimizing certain specific quality indicators (such as giving priority to latency in a network with high real-time requirements).

[0030] When building a network topology graph, you first need to map devices and links to nodes and edges of the graph, and then assign an initial weight to each edge based on the collected link quality data. Specifically, the process of building a network topology graph includes the following steps: Node identification: Assign a node to each device through the device's unique identifier (such as device IP, device ID, etc.).

[0031] Link identification: Determine the link and assign a weight based on the connection status between devices (such as physical link, virtual link, etc.). The weight is calculated using the above formula.

[0032] Graph construction: Devices and links are constructed into a network topology graph. The weight of each link represents the quality information of the link. The topology graph is a directed graph, in which the direction of the edge represents the direction of the data flow, and the weight of the edge represents the quality of the link.

[0033] In one possible implementation, the generation of the network topology map can be automated. Based on the performance data of devices and links, a network monitoring platform or a network topology discovery tool (such as Cisco Prime, Nagios, etc.) can be used to automatically generate a topology map and update it in real time.

[0034] After completing the construction of the network topology, the next step is data fusion and network status representation. The quality indicator data of devices and links can be fused in a multi-dimensional way to form a unified network status representation. This representation includes the health status of each device and link, as well as the interconnection information between devices.

[0035] Specifically, the network state vector x i It can be represented as a multi-dimensional data vector, where each dimension represents a specific network quality indicator (such as latency, bandwidth, load, etc.). These network state vectors provide reliable data support for subsequent algorithm analysis, so that subsequent path optimization, load evaluation and anomaly location can be based on the actual network operation status.

[0036] By combining multi-source data collection and network topology modeling, comprehensive collection and modeling of network status are completed.

[0037] S2. Shortest path calculation and abnormal path identification based on graph theory: Use the shortest path algorithm in graph theory to calculate the shortest path from the source device to the target device, and identify abnormal links and devices in the path; The shortest path calculation and abnormal path identification steps based on graph theory came into being. The purpose is to calculate the shortest path in the network, analyze whether there are abnormal links or devices, and further identify the root causes that may cause network quality problems. The shortest path algorithm is the core method of this step. By evaluating the quality of each link, it can effectively identify potential bottlenecks in the network and help operation and maintenance personnel accurately locate abnormalities.

[0038] Specifically, the Dijkstra algorithm is used as the main method for shortest path calculation in this embodiment. The Dijkstra algorithm calculates the shortest path from the source device to the target device by assigning weights to each link, and determines whether there are abnormal links based on the link weights in the path. This calculation method takes into account multiple performance indicators of the link (such as latency, packet loss rate, and bandwidth utilization), and can provide a comprehensive assessment of the health of the network.

[0039] In the network topology, each link has a corresponding weight. Next, the Dijkstra algorithm is used to calculate the shortest path from the source device to the target device. The Dijkstra algorithm is a classic graph theory algorithm used to calculate the shortest path from a source node to all other nodes in a weighted graph.

[0040] Alternatively, the shortest path cost can be defined as: Where: d(v s ,v t ) indicates that from the source node vs To the target node v t The shortest path cost of s To the target node v t The set of all paths of ij for each link on the path; w ij Link e ij The weight of; π: represents the weight from the source node v s To the target node v t Any path of represents the sum of the weights of all links on the path π.

[0041] In one possible implementation, the Dijkstra algorithm traverses the nodes and links in the graph, gradually updating the shortest path cost of each node until the shortest path from the source device to the target device is finally found. During the path calculation process, the algorithm compares the costs of different paths and selects the path with the minimum cost. The weight of the link determines the cost of the path. If the weight of a link is too high, the cost of the path will increase, thus affecting the selection of the shortest path.

[0042] After calculating the shortest path from the source device to the target device, the next step is to analyze the path and identify possible abnormal links or devices. Abnormal paths usually have high link weights, indicating that there are problems with the link's latency, packet loss rate, or bandwidth utilization, which may be the cause of the decline in network quality.

[0043] Specifically, if the weight of some links in the path exceeds the preset threshold, it can be determined as an abnormal link. The threshold can be set based on the actual needs of the network or experience. For example, if the packet loss rate of a link is higher than 2% or the delay exceeds 100ms, it is considered that the link may be abnormal.

[0044] In one possible implementation, a dynamic tolerance range can be set for the weight of each link, and the threshold can be adjusted dynamically. For example, during peak network hours, the load of the link may temporarily increase, resulting in a slight increase in latency and packet loss rate. At this time, the tolerance of anomaly detection can be appropriately increased. In this way, false alarms can be avoided and the accuracy of anomaly detection can be improved.

[0045] Based on the weight of the links in the path, determine whether there are abnormal links. The criterion for determining abnormal links is usually that the weight of the link exceeds a preset threshold. If abnormal links are found, further analyze the device load and link status behind these links to help quickly locate the root cause of the network abnormality.

[0046] For example, if the weight of a link in the path is w ijIf the preset threshold value θ is exceeded, the following formula can be used to determine whether the link is abnormal: w ij >θ Among them, θ is a preset threshold, which can usually be adjusted according to the actual situation of the network.

[0047] For example, in a real network, suppose there is a link between device A and device B, with a latency of 120ms, a packet loss rate of 2%, and a bandwidth utilization rate of 90%. According to the preset weight coefficients α=1, β=2, γ=1, the weight of the link is calculated as: w AB =1·120+2·2+1·(1-0.9)=120+4+0.1=124.1 If the preset threshold value θ=100, the weight of the link exceeds the threshold value, and it is considered that the link is abnormal.

[0048] As another option, if the abnormality of latency or packet loss rate is very serious in a particular case, a stricter threshold can be further set. For example, if the packet loss rate exceeds 5% or the latency exceeds 200ms, it is directly marked as a serious abnormality.

[0049] By calculating the shortest path in the network topology graph, the graph-theory-based shortest path calculation and abnormal path identification steps effectively help identify abnormal links and devices that may cause network quality degradation.

[0050] S3, Load evaluation and bottleneck analysis based on information theory: Calculate information entropy based on the performance status of devices and links to evaluate the load of devices or links and identify performance bottlenecks in the network; Through information theory entropy calculation and link redundancy analysis, we can further evaluate the load conditions of devices and links and identify possible performance bottlenecks. This step can provide a more comprehensive understanding of the network's operating status, accurately locate bottleneck devices or links in the network, and ensure the accuracy of abnormal location.

[0051] Specifically, in this embodiment, information entropy in information theory is used to measure the load status of a device or link, and redundancy analysis helps determine whether a link is likely to become a performance bottleneck. This step is further connected with the aforementioned steps S1 (multi-source data collection and network topology modeling) and S2 (shortest path calculation and abnormal path identification), providing an in-depth analysis of the network status and helping to further optimize path and resource allocation.

[0052] In this embodiment, information entropy is used to evaluate the load level of devices and links in the network. Information entropy is a measure of system uncertainty or information volume. Devices or links with higher loads usually show higher entropy values, which indicates that the device or link may face performance bottlenecks when processing large amounts of traffic.

[0053] In general, the formula for calculating information entropy is as follows: Among them, H(X): represents information entropy, which is used to measure the uncertainty of system X; p(s i ): represents a state s in system X i The probability of state s i The probability of occurrence in the system; i : represents a specific state in system X; n: represents the number of all possible states in the system; log 2 : represents the logarithm operation with base 2; In the network, the load status of a device or link (such as latency, bandwidth utilization, packet loss rate, etc.) is obtained through sampling data. When a device or link is under a high load, p(s i ) will tend to be evenly distributed, resulting in an increase in information entropy, indicating that the device or link is under high pressure. If the entropy value of a device or link is high, it means that it is unstable when processing traffic and may become a bottleneck for network performance.

[0054] As an option, if the load status of the device changes frequently or the load distribution of the link is relatively uniform, the calculated value of the information entropy may increase, indicating that the device or link needs to be further optimized to avoid performance degradation.

[0055] Link redundancy analysis is another important part of this step, which aims to evaluate whether the link has a performance bottleneck. The lower the redundancy of the link, the more traffic it carries and may become a bottleneck.

[0056] In a possible implementation, the link redundancy R ij Calculated by the following formula: Its, R ij Link e ij The redundancy of the link indicates the remaining ratio of the link bandwidth; BandwidthUtilization ij Link e ij Bandwidth utilization; MaxBandwidth ij Link e ij The maximum bandwidth of The lower the redundancy of a link, the closer the bandwidth usage of the link is to its maximum carrying capacity, and there is a risk of overload. If the redundancy is lower than a certain set threshold, the link may become a bottleneck, thus affecting the performance of the overall network.

[0057] In general, if the redundancy of a link Rij If the redundancy of a link is lower than the preset threshold θ, it is considered that the link has a performance bottleneck. The specific threshold setting can be adjusted according to the network usage. For example, if the redundancy of a link is lower than 20%, it can be considered that the load of the link is high and may cause a network bottleneck.

[0058] Equipment load and bottleneck identification The load status of a device is usually closely related to its CPU and memory usage. By calculating the entropy value of the device's load status, we can determine whether the device is at risk of overload. If the device's load entropy value is high, it means that the device may be processing a large amount of data traffic and is at risk of overload, which may become a bottleneck of the network.

[0059] As an alternative, during peak network load periods, the load on some devices may increase, resulting in an increase in information entropy. At this time, the calculation of entropy values ​​can help identify which devices are facing a larger load, so that resource adjustments or path optimization can be made in a timely manner.

[0060] For example, suppose the bandwidth utilization of a link is 85% and the maximum bandwidth is 100 Mbps. According to the redundancy calculation formula, the link redundancy is: If the redundancy threshold is set to 0.2, the redundancy of the link is lower than the threshold, indicating that the bandwidth of the link is close to saturation and may become a performance bottleneck.

[0061] Another example, assuming that the CPU utilization of a device is 95% and the memory usage is 90%, the load entropy of the device may be high, indicating that the device is processing a large amount of data traffic. According to the information entropy formula, it is possible to further evaluate whether the device is overloaded and adjust the network configuration based on the entropy value.

[0062] This step effectively evaluates the load status of devices and links through information entropy calculation and link redundancy analysis in information theory, and further identifies bottlenecks in the network. In some cases, when the redundancy of the link is low or the load of the device is high, it may become a performance bottleneck and affect the overall network quality. Through these analyses, network bottlenecks can be discovered in time, network overload can be avoided, and the stability and reliability of the network can be improved.

[0063] Generally speaking, this method can help accurately identify high-load devices or links in the network, optimize and adjust them in a timely manner, avoid performance degradation and quality abnormalities, and provide reliable data support for subsequent path optimization and resource adjustments.

[0064] S4. Resource allocation and optimization based on game theory: Use game theory models to optimize the resources of devices and links in the network, avoid resource conflicts in the network, and improve network performance; The resource allocation and optimization steps based on game theory will optimize the allocation of resources in the network through the game theory model, avoid resource conflicts, and improve the overall network performance. The game theory model can simulate the resource competition and cooperation between multiple devices or links, so as to find the global optimal resource allocation solution.

[0065] Specifically, in this embodiment, we regard the devices and links in the network as participants in the game, and the strategy of each device is to adjust its bandwidth allocation and load balancing to maximize its own network performance. Under the framework of game theory, by solving the Nash equilibrium, each device can choose the optimal resource allocation solution when the strategies of other devices are fixed. The ultimate goal is to optimize the use of network resources, avoid overload of devices or links, and ensure the stability of network quality.

[0066] In this embodiment, we first need to establish a game model for each device and link. The strategy of each device includes operations such as bandwidth allocation, data forwarding strategy, and link selection. The participants in the game are devices or links in the network, and they optimize their respective network performance by selecting different strategies.

[0067] In general, each device wants to minimize its own performance indicators such as latency and packet loss rate, and the strategies of other devices will also affect its performance. Therefore, the interaction between devices can be described by a game theory model. When each device chooses a strategy, it considers how to choose the most advantageous strategy while keeping the strategies of other devices unchanged.

[0068] Specifically, the payment function of the device can be expressed as: U i (x 1 ,x 2 ,…,x n )=-(Delay i (x i )+PacketLoss i (x i )) Among them: U i (x 1 ,x 2 ,…,x n ) is the payment function of device i, which represents the utility value after device i selects the strategy; Delay i (x i ) for device i in policy x i Delay under PacketLoss i (x i ) for device i in policy x i Packet loss rate under x 1 ,x 2,…,x n is the strategy for devices 1, 2, …, n.

[0069] As an alternative, the payoff function U i (x 1 ,x 2 ,…,x n )The performance of the device is measured by the weighted sum of latency and packet loss rate. The goal of the device is to select a strategy that maximizes its payment function, that is, to minimize latency and packet loss rate while keeping other device strategies unchanged.

[0070] The core idea of ​​game theory is Nash equilibrium, which means that in a multi-party game, the strategy choices of all participants have reached the optimal state, and no participant can improve its own payment by unilaterally changing the strategy. In the context of network resource allocation, Nash equilibrium means that each device chooses an optimal resource allocation strategy to achieve the best performance under the premise that other devices choose a fixed strategy.

[0071] In general, we determine the optimal resource allocation strategy for each device by solving Nash equilibrium. In practice, the methods for solving Nash equilibrium usually include iteration and best response dynamics. In the iteration method, each device updates its strategy according to the current state of the network until the strategies of all devices reach a stable state, which is Nash equilibrium.

[0072] Specifically, when the strategies of other devices are fixed, device i selects the optimal resource allocation strategy Maximize the payment function. By solving the Nash equilibrium in the game, we can get the optimal resource allocation plan for the entire network.

[0073] By solving the game theory model, the optimal allocation of network resources is finally achieved. When each device chooses a strategy, it considers how to optimize its own performance while avoiding competing for resources with other devices. Under the framework of game theory, resource competition and coordination between devices are achieved by solving Nash equilibrium, thereby ensuring the effective use of network resources and avoiding the emergence of network bottlenecks.

[0074] In one possible implementation, when resources are scarce, there may be game conflicts between devices. The result of the game may be that some devices choose non-optimal strategies, causing network congestion or increased latency. Through the game theory model, devices can be guided to choose more reasonable resource allocation strategies, reduce resource conflicts, and thus improve overall network performance.

[0075] For example, suppose there are two devices A and B in the network. They share a link and need to allocate the bandwidth of the link. If devices A and B choose different bandwidth allocation strategies x A and xB , their performance (such as delay and packet loss rate) will be affected. Device A and device B choose the optimal strategy through game to maximize the total utility of both.

[0076] For example, if device A selects the bandwidth allocation strategy Device B selects a bandwidth allocation strategy Then both can obtain lower latency and packet loss rate, thus achieving optimal network performance.

[0077] This step optimizes the resource allocation in the network through game theory, avoiding resource competition between devices or links in the network. By solving the Nash equilibrium, the device can choose the optimal strategy when the strategies of other devices are fixed, ensuring the efficient use of network resources. The game theory model can handle the complex interaction between multiple devices and provide an effective network optimization method.

[0078] In general, this method can effectively avoid device overload and bandwidth waste, helping to improve the overall performance and stability of the network. When bottlenecks occur in the network, game theory optimization can guide devices to adjust resource allocation strategies, thereby improving network quality.

[0079] S5. Dynamic path adjustment and load optimization based on deep reinforcement learning: Through deep reinforcement learning training, network path selection and load distribution strategies are dynamically adjusted to optimize network performance and reduce quality anomalies; This step uses deep reinforcement learning (DRL) technology to enable the network to dynamically adjust path selection and load distribution strategies based on real-time data, further improving the stability and overall performance of the network.

[0080] Specifically, this embodiment transforms the network state into the state space in deep reinforcement learning, transforms the device's path selection, load distribution and other decisions into the action space, and trains the agent through the reinforcement learning model so that it automatically selects the optimal strategy based on the real-time network state. The agent continuously optimizes the strategy through interaction with the environment (i.e., making decisions based on the network state and obtaining feedback) to achieve dynamic adjustment of the path and optimal load distribution.

[0081] In the deep reinforcement learning framework, we first need to define the state space and action space, which is the basis of model training.

[0082] Generally speaking, the devices and link performance in the network constitute the state space. The device status can include parameters such as CPU utilization, memory usage, link bandwidth utilization, latency, and packet loss rate. All these parameters together describe the current operating status of the network.

[0083] Specifically, in this embodiment, the state space s can be represented as a multidimensional vector, which includes the performance indicators of all devices and links in the network. For each device i, its state vector x i It can be defined as: Where: Delay i Indicates the delay of device i; PacketLoss i Indicates the packet loss rate of device i; BandwidthUtilization i Indicates the bandwidth utilization of device i; CPU i and Memory i Represent the CPU and memory load of device i respectively.

[0084] As an option, in addition to device performance, network topology, link status, etc. can also be incorporated into the state space to further enhance the agent's perception of the network environment.

[0085] The action space refers to the operations that a device in the network can choose. Specifically, the actions of a device can include: Path selection: determines which link the data flows through.

[0086] Load adjustment: Adjust the load distribution according to the link load conditions.

[0087] Generally, the action space of a device includes selecting a suitable path or adjusting bandwidth, load, etc. The choice of action directly affects network performance, so it is necessary to optimize the strategy through reinforcement learning algorithms.

[0088] In deep reinforcement learning, the reward function is used to evaluate the quality of each action and help the agent learn how to choose the optimal strategy during training. The reward function is usually composed of indicators such as network latency, packet loss rate, and bandwidth utilization. The purpose is to improve network performance by optimizing these indicators.

[0089] Specifically, the reward function R(s,a) in this embodiment is designed as follows: R(s,a)=-(Delay(s,a)+PacketLoss(s,a)) Where: Delay(s,a) represents the network delay after action a is selected in state s; PacketLoss(s,a) represents the network packet loss rate after action a is selected in state s.

[0090] As an option, the reward function can also include other indicators such as bandwidth utilization, taking into account the load balancing and resource utilization of the network. When designing the reward function, the goal is to improve the overall network performance by encouraging the selection of low-latency, low-packet-loss rate paths and load balancing strategies.

[0091] In this embodiment, the Q-learning algorithm and the deep Q network (DQN) are used to train the agent. Q-learning is a reinforcement learning method that finds the optimal strategy by updating the Q value. The Q value represents the expected benefit of selecting a certain action in a certain state. Through repeated trials and updates, the agent can gradually learn the optimal path selection and load distribution strategy.

[0092] In general, the update formula of Q-learning is as follows: Q(s t ,a t )=Q(s t ,a t )+α(r t +γmax a′ Q(s t+1 ,a′)-Q(s t ,a t )) Where: Q(s t ,a t ) means that at time step t, the state s t Next select action a t Q value; r t Indicates that at time step t, action a is selected t The immediate reward obtained after the reward is obtained; γ is the discount factor, which indicates the discount degree of future rewards; α is the learning rate, which controls the amplitude of the update; max a′ Q(s t+1 ,a′) is the next state s t+1 The maximum Q-value of all possible actions represents an estimate of future rewards.

[0093] Specifically, the agent learns through training how to select the optimal path and load distribution strategy under different network conditions, ultimately optimizing the overall performance of the network. The agent will update its path selection and load distribution strategy according to the real-time network status, continuously optimizing network performance.

[0094] S6. Multi-objective optimization and comprehensive anomaly location: Use a multi-objective optimization algorithm to balance and optimize multiple quality indicators, locate the cause of anomalies in the network, and provide network configuration optimization suggestions; The multi-objective optimization and comprehensive anomaly location step will combine multiple optimization goals, comprehensively analyze the network's quality indicators through a multi-objective optimization algorithm, locate the source of anomalies, and provide adjustment suggestions for network configuration based on the optimization results. The core purpose of this step is to balance multiple network quality indicators and find the optimal solution that does not harm other goals, thereby ensuring the overall healthy operation of the network.

[0095] Specifically, the multi-objective optimization algorithm in this embodiment not only considers indicators such as network delay, packet loss rate and bandwidth utilization, but also integrates factors such as network topology and equipment load to ensure that while improving network performance, it will not cause excessive negative impact on other performance indicators. By solving the Pareto optimal solution, this step can locate the root cause of network anomalies and provide the most appropriate optimization solution.

[0096] In a network, there may be contradictions between different quality indicators. For example, reducing latency may require increasing bandwidth utilization, but this may lead to an increase in packet loss rate. Therefore, this embodiment uses multi-objective optimization to balance these objectives and ensure optimal overall network performance.

[0097] Generally, we consider various quality indicators by constructing a multi-objective optimization function. These objectives include: Latency: Optimize the response time of data transmission.

[0098] Packet loss rate: Minimize data loss and ensure data integrity.

[0099] Bandwidth utilization: Improve the efficiency of link bandwidth utilization and avoid link idleness or overload.

[0100] As an option, the multi-objective optimization function can be adjusted according to the actual application scenario. For example, in a real-time communication network, a higher weight can be given to latency, while in a network with large-volume data transmission, more attention can be paid to bandwidth utilization and packet loss rate.

[0101] Pareto optimal solution is a very important concept in multi-objective optimization. A solution is Pareto optimal if and only if there is no other solution that is better than it in all objectives. That is, if choosing a solution will make one objective better, other objectives may become worse. Therefore, the Pareto optimal solution represents the best solution to balance between multiple given objectives.

[0102] Specifically, in this embodiment, by calculating the Pareto optimal solution, we can find an optimal resource allocation solution for each network state. These optimal solutions may not optimize all objectives at the same time, but they achieve the best balance between multiple objectives.

[0103] In general, in order to calculate the Pareto optimal solution, we need to evaluate all possible network configurations, calculate their performance in terms of latency, packet loss rate, and bandwidth utilization, and then select those configurations that cannot be dominated by other solutions in all objectives. Through multiple iterations and optimization, we finally get a set of Pareto optimal solutions.

[0104] After obtaining the Pareto optimal solution through multi-objective optimization, the next step is to locate network anomalies and provide configuration adjustment suggestions based on the optimization results. Comprehensive anomaly location analyzes the performance of each device and link to find out the bottlenecks that may affect the overall network performance.

[0105] Specifically, by comparing the latency, packet loss rate, bandwidth utilization and other indicators of each network path, it is possible to quickly identify which paths or links have abnormalities. When the performance indicators of a path are much higher than those of other paths, it can be considered that the path or link has an abnormality, which may be the cause of the deterioration of network quality.

[0106] As an option, if the CPU or memory load of a device is high, it may cause data forwarding delays and increase latency. Therefore, when optimizing network configuration, you can consider resource expansion or load sharing for these high-load devices.

[0107] For example, suppose in a multi-path network, latency, packet loss rate, and bandwidth utilization are the objectives to be optimized. Through multi-objective optimization, we may obtain the following optimization results: On path P 1 The latency is low, but the bandwidth utilization is low and the packet loss rate is high; On path P 2 The latency is high, but the bandwidth utilization and packet loss rate are low.

[0108] By calculating the Pareto optimal solution, we found that choosing path P 2 It can achieve the best balance between latency and packet loss rate and transmit data according to this path. At this time, the network traffic can be adjusted to path P 2 , and optimizes the devices and links in the path based on load evaluation to ensure network quality.

[0109] This step achieves a balance between multiple performance objectives through multi-objective optimization and Pareto optimal solution calculation, ensuring network optimization under different quality indicators. This not only helps the network locate the source of abnormalities and improve the stability and robustness of the network, but also provides actionable suggestions for performance optimization in actual network environments.

[0110] Generally speaking, this method can effectively find the best balance among multiple quality objectives, avoid the adverse side effects brought by a single optimization objective, improve the utilization efficiency of network resources, and ensure the efficient operation of the network.

[0111] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for locating the cause of abnormal network dialing quality indicators, characterized in that: The method comprises the following steps: Multi-source data collection and network topology modeling: By collecting the performance data of each device in the network and the traffic data of the link, the status information of the network device and link is generated, and the network topology model is established; Shortest path calculation and abnormal path identification based on graph theory: Use the shortest path algorithm in graph theory to calculate the shortest path from the source device to the target device and identify abnormal links and devices in the path; Load evaluation and bottleneck analysis based on information theory: Information entropy is calculated based on the performance status of devices and links to evaluate the load of devices or links and identify performance bottlenecks in the network; Resource allocation and optimization based on game theory: Use game theory models to optimize the resources of devices and links in the network, avoid resource conflicts in the network, and improve network performance; Dynamic path adjustment and load optimization based on deep reinforcement learning: Through deep reinforcement learning training, network path selection and load distribution strategies are dynamically adjusted to optimize network performance and reduce quality anomalies; Multi-objective optimization and comprehensive anomaly location: A multi-objective optimization algorithm is used to balance and optimize multiple quality indicators, locate the causes of anomalies in the network, and provide network configuration optimization suggestions.

2. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The multi-source data collection and network topology modeling further include: Collect real-time performance data of each device through network management protocols, including CPU utilization, memory usage and motherboard temperature; Collect link bandwidth utilization, packet loss rate, and latency data through network traffic monitoring protocols; Based on the collected device and link data, a network topology diagram is constructed, the devices and links are abstracted as nodes and edges in the diagram, and each link is assigned a corresponding quality indicator weight.

3. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The shortest path calculation and abnormal path identification based on graph theory further include: For each link, its weight is calculated based on its latency, packet loss rate, and bandwidth performance indicators. The link weight is obtained through a weighted calculation formula. The Dijkstra algorithm is used to calculate the shortest path from the source device to the target device, identify potential abnormal paths, and make abnormal judgments based on the weights of the links in the path.

4. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The load evaluation and bottleneck analysis based on information theory further includes: By calculating the load status information entropy of the device or link, the current load level is evaluated; By analyzing the redundancy of devices or links, possible performance bottlenecks are identified. If the redundancy of a device or link is lower than a preset threshold, it is considered a bottleneck.

5. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The resource allocation and optimization based on game theory further includes: The resource allocation model between devices is constructed through the framework of game theory. The device strategies include adjusting bandwidth allocation and load balancing. Solve Nash equilibrium and determine the optimal resource allocation strategy for each device under the condition that the strategies of other devices are fixed, so as to optimize the use of network resources.

6. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The dynamic path adjustment and load optimization based on deep reinforcement learning further includes: The state space is defined to include the performance status of the device and the traffic status of the link, and the action space includes path selection and load adjustment; Design a reward function to evaluate the reward for each action based on the network's latency and packet loss rate, and encourage the selection of actions that reduce network quality anomalies; The Q-learning algorithm is used to train the intelligent agent, and the path selection and load distribution strategies are continuously optimized based on historical network status and action feedback.

7. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The multi-objective optimization and comprehensive anomaly positioning further include: Construct a multi-objective optimization function that comprehensively considers multiple network quality indicators, with the goal of minimizing the overall performance loss of the network; Through the multi-objective optimization algorithm, the Pareto optimal solution is calculated and an optimal solution is selected that does not harm other objectives.

8. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The shortest path calculation further comprises: When calculating the shortest path, the link load and real-time performance changes are taken into account and the path selection is dynamically adjusted.

9. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The game model establishment further includes: Define the strategies of devices in the game, including bandwidth allocation, data forwarding strategy, and link selection, to reduce network resource conflicts.

10. The method for locating the cause of abnormal network dialing quality indicators according to claim 1, characterized in that: The deep reinforcement learning training further includes: The training process of deep reinforcement learning is optimized, and a more efficient learning rate adjustment method and reward mechanism are adopted to ensure that the intelligent agent can quickly adapt to changes in the network environment and optimize path selection and load distribution.

Citation Information

Cited By

  • Automatic oiling machine communication network intelligent diagnosis system

    CN120186004A

  • An intelligent diagnostic system for the communication network of an automatic fuel dispenser

    CN120186004B

  • Unmanned aerial vehicle cluster low-cost control system and method based on SDN and NFV

    CN120802723A

  • Network testing method and equipment for computing power cluster

    CN121217615A

  • Network testing method and device of computing power cluster

    CN121217615B