SDN intelligent delay monitoring system and method

By introducing an intelligent module to the SDN control plane, optimizing the detection interval and monitoring strategy, the problem of inability to weigh the control plane overhead and delay monitoring accuracy in the SDN network is solved, and efficient delay monitoring and network optimization are achieved.

CN120301799APending Publication Date: 2025-07-11QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120474.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing software-defined network (SDN) delay monitoring methods increase network size, control plane overhead and delay monitoring accuracy cannot be traded down, resulting in increased measurement errors and inability to intelligently learn network behavior.

Method used

The intelligent module is introduced to the control plane, and the switch and link delays are calculated through the delay measurement module, combined with the data conversion module and the data storage module, and the Q-learning algorithm is used to optimize the detection interval and monitoring strategy to reduce the control plane overhead and switch delay.

Benefits of technology

While ensuring network link delay monitoring accuracy, it reduces control plane overhead and switch delay, and can quickly adjust policies to deal with network status abnormalities, optimize routing and improve data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301799A_ABST
    Figure CN120301799A_ABST
Patent Text Reader

Abstract

The invention discloses an SDN (Software Defined Network) intelligent delay monitoring system and method, belongs to the technical field of network monitoring, and aims to solve the technical problem that control plane overhead and delay monitoring precision in a delay monitoring method cannot be balanced along with increase of a network scale. Comprising a delay measurement module which is used for obtaining a detection interval from a data storage module, carrying out delay monitoring on a switch and each link of a data plane based on the detection interval, and sending switch time delay, link time delay of each link and CPU overhead in a network to a data conversion module; the data conversion module is used for converting CPU overhead generated by current delay monitoring, switch delay in a network and a ratio of link delay of m links with the highest delay to a detection interval into a state required by the agent module, and sending the state to the data storage module for storage; an action set and two initialized Q tables are arranged in the agent module, and the agent module is used for executing decision learning and state updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network monitoring, and in particular to an SDN intelligent delay monitoring system and method. Background Art

[0002] Software Defined Network (SDN) is a new architecture that separates the data plane from the control plane. Such a characteristic endows it with more network programmability, serviceability, heterogeneity, and maintainability. For this reason, this architecture has received increasing attention from researchers and scholars.

[0003] Monitoring the delay in the network is very important. It can assist in finding network congestion, better help managers optimize traffic scheduling and load balancing, and make full use of network resources. Delay monitoring can also determine the fault points in the network, help managers quickly detect and repair the fault points in a timely manner, and SDN brings new opportunities for delay measurement.

[0004] However, the current delay monitoring methods for software defined networks are usually based on the control plane, injecting timestamp packets as probe packets to measure the network delay at specific time points. But when continuous network delay monitoring is required, there are the following disadvantages: the increase in control plane overhead and the increase in switch delay. As the network scale grows, the measurement error of using OpenFlow messages to measure the time from the controller to the switch becomes larger and larger. These methods cannot well balance the relationship between overhead and measurement accuracy, and the control plane in these methods cannot intelligently learn the network behavior to weigh the relationship between this overhead and measurement accuracy.

[0005] The problem that the control plane overhead and the delay monitoring accuracy cannot be weighed in the delay monitoring method as the network scale increases is a technical problem that needs to be solved. Summary of the Invention

[0006] The technical task of the present invention is to provide an SDN intelligent delay monitoring system and method for the above deficiencies to solve the technical problem that the control plane overhead and the delay monitoring accuracy cannot be weighed in the delay monitoring method as the network scale increases.

[0007] In a first aspect, an SDN intelligent delay monitoring system of the present invention includes a delay measurement module, a data conversion module, an agent module, and a data storage module configured in an SDN controller;

[0008] An initialized detection interval is configured in the data storage module;

[0009] The delay measurement module is used to obtain the detection interval from the data storage module, monitor the delay of the switches in the data plane and each link based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module;

[0010] The data conversion module is used to convert the CPU overhead generated by the current delay monitoring, the switch delay in the network, and the ratio of the link delay of the m links with the highest delay to the detection interval into the state required by the agent module, and send the state to the data storage module for storage;

[0011] An action set and two initialized Q-tables are set in the agent module, which are used to perform the following:

[0012] Decision-making and learning: Based on the current state and policy, select an action, execute the action in the environment, obtain the next state and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q-tables with the same probability based on the reward value. Among them, when selecting an action based on the current state and policy, if an action is selected based on one of the Q-tables, the Q-value of the other Q-table is updated;

[0013] State update: Use the next state as the current state and enter the next decision-making and learning cycle until the termination condition is met. Among them, if the next state does not meet the preset conditions, reset the detection interval and send the newly set detection interval to the data storage module.

[0014] Preferably, the delay measurement module is used to send Packet-out messages and ECHO-REQUEST messages to all switches in the data plane based on the detection interval and receive the Packet-in messages and ECHO-RESPONSE messages returned by each switch, calculate the switch delay in the network based on the time difference between the Packet-out message and the Packet-in message and the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message, and send Packet-out messages to all links in the network topology based on the detection interval and receive the Packet-in messages returned by each link, calculate the link delay of each link in the network based on the time difference between the Packet-out message and the Packet-in message.

[0015] Preferably, the switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. The delay measurement module is used to calculate the switch delay by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows:

[0016] The time when the delay measurement module sends a data packet to the switch through a Packet-out message is T a , the switch processes the received data packet and returns the data packet to the delay measurement module through a Packet-in message. The time when the delay measurement module receives the data packet is T b , and the total delay U is recorded as U = T b -T a ;

[0017] The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c , the delay measurement module receives the ECHO-RESPONSE message returned by the switch and records the current time as T d , and the delay C between the controller and the switch is C = T d -T c , where the SDN controller and the switch are time-synchronized;

[0018] Calculate the switch delay S = U - C;

[0019] Among them, the delay measurement module is used to perform the following calculation of the link delay:

[0020] The delay measurement module sends a Packet-out message to one of the switches s i and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After the delay measurement module receives the Packet-in message, it records the current receiving time and calculates the difference T between the sending time and the receiving time csi , where s i represents the i-th switch;

[0021] The delay measurement module sends a Packet-out message to one of the switches s j and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After the delay measurement module receives the Packet-in message, it records the current receiving time and calculates the difference T between the sending time and the receiving time csj , where s j represents the j-th switch;

[0022] The delay measurement module sends a Packet - out message to one of the switches s i and records the current sending time. After the switch s i receives the Packet - out message, it forwards it to the switch s according to the routing table j . After the switch s j receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the returned Packet - in message, it records the current reception time. The total time from the delay measurement module to the switch s i , then to the switch s j and back to the delay measurement module is T;

[0023] The link delay D i between the switch s j and the switch s ij is calculated by the formula:

[0024]

[0025] Preferably, the agent module is used to select the most valuable action with a probability of 1 - ε to perform delay monitoring. Among them, the most valuable action refers to the action with the largest Q - value corresponding to the addition of two Q - tables in the current same state. The value range of ε is [0,1].

[0026] Preferably, the agent module is used to evaluate the reward value of an action based on the following reward rules: If the change in the delay monitoring strategy can reduce the overhead of the SDN controller during the monitoring process, the reward value is positive, otherwise the reward value is positive or negative; If the change in the delay monitoring strategy can reduce the switch delay, the reward value is positive, otherwise the reward value is negative; If the change in the strategy can make the ratio of the link delay value to the detection interval larger, the reward value is positive, otherwise it is negative; The change in other strategies is determined according to the specific situation;

[0027] Among them, the calculation formula of the reward value is as follows:

[0028]

[0029] Among them, V represents the ratio of the monitored link delay value in the network to the detection interval, S represents the switch delay, U represents the CPU overhead caused by the SDN controller due to delay monitoring, and represents the weight, D ij represents the link delay of the monitored link, and t represents the detection interval;

[0030] Based on the reward value, when updating the Q1 table and the Q2 table with the same probability, if the Q1 table is updated, the update formula is as follows:

[0031] a * = argmax a Q1(s ′ , a),

[0032] Q1(s, a) ← Q1(s, a) + α(R + γQ2(s ′ , a * ) - Q1(s, a)),

[0033] where a * represents the most valuable action in Q1 table for state s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in Q1 table for state s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to state s ′ and action a in Q1 table, Q1(s, a) represents the return value corresponding to state s and action a in Q1 table, Q2(s ′ , a * ) represents the return value corresponding to state s ′ and action a * in Q2 table;

[0034] If Q2 table is updated, the update formula is as follows:

[0035] a * = argmax a Q2(s ′ , a),

[0036] Q2(s, a) ← Q2(s, a) + α(R + γQ1(s ′ , a * ) - Q2(s, a)),

[0037] where a * represents the most valuable action in Q2 table for state s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in Q2 table for state s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ , a) represents the return value corresponding to state s ′ and action a in Q2 table, Q2(s, a) represents the return value corresponding to state s and action a in Q2 table, Q1(s ′ , a * ) represents the return value corresponding to state s ′ and action a * in Q1 table.

[0038] In a second aspect, an SDN intelligent delay monitoring method of the present invention is used for delay monitoring through an SDN intelligent delay monitoring system as described in any item of the first aspect. The method includes the following steps:

[0039] Configure an initialized detection interval in the data storage module;

[0040] Obtain the detection interval from the data storage module through the delay measurement module, perform delay monitoring on the switches in the data plane and each link based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module;

[0041] Convert the CPU overhead generated by the current delay monitoring, the switch delay in the network, and the ratio of the link delay of the m links with the highest delay to the detection interval into the state required by the agent module through the data conversion module, and send the state to the data storage module for storage;

[0042] Set an action set and two initialized Q tables in the agent module, and execute decision learning and state update through the agent module;

[0043] Among them, decision learning: select an action based on the current state and policy, execute the action in the environment, obtain the next state and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q tables with the same probability based on the reward value. When selecting an action based on the current state and policy, if an action is selected based on one of the Q tables, the Q value of the other Q table is updated;

[0044] State update: Use the next state as the current state and enter the next decision learning cycle until the termination condition is met. If the next state does not meet the preset condition, reset the detection interval and send the newly set detection interval to the data storage module.

[0045] Preferably, based on the detection interval, the delay measurement module sends ECHO-REQUEST messages to all switches in the data plane and receives the ECHO-RESPONSE messages returned by each switch. The switch delay in the network is calculated based on the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message. Packet-out messages are sent to all links in the network topology based on the detection interval, and Packet-in messages returned by each link are received. The link delay of each link in the network is calculated based on the time difference between the Packet-out message and the Packet-in message.

[0046] Preferably, the switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. The switch delay is calculated by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows:

[0047] The time when the delay measurement module sends a data packet to the switch through the Packet-out message is T a , the switch processes the received data packet and returns the data packet to the delay measurement module through the Packet-in message. The time when the delay measurement module receives the data packet is T b , and the total delay U is recorded as U = T b -T a ;

[0048] The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c , the delay measurement module receives the ECHO-RESPONSE message returned by the switch and records the current time as T d , and the delay C between the controller and the switch is C = T d -T c , where the SDN controller and the switch are time-synchronized;

[0049] Calculate the switch delay S = U - C;

[0050] Among them, the following is used to calculate the link delay:

[0051] The delay measurement module sends a Packet-out message to one of the switches s i and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After the delay measurement module receives the Packet-in message, it records the current receiving time and calculates the difference T between the sending time and the receiving time csi , where s i represents the i-th switch;

[0052] The delay measurement module sends a Packet-out message to one of the switches s j and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time csj , where s j represents the j-th switch;

[0053] The delay measurement module sends a Packet-out message to one of the switches s i and records the current sending time. After the switch s i receives the Packet-out message, it forwards it to the switch s according to the routing table j . After the switch s j receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message returned to the delay measurement module, the delay measurement module records the current acceptance time. The total time from the delay measurement module to the switch s i , then to the switch s j , and back to the delay measurement module is T;

[0054] The link delay D i between the switch s j and the switch s ij is calculated as follows:

[0055]

[0056] Preferably, the agent module selects the most valuable action with a probability of 1 - ε to perform delay monitoring. Among them, the most valuable action refers to the action with the largest Q value corresponding to the addition of two Q tables in the current same state. The value range of ε is [0,1].

[0057] Preferably, the reward value of the action is evaluated based on the following reward rules: If the change in the delay monitoring strategy can reduce the SDN controller overhead during the monitoring process, the reward value is positive, otherwise the reward value is positive or negative; If the change in the delay monitoring strategy can reduce the switch delay, the reward value is positive, otherwise the reward value is negative; If the change in the strategy can make the ratio of the link delay value to the detection interval larger, the reward value is positive, otherwise it is negative; The changes in other strategies are determined according to specific situations;

[0058] Among them, the calculation formula of the reward value is as follows:

[0059]

[0060] Wherein, V represents the ratio of the monitored link delay value to the probing interval in the network, S represents the switch delay, and U represents the CPU overhead of the SDN controller caused by delay monitoring. And represents the weight, D ij represents the link delay of the monitored link, and t represents the probing interval.

[0061] When updating the Q1 table and the Q2 table with the same probability based on the reward value, if the Q1 table is updated, the update formula is as follows:

[0062] a * = argmax a Q1(s ′ , a),

[0063] Q1(s, a) ← Q1(s, a) + α(R + γQ2(s ′ , a * )) - Q1(s, a)),

[0064] Wherein, a * represents the most valuable action in the Q1 table for state s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in the Q1 table for state s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to state s ′ and action a in the Q1 table, Q1(s, a) represents the return value corresponding to state s and action a in the Q1 table, and Q2(s ′ , a * ) represents the return value corresponding to state s ′ and action a * in the Q2 table;

[0065] If the Q2 table is updated, the update formula is as follows:

[0066] a * = argmax a Q2(s ′ , a),

[0067] Q2(s, a) ← Q2(s, a) + α(R + γQ1(s ′ , a * ) - Q2(s, a)),

[0068] Wherein, a * represents the most valuable action in the Q2 table for state s ′ , s ′Represents the next state, s represents the current state, a represents the most valuable action in state s of Q2 table, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ ,a) represents the return value corresponding to state s ′ and action a in the Q2 table, Q2(s,a) represents the return value corresponding to state s and action a in the Q2 table, Q1(s ′ ,a * ) represents the return value corresponding to state s ′ and action a * in the Q1 table.

[0069] The SDN intelligent delay monitoring system and method of the present invention have the following advantages: introducing an agent module into the control plane, continuously interacting with the network according to the link delay in the data plane and the overhead in the control plane, intelligently monitoring the link delay in the network under the condition of ensuring no adverse impact on the overall network operation, reducing the control plane overhead and switch delay in the delay monitoring process while ensuring the monitoring accuracy of the link delay in the network, and being able to quickly adjust the delay monitoring strategy when the network state is abnormal, providing effective help for quickly finding congestion in the network, optimizing routing, and improving data transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0071] The present invention will be further described below with reference to the drawings.

[0072] Figure 1 It is a structural block diagram of an SDN intelligent delay monitoring system for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0074] The embodiments of the present invention provide an SDN intelligent delay monitoring system and method, which are used to solve the technical problem that the control plane overhead and the delay monitoring accuracy cannot be balanced in the delay monitoring method as the network scale increases.

[0075] Embodiment 1:

[0076] An SDN intelligent delay monitoring system of the present invention includes a delay measurement module, a data conversion module, an agent module, and a data storage module configured in an SDN controller.

[0077] An initialized detection interval is configured in the data storage module.

[0078] The delay measurement module is used to obtain the detection interval from the data storage module, perform delay monitoring on the switches in the data plane and each link based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module.

[0079] As a specific implementation of the delay measurement module, this module is used to send Packet-out messages and ECHO-REQUEST messages to all switches in the data plane based on the detection interval and receive the Packet-in messages and ECHO-RESPONSE messages returned by each switch, calculate the switch delay in the network based on the time difference between the Packet-out message and the Packet-in message and the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message, and send Packet-out messages to all links in the network topology based on the detection interval and receive the Packet-in messages returned by each link, calculate the link delay of each link in the network based on the time difference between the Packet-out message and the Packet-in message.

[0080] The switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. The delay measurement module is used to calculate the switch delay by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows:

[0081] (1) The time when the delay measurement module sends a data packet to the switch through the Packet-out message is T a , the switch receives the data packet for processing and returns the data packet to the delay measurement module through the Packet-in message. The time when the delay measurement module receives the data packet is T b , and the total delay U is recorded as U = T b -T a ;

[0082] (2) The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c, the delay measurement module receives the ECHO - RESPONSE message returned by the switch and records the current time as T d , the delay C between the controller and the switch is C = T d -T c , where the SDN controller and the switch are time - synchronized;

[0083] (3) Calculate the switch delay S = U - C.

[0084] Among them, the delay measurement module is used to perform the following calculation of the link delay:

[0085] (1) The delay measurement module sends a Packet - out message to one of the switches s i and records the current sending time. When the switch receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the Packet - in message, it records the current receiving time and calculates the difference T between the sending time and the receiving time csi , where s i represents the i - th switch;

[0086] (2) The delay measurement module sends a Packet - out message to one of the switches s j and records the current sending time. When the switch receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the Packet - in message, it records the current receiving time and calculates the difference T between the sending time and the receiving time csj , where s j represents the j - th switch;

[0087] (3) The delay measurement module sends a Packet - out message to one of the switches s i and records the current sending time. After the switch s i receives the Packet - out message, it forwards it to the switch s according to the routing table j , after the switch s j receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the Packet - in message returned to it, it records the current acceptance time. The total time from the delay measurement module to the switch s i , then to the switch s j and back to the delay measurement module is T;

[0088] (4) The link delay D between the switch s i and the switch s j ​ij The calculation formula is expressed as:

[0089]

[0090] The data conversion module is used to convert the CPU overhead generated by the current latency monitoring, the switch latency in the network, and the ratio of the latency of the m links with the highest latency to the detection interval into the states required by the agent module, and send the states to the data storage module for storage.

[0091] The agent module is set with an action set and two initialized Q-tables, and is used to perform the following:

[0092] (1) Decision-making learning: Based on the current state and policy, select an action and execute the action in the environment to obtain the state at the next moment and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q-tables with the same probability based on the reward value. Among them, when selecting an action based on the current state and policy, if an action is selected based on one of the Q-tables, the Q-value of the other Q-table is updated;

[0093] (2) State update: Use the state at the next moment as the current state and enter the next decision-making learning cycle until the termination condition is met. Among them, if the state at the next moment does not meet the preset conditions, reset the detection interval and send the newly set detection interval to the data storage module.

[0094] As a specific implementation of the agent module, the action set is set to {0.2, 0, -0.2}. The elements in this set can be selected multiple times. "0.2" in the set means that the duration of latency monitoring for the link is increased by 0.2 seconds, "0" means maintaining the current latency monitoring duration, and "-0.2" means that the latency monitoring duration is shortened by 0.2 seconds.

[0095] In specific implementation, the agent module randomly selects an action from the action set to perform latency monitoring based on the probability of ε, and selects the most valuable action to perform latency monitoring based on the probability of 1 - ε. Among them, the most valuable action refers to the action corresponding to the largest Q-value obtained by adding the two Q-tables in the current same state. The value range of ε is [0, 1].

[0096] In this embodiment, the agent module is used to evaluate the reward value of the action based on the following reward rules: If the change in the latency monitoring policy can reduce the SDN controller overhead during the monitoring process, the reward value is positive, otherwise the reward value is positive or negative; if the change in the latency monitoring policy can reduce the switch latency, the reward value is positive, otherwise the reward value is negative; if the change in the policy can make the ratio of the link latency value to the detection interval larger, the reward value is positive, otherwise it is negative; the changes in other policies are determined according to specific circumstances.

[0097] Among them, the calculation formula of the reward value is as follows:

[0098]

[0099] Among them, V represents the ratio of the monitored link delay value in the network to the detection interval, S represents the switch delay, U represents the CPU overhead caused by the SDN controller due to delay monitoring, and represents the weight, D ij represents the link delay of the monitored link, and t represents the detection interval;

[0100] When updating the Q1 table and the Q2 table with the same probability based on the reward value, if the Q1 table is updated, the update formula is as follows:

[0101] a * = argmax a Q1(s ′ , a),

[0102] Q1(s, a) ← Q1(s, a) + α(R + γQ2(s ′ , a * ) - Q1(s, a)),

[0103] where a * represents the most valuable action in the Q1 table in the s ′ state, s ′ represents the next state, s represents the current state, a represents the most valuable action in the Q1 table in the s state, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to the state s ′ and the action a in the Q1 table, Q1(s, a) represents the return value corresponding to the state s and the action a in the Q1 table, Q2(s ′ , a * ) represents the return value corresponding to the state s ′ and the action a * in the Q2 table;

[0104] If the Q2 table is updated, the update formula is as follows:

[0105] a * = argmax a Q2(s ′ , a),

[0106] Q2(s, a) ← Q2(s, a) + α(R + γQ1(s ′ , a * ) - Q2(s, a)),

[0107] where a* Represents the Q2 table s ′ The most valuable action in the state, s ′ Represents the next state, s represents the current state, a represents the most valuable action in the Q2 table s state, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ ,a) represents the return value corresponding to state s in the Q2 table ′ And the return value corresponding to action a, Q2(s,a) represents the return value corresponding to state s and action a in the Q2 table, Q1(s ′ ,a * ) represents the return value corresponding to state s in the Q1 table ′ And action a * The corresponding return value.

[0108] In this embodiment, the agent module adjusts the delay monitoring strategy in the network, considering the following:

[0109] (1) When performing delay monitoring, the overhead of the control plane in the software-defined network architecture should be considered, aiming to reduce the overhead;

[0110] (2) When performing delay monitoring, the change in switch delay during the delay process should be considered. Excessive probe messages will cause the switch delay to increase, aiming to reduce the switch delay:

[0111] (3) When performing delay monitoring, the ratio of the link delay value to the probe interval should be considered, and as much maximum benefit as possible should be obtained under the hardware performance limit.

[0112] That is, when adjusting the delay monitoring scheme, the load situation of the control plane should be considered, because excessive load will affect the overall performance of the network; secondly, the switch delay should be considered, because frequent sending of probe messages will cause the switch delay to increase and affect the performance of the switch; finally, the ratio of the link delay value to the probe interval should also be considered to be proportional.

[0113] Considering the above requirements, the states required by the agent in the delay monitoring scheme should be set to "controller load in the control plane", "switch delay", and "ratio of the delay values of the n links with the largest delay values in the network topology to the probe interval".

[0114] The agent will learn an approximate optimal solution for adjusting the delay monitoring scheme due to traffic fluctuations in the network. While ensuring the accuracy of link delay monitoring in the network, it reduces the control plane overhead and switch delay during the delay monitoring process, and can also quickly adjust the delay monitoring strategy when the network state is abnormal, providing effective help for quickly finding congestion in the network, optimizing routing, and improving data transmission efficiency.

[0115] Introduce the agent decision-making module into the control plane. According to the link delay in the data plane and the overhead of the control plane, continuously interact with the network, and intelligently monitor the link delay in the network under the condition of ensuring that it will not have an adverse impact on the overall network operation. While ensuring the monitoring accuracy of the link delay in the network, it reduces the control plane overhead and switch delay during the delay monitoring process.

[0116] Embodiment 2:

[0117] A method for intelligent delay monitoring of SDN in the present invention performs delay monitoring through the system disclosed in Embodiment 1. The method includes the following steps:

[0118] Step S100: Configure an initialized detection interval in the data storage module;

[0119] Step S200: Obtain the detection interval from the data storage module through the delay measurement module, perform delay monitoring on the switches and each link in the data plane based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module;

[0120] Step S300: Convert the CPU overhead generated by the current delay monitoring, the switch delay in the network, and the ratio of the link delay of the m links with the highest delay to the detection interval into the state required by the agent module through the data conversion module, and send the state to the data storage module for storage;

[0121] Step S400: Set an action set and two initialized Q tables in the agent module, and perform decision learning and state update through the agent module;

[0122] Among them, decision learning: Select an action based on the current state and policy, execute the action in the environment, obtain the next state and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q tables with the same probability based on the reward value. Among them, when selecting an action based on the current state and policy, if an action is selected based on one of the Q tables, the Q value of the other Q table is updated;

[0123] State update: Use the next state as the current state and enter the next decision learning cycle until the termination condition is met. Among them, if the next state does not meet the preset condition, reset the detection interval and send the newly set detection interval to the data storage module.

[0124] In step S200 of this embodiment, based on the detection interval, Packet-out messages and ECHO-REQUEST messages are sent to all switches in the data plane, and Packet-in messages and ECHO-RESPONSE messages returned by each switch are received. The switch delay in the network is calculated based on the time difference between the Packet-out message and the Packet-in message and the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message. Based on the detection interval, Packet-out messages are sent to all links in the network topology, and Packet-in messages returned by each link are received. The link delay of each link in the network is calculated based on the time difference between the Packet-out message and the Packet-in message.

[0125] The switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. In this embodiment, the switch delay is calculated by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows:

[0126] (1) The time when the delay measurement module sends a data packet to the switch through the Packet-out message is T a , the switch receives the data packet for processing and returns the data packet to the delay measurement module through the Packet-in message. The time when the delay measurement module receives the data packet is T b , and the total delay U is recorded as U = T b -T a ;

[0127] (2) The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c , the delay measurement module receives the ECHO-RESPONSE message returned by the switch and records the current time as T d , and the delay C between the controller and the switch is C = T d -T c , where the SDN controller and the switch are time-synchronized;

[0128] (3) Calculate the switch delay S = U - C.

[0129] Among them, the delay measurement module is used to perform the following calculation of the link delay:

[0130] (1) The delay measurement module sends a packet to one of the switches s iSend a Packet-out message and record the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time. csi , where, s i represents the i-th switch;

[0131] (2) The delay measurement module sends a Packet-out message to one of the switches s j Send a Packet-out message and record the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time. csj , where, s j represents the j-th switch;

[0132] (3) The delay measurement module sends a Packet-out message to one of the switches s i Send a Packet-out message and record the current sending time. After switch s i receives the Packet-out message, it forwards it to switch s j according to the routing table. After switch s j receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message returned to the delay measurement module, the delay measurement module records the current receiving time. The total time from the delay measurement module to switch s i , then to switch s j , and back to the delay measurement module is T;

[0133] (4) The link delay D i between switch s j and switch s ij is calculated as:

[0134]

[0135] In step S400 of this embodiment, the action set is set to {0.2, 0, -0.2}. The elements in this set can be selected multiple times. "0.2" in the set means that the duration of link delay monitoring is increased by 0.2 seconds, "0" means maintaining the current delay monitoring duration, and "-0.2" means that the duration of delay monitoring is shortened by 0.2 seconds.

[0136] During specific implementation, the agent module randomly selects an action from the action set to perform delay monitoring with a probability of ε, and selects the most valuable action to perform delay monitoring with a probability of 1 - ε. Among them, the most valuable action refers to the action with the largest corresponding Q value obtained by adding the two Q tables in the current same state. The value range of ε is [0, 1].

[0137] In this embodiment, the reward value of the action is evaluated based on the following reward rules: If the policy change of delay monitoring can reduce the overhead of the SDN controller during the monitoring process, the reward value is positive; otherwise, the reward value is positive or negative; if the policy change of delay monitoring can reduce the switch delay, the reward value is positive; otherwise, the reward value is negative; if the policy change can increase the ratio of the link delay value to the detection interval, the reward value is positive; otherwise, it is negative; the changes of other policies are determined according to specific situations.

[0138] Among them, the calculation formula of the reward value is as follows:

[0139]

[0140] Among them, V represents the ratio of the monitored link delay value to the detection interval in the network, S represents the switch delay, U represents the CPU overhead caused by the SDN controller due to delay monitoring, and represents the weight, D ij represents the link delay of the monitored link, and t represents the detection interval;

[0141] Based on the reward value, when updating the Q1 table and the Q2 table with the same probability, if the Q1 table is updated, the update formula is as follows:

[0142] a * = argmax a Q1(s ′ , a),

[0143] Q1(s, a) ← Q1(s, a) + α(R + γQ2(s ′ , a * )) - Q1(s, a)),

[0144] Among them, a * represents the most valuable action in the Q1 table in the s ′ state, s ′ represents the next state, s represents the current state, a represents the most valuable action in the Q1 table in the s state, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to the state s ′ and the action a in the Q1 table, Q1(s, a) represents the return value corresponding to the state s and the action a in the Q1 table, Q2(s ′,a * ) represents the state s in the Q2 table ′ and the action a * for the corresponding return value;

[0145] If the Q2 table is updated, the update formula is as follows:

[0146] a * = argmax a Q2(s ′ ,a),

[0147] Q2(s,a) ← Q2(s,a) + α(R + γQ1(s ′ ,a * )) - Q2(s,a)),

[0148] where a * represents the most valuable action in the Q2 table for state s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in the Q2 table for state s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ ,a) represents the return value corresponding to state s ′ and action a in the Q2 table, Q2(s,a) represents the return value corresponding to state s and action a in the Q2 table, Q1(s ′ ,a * ) represents the return value corresponding to state s ′ and action a * in the Q1 table.

[0149] The above has introduced the SDN intelligent delay monitoring system and method provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An SDN intelligent delay monitoring system, characterized in that, It includes a delay measurement module, a data conversion module, an agent module, and a data storage module configured in the SDN controller; The data storage module is configured with an initialized detection interval; The delay measurement module is used to obtain the detection interval from the data storage module, perform delay monitoring on the switches in the data plane and each link based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module; The data conversion module is used to convert the CPU overhead generated by the current delay monitoring, the switch delay in the network, and the ratio of the link delay of the m links with the highest delay to the detection interval into the state required by the agent module, and send the state to the data storage module for storage; The agent module is set with an action set and two initialized Q tables, and is used to perform the following: Decision learning: Based on the current state and policy, select an action and execute the action in the environment to obtain the next state and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q tables with the same probability based on the reward value. When selecting an action based on the current state and policy, if an action is selected based on one of the Q tables, the Q value of the other Q table is updated; State update: Use the next state as the current state and enter the next decision learning cycle until the termination condition is met. If the next state does not meet the preset condition, reset the detection interval and send the newly set detection interval to the data storage module.

2. The SDN intelligent delay monitoring system according to claim 1, wherein The delay measurement module is used to send Packet-out messages and ECHO-REQUEST messages to all switches in the data plane based on the detection interval and receive the Packet-in messages and ECHO-RESPONSE messages returned by each switch, calculate the switch delay in the network based on the time difference between the Packet-out message and the Packet-in message and the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message, and send Packet-out messages to all links in the network topology based on the detection interval and receive the Packet-in messages returned by each link, calculate the link delay of each link in the network based on the time difference between the Packet-out message and the Packet-in message.

3. The SDN intelligent delay monitoring system according to claim 1 or 2, characterized in that, The switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. The delay measurement module is used to calculate the switch delay by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows: The time when the delay measurement module sends a data packet to the switch through a Packet-out message is T a , the switch receives the data packet for processing, returns the data packet to the delay measurement module through a Packet-in message, and the time when the delay measurement module receives the data packet is T b , the total delay U is recorded as U = T b -T a ; The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c , the delay measurement module receives the ECHO-RESPONSE message returned by the switch and records the current time as T d , the latency C between the controller and the switch is C = T d -T c , where the SDN controller is time-synchronized with the switch; Calculate the switch delay S = U - C; Among them, the delay measurement module is used to perform the following calculation of the link delay: The delay measurement module sends a Packet-out message to one of the switches s i and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time csi , where s i represents the i-th switch; The delay measurement module sends a Packet-out message to one of the switches s j and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time csj , where s j represents the j-th switch; The delay measurement module sends a Packet - out message to one of the switches s i and records the current sending time. After switch s i receives the Packet - out message, it forwards it to switch s j according to the routing table. After switch s j receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the returned Packet - in message, it records the current reception time. The total time from the delay measurement module to switch s i , then to switch s j and back to the delay measurement module is T; Switch s i With Switch s j Link delay D between ij The calculation formula is expressed as:

4. The SDN intelligent delay monitoring system according to claim 3, characterized in that The agent module is used to select the most valuable action with a probability of 1 - ε to perform delay monitoring. Among them, the most valuable action refers to the action with the largest Q value corresponding to the sum of two Q - tables in the current same state. The value range of ε is [0, 1].

5. The SDN intelligent delay monitoring system according to claim 1, wherein The agent module is used to evaluate the reward value of an action based on the following reward rules: If the policy change of delay monitoring can reduce the SDN controller overhead during monitoring, the reward value is positive; otherwise, the reward value is positive or negative; if the policy change of delay monitoring can reduce the switch delay, the reward value is positive; otherwise, the reward value is negative; if the policy change can make the ratio of the link delay value to the detection interval larger, the reward value is positive; otherwise, it is negative; the change of other policies is determined according to specific situations; Among them, the calculation formula of the reward value is as follows: Among them, V represents the ratio of the monitored link delay value in the network to the detection interval, S represents the switch delay, and U represents the CPU overhead caused by the SDN controller due to delay monitoring. And represents the weight, D ij represents the link delay of the monitored link, and t represents the detection interval. Based on the reward value, when updating the Q1 table and the Q2 table with the same probability, if the Q1 table is updated, the update formula is as follows: a * = argmax a Q1(s ′ , a), Q1(s,a) ← Q1(s,a) + α(R + γQ2(s ′ ,a * ) - Q1(s,a)), Among them, a * represents the most valuable action in the state of Q1 table s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in the state of Q1 table s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to the state s ′ and the action a in the Q1 table, Q1(s, a) represents the return value corresponding to the state s and the action a in the Q1 table, Q2(s ′ , a * ) represents the return value corresponding to the state s ′ and the action a * in the Q2 table; If the Q2 table is updated, the update formula is as follows: a * = argmax a Q2(s ′ , a), Q2(s,a)←Q2(s,a)+α(R+γQ1(s ′ ,a * )-Q2(s,a)), where a * represents the most valuable action in the state of Q2 table s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in the state of Q2 table s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ , a ′ represents the return value corresponding to the state s ′ and the action a in the Q2 table, Q2(s, a) represents the return value corresponding to the state s and the action a in the Q2 table, Q1(s * , a ′ ) represents the return value corresponding to the state s * and the action a in the Q1 table.

6. An SDN intelligent delay monitoring method, characterized in that, It is used to perform delay monitoring through an SDN intelligent delay monitoring system as described in any one of claims 1 - 5. The method includes the following steps: Configure the initialized detection interval in the data storage module; Obtain the detection interval from the data storage module through the delay measurement module, perform delay monitoring on the switches in the data plane and each link based on the detection interval, calculate the switch delay in the network, the link delay of each link, and the CPU overhead of the SDN controller caused by the delay through the delay monitoring, and send the switch delay in the network, the link delay of each link, and the CPU overhead to the data conversion module; Through the data conversion module, convert the CPU overhead generated by the current delay monitoring, the switch delay in the network, and the ratio of the link delay of the m links with the highest delay to the detection interval into the state required by the agent module, and send the state to the data storage module for storage; Set an action set and two initialized Q - tables in the agent module, and execute decision - making learning and state update through the agent module; Among them, decision - making learning: Based on the current state and policy, select an action, execute the action in the environment, obtain the next - moment state and the reward value for evaluating the action, send the reward value to the data storage module for storage, and update the two Q - tables with the same probability based on the reward value. Among them, when selecting an action based on the current state and policy, if an action is selected based on one of the Q - tables, the Q value of the other Q - table is updated; State update: Use the next - moment state as the current state, enter the next decision - making learning cycle until the termination condition is met. Among them, if the next - moment state does not meet the preset condition, reset the detection interval and send the newly set detection interval to the data storage module.

7. The SDN intelligent delay monitoring method according to claim 6, characterized in that, Based on the detection interval, the Packet-out message and the ECHO-REQUEST message are sent to all switches in the data plane through the delay measurement module, and the Packet-in message and the ECHO-RESPONSE message returned by each switch are received. The switch delay in the network is calculated based on the time difference between the Packet-out message and the Packet-in message and the time difference between the ECHO-REQUEST message and the ECHO-RESPONSE message. The Packet-out message is sent to all links in the network topology based on the detection interval, and the Packet-in message returned by each link is received. The link delay of each link in the network is calculated based on the time difference between the Packet-out message and the Packet-in message.

8. The SDN intelligent delay monitoring method according to claim 6, characterized in that The switch delay in the network includes the queuing delay inside the switch and the processing delay of the switch. The switch delay is calculated by calculating the total delay and the delay between the controller and the switch. The calculation method is as follows: The time when the delay measurement module sends a data packet to the switch through a Packet-out message is T a , the switch receives the data packet for processing, returns the data packet to the delay measurement module through a Packet-in message, and the time when the delay measurement module receives the data packet is T b , the total delay U is recorded as U = T b -T a ; The delay measurement module sends an ECHO-REQUEST message to the switch in the data plane and records the current time T c , the delay measurement module receives the ECHO-RESPONSE message returned by the switch and records the current time as T d , the delay C between the controller and the switch is C = T d - T c , where the SDN controller and the switch are time synchronized; Calculate the switch delay S = U - C; Among them, the following is used to calculate the link delay: The delay measurement module sends a Packet-out message to one of the switches s i and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time csi , where s i represents the i-th switch; The delay measurement module sends a Packet-out message to one of the switches s j and records the current sending time. When the switch receives the Packet-out message, it returns a Packet-in message to the delay measurement module. After receiving the Packet-in message, the delay measurement module records the current receiving time and calculates the difference T between the sending time and the receiving time csj , where s j represents the j-th switch; The delay measurement module sends a Packet - out message to one of the switches s i and records the current sending time. After switch s i receives the Packet - out message, it forwards it to switch s according to the routing table j . After switch s j receives the Packet - out message, it returns a Packet - in message to the delay measurement module. After the delay measurement module receives the returned Packet - in message, it records the current reception time. The total time from the delay measurement module to switch s i , then to switch s j and back to the delay measurement module is T; Switch s i With switch s j Link delay D between ij The calculation formula is expressed as:

9. The SDN intelligent delay monitoring method according to claim 8, characterized in that, The agent module selects the most valuable action with a probability of 1 - ε to perform delay monitoring. Among them, the most valuable action refers to adding two Q-tables in the current same state, and the action corresponding to the largest Q value. The value range of ε is [0, 1].

10. The SDN intelligent delay monitoring method according to claim 6, wherein Evaluate the reward value of the action based on the following reward rules: If the policy change of the delay monitoring can reduce the SDN controller overhead during the monitoring process, the reward value is positive, otherwise the reward value is positive or negative; if the policy change of the delay monitoring can reduce the switch delay, the reward value is positive, otherwise the reward value is negative; if the policy change can make the ratio of the link delay value to the detection interval larger, the reward value is positive, otherwise it is negative; the changes of other policies are determined according to specific situations; Among them, the calculation formula of the reward value is as follows: Among them, V represents the ratio of the monitored link delay value to the detection interval in the network, S represents the switch delay, and U represents the CPU overhead caused by the SDN controller due to delay monitoring. and represents the weight, D ij represents the link delay of the monitored link, and t represents the detection interval. Based on the reward value, when updating the Q1 table and the Q2 table with the same probability, if the Q1 table is updated, the update formula is as follows: a * = argmax a Q1(s ′ , a), Q1(s,a)←Q1(s,a)+α(R+γQ2(s ′ ,a * )-Q1(s,a)), Among them, a * represents the most valuable action in the state of Q1 table s ′ The state, s ′ represents the next state, s represents the current state, a represents the most valuable action in the state of Q1 table s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q1(s ′ , a) represents the return value corresponding to the state s ′ and the action a in the Q1 table, Q1(s, a) represents the return value corresponding to the state s and the action a in the Q1 table, Q2(s ′ , a * ) represents the return value corresponding to the state s ′ and the action a * in the Q2 table; If the Q2 table is updated, the update formula is as follows: a * = argmax a Q2(s ′ , a), Q2(s,a) ← Q2(s,a) + α(R + γQ1(s ′ ,a * ) - Q2(s,a)), where a * represents the most valuable action in the state of Q2 table s ′ , s ′ represents the next state, s represents the current state, a represents the most valuable action in the state of Q2 table s, R represents the reward value, α represents the learning rate, γ represents the discount factor, Q2(s ′ , a) represents the return value corresponding to the state s ′ and the action a in the Q2 table, Q2(s, a) represents the return value corresponding to the state s and the action a in the Q2 table, Q1(s ′ , a * ) represents the return value corresponding to the state s ′ and the action a * in the Q1 table.