SD-WAN flow distribution method and system based on deep reinforcement learning
Through the deep reinforcement learning method, combined with active and passive detection technology, the SD-WAN traffic allocation is optimized, which solves the problem of low resource allocation efficiency in dynamic network environments, and improves the availability and reliability of the network.
Patent Information
- Application Number
- CN202510867303.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional SD-WAN traffic engineering methods are difficult to achieve efficient and adaptive resource allocation in dynamic network environments, especially in heterogeneous links and complex network states, resulting in network performance degradation.
Using a deep reinforcement learning method, the network basic delay and application layer delay data are collected through active and passive detection technology, the delay threshold is calculated using quartiles, reward function and channel switching model are constructed, and traffic allocation is optimized.
Improves the overall availability and service reliability of the SD-WAN network, reduces service downtime, and enables proactive adjustment of traffic paths to deal with network problems.
Smart Images

Figure CN120567786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic engineering technology, and in particular to a SD-WAN traffic distribution method and system based on deep reinforcement learning. Background Art
[0002] SD-WAN improves network flexibility through centralized control and multipath forwarding. However, dynamic network loads and heterogeneous links such as MPLS, 4G / 5G, and broadband make it difficult for traditional rule-based traffic engineering (such as static routing and heuristic algorithms) to achieve efficient and adaptive resource allocation. Traditional mathematical optimization methods, for example, are computationally complex and lack real-time response capabilities. Heuristic algorithms also struggle to adapt to complex network conditions. Summary of the Invention
[0003] The present invention provides an SD-WAN traffic distribution method based on deep reinforcement learning, which mainly includes:
[0004] Step S1: Collect network basic delay data and application layer delay data of two communication area edge devices of SD-WAN, where the two communication area edge devices are CPE and uCPE;
[0005] Step S2: calculating the actual delay data based on the network basic delay data and the application layer delay data and dividing the delay data into peak time periods and non-peak time periods;
[0006] Step S3: Calculate the delay thresholds of the peak period and the non-peak period using quartiles based on the basic network delay data;
[0007] Step S4: constructing a reward function using the real delay data and the delay threshold, and constructing a channel switching model based on the reinforcement learning architecture, environmental conditions, and the reward function;
[0008] Step S5: training and optimizing the channel switching model using network basic delay data, application layer delay data, real delay data, peak time periods and non-peak time periods to obtain a trained channel switching model;
[0009] Step S6: Input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
[0010] Optionally, in step S1, the two communication area edge devices are CPE and uCPE, and an active monitoring method is used to obtain network basic delay data, and a passive detection method is used to obtain application layer delay data.
[0011] Optionally, the active detection method is specifically:
[0012] Send a detection packet to the peer gateway at a preset sending period, calculate the delay time, and obtain the initial delay time data;
[0013] obtaining an original time series based on the initial delay time data;
[0014] Based on the sequence value of the original time series, the fluctuation value at time t is obtained using the exponentially weighted moving average method;
[0015] A fluctuation delay value is calculated based on the fluctuation value and recorded as network basic delay data.
[0016] Optionally, the passive detection method includes:
[0017] Extract the message features from layer 3 to layer 7 in TCP / IP;
[0018] Calculating the difference between the first and last packets based on the message characteristics;
[0019] Eliminate interference from retransmitted packets based on the calculation result of the difference between the first and last packets to obtain a corrected data packet;
[0020] Calculating the effective delay of the modified data packet by jitter;
[0021] Based on the effective delay result, an exponentially weighted moving average method is used to update the effective delay estimate, which is recorded as application layer delay data.
[0022] Optionally, the operation process of step S3 is specifically as follows:
[0023] Based on the network basic delay data, RTT raw,1 and RTT raw,2 , using the quartile calculation method, calculate the upper delay threshold d th :
[0024] d th =Q3+k Q *IQR
[0025] Among them, Q3 is the upper limit position, k Q is the lower limit, and IQR is the interquartile range.
[0026] Optionally, in step S4, the reward function r(t) is specifically:
[0027]
[0028] Where N represents the N time points to be reviewed with t as the timestamp, C N Indicates the number of channels switched within N time points, where n is the time position from the 0th time point to the Nth time point.
[0029] Optionally, the channel switching model is specifically:
[0030] V π (S)=E π (r t (s t , a t )+δV π (s t+1)| s0=s]
[0031] Q t+1 (s, a) = Q t (s,a)+β[r t (s, a)+δmax a′ Q t (s, a′)-Q t (s, a)]
[0032] Among them, a represents whether to switch the channel action, the subscript t represents the corresponding time, S represents the current state, Q(s,a) represents the value of the Q function when the strategy is π, V π (S) represents the value function with π as the strategy in reinforcement learning, E π Expresses expectation, r t is the reward value at time point t, δ represents the hyperparameter of the next state value function, a' represents the set of possible actions in state t, Q t represents the expected cumulative reward of taking action a at time t, β represents the learning rate of the expected cumulative reward of the Q function, and s0 represents the initial state.
[0033] The present invention also discloses an SD-WAN traffic distribution system based on deep reinforcement learning, the system comprising:
[0034] The data collection module is used to collect network basic delay data and application layer delay data of edge devices in two communication areas of SD-WAN;
[0035] A time period division module, configured to calculate the real delay data based on the network basic delay data and the application layer delay data and then divide the delay data into peak time periods and non-peak time periods;
[0036] a delay threshold calculation module, configured to calculate the delay thresholds of the peak period and the non-peak period using quartiles based on the basic network delay data;
[0037] A model building module, configured to build a reward function using the real delay data and the delay threshold, and to build a channel switching model based on a reinforcement learning architecture, environmental conditions, and the reward function;
[0038] A model training module is used to train and optimize the channel switching model using network basic delay data, application layer delay data, real delay data, peak period and non-peak period to obtain a trained channel switching model;
[0039] The traffic distribution module is used to input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
[0040] Optionally, the two communication area edge devices are CPE and uCPE, and an active monitoring method is used to obtain network basic delay data, and a passive detection method is used to obtain application layer delay data.
[0041] Optionally, the workflow of the delay threshold calculation module is specifically as follows:
[0042] Based on the network basic delay data, RTT raw,1 and RTT raw,2 , using the quartile calculation method, calculate the upper delay threshold d th :
[0043] d th =Q3+k Q *IQR
[0044] Among them, Q3 is the upper limit position, k Q is the lower limit, and IQR is the interquartile range.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The traffic engineering involved in the present invention aims to proactively orchestrate traffic between edge nodes to maximize availability. Actively modify traffic paths between edge nodes, minimize service downtime, predict network problems and adjust traffic routing accordingly. The present invention focuses on optimizing traffic engineering that uses broadband Internet to connect edge devices, and develops a traffic engineering method based on reinforcement learning to improve the performance of SD-WAN-based networks in terms of service availability. In SD-WAN, traffic engineering allows enterprises to arrange traffic based on detection measurements of WAN performance (such as packet delay, packet loss, jitter, and service demand). The present invention introduces active and passive measurement technologies, constructs reinforcement learning states based on the collected delay data, trains reinforcement learning models to implement channel switching, and improves the overall network availability of SD-WAN. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 This is a method step diagram of a SD-WAN traffic distribution method based on deep reinforcement learning according to an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of a hybrid method for measuring inter-channel delay according to an embodiment of the present invention;
[0050] Figure 3 Schematic diagram of a reinforcement learning network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0052] Example 1
[0053] A SD-WAN traffic distribution method based on deep reinforcement learning, such as Figure 1 As shown, the method includes:
[0054] Step S1: Collect network basic delay data and application layer delay data of edge devices in two communication areas of SD-WAN.
[0055] For edge devices in both SD-WAN zones, channel performance data is continuously collected through active probing (such as ICMP delay measurement) and passive monitoring (including but not limited to deep packet inspection and analysis of real-time traffic). This embodiment combines the network base delay provided by ICMP delay measurement with the application layer delay provided by deep packet inspection to record delay indicators and define delay thresholds.
[0056] Passive detection (such as deep packet inspection): Based on L3-L7 layer packet characteristics, identify application types and mark service priorities for use in step 2. Combined with TCP / UDP flow state tracking, record the round-trip time of the first / last packet of the session and eliminate retransmission packet interference. The specific steps are as follows:
[0057] Calculation method of the RTT difference between the first and last packets:
[0058] First packet RTT: RTT 首包 =time 收到首包 -time 发送首包 ;
[0059] Last packet RTT: RTT 末包 =time 收到末包 -time 发送末包 ;
[0060] RTT difference between the first and last packets: RTT 差值 =|RTT 首包 -RTT 末包 |;
[0061] Dynamic threshold setting: RTT threshold =0.2*RTT 首包 ;
[0062] When RTT 插值 >RTT threshold It is determined to be an abnormal retransmission. The real-time audio and video streams are subjected to packet-by-packet jitter measurement, and the effective delay is dynamically calculated through jitter calculation. The specific steps are as follows:
[0063] Calculate the transmission delay difference between adjacent packets i and j:
[0064] ΔD i,j =|D i -D j | (1)
[0065] Use the exponentially weighted moving average (EWMA) to update the effective delay estimate Leff and report it to the server in step 2.
[0066]
[0067] Active detection (including but not limited to ICMP delay measurement): 1. Send a detection packet to the peer gateway every 5 seconds and calculate the delay time as the delay benchmark. 2. The original time series of the acquired active detection data is expressed as:
[0068] RTT raw =[rtt1, rtt2, ..., rtt L ] (3)
[0069] L represents the time period for the delayed system to review. 3. Use the exponentially weighted moving average to obtain the fluctuation value at the current time t:
[0070] E t =γE t-1 +(1-γ)rtt t (4)
[0071] Among them, rtt tis the actual measurement delay at time t. E0 can be taken as the average of the same time period over the past four weeks. γ can be between 0.7 and 0.9. A value of 0.7 indicates sensitivity to rapid changes, while a value of 0.9 indicates smoothness and resistance to short-term fluctuations.
[0072] d(c, t) = E t +k*σ t (5)
[0073] Where d(c, t) is the fluctuation delay value calculated by the cth channel at time t, k is the sensitivity coefficient, which can be 3 when the network environment is stable and 2 when the network environment fluctuates greatly; σ t For History t The standard deviation of . Combining ICMP measurement and deep packet inspection, the delay value measurement solution between channels is as follows Figure 1 shown
[0074] Step S2: Calculate the real delay data based on the network basic delay data and the application layer delay data and then divide the data into peak time periods and non-peak time periods.
[0075] Using the link latency continuously measured by ICMP active probing in step S1 (here, the link latency is the round-trip time of the ICMP packet), combined with an exponentially weighted moving average sliding window algorithm, d(c, t) is obtained. Data is collected over M days, and the average RTT and packet loss rate are calculated. Using a spike exceeding 150% of the average RTT or a packet loss rate exceeding 3% as the peak criteria, periodic high latency periods are identified and marked as the base spike area.
[0076] Time period division logic: (1) If a period has a delay trigger greater than the peak standard for 3 or more times in the past 5 working days, it is marked as a peak period; otherwise, it is a candidate non-peak period. (2) When deep packet inspection detects applications such as video conferencing, γ can be temporarily increased to 0.6 to improve sensitivity and re-identify candidate non-peak periods.
[0077] Step S3: Based on the network basic delay data, calculate the delay thresholds of the peak period and the non-peak period using quartiles. For both peak and non-peak periods, use the historical delay data RTT obtained by active measurement. raw,1 and RTT raw,2 , and use the quartiles to calculate the upper limit threshold of the delay.
[0078] According to the quartile calculation method, the upper limit position Q3=0.75*(nRTT raw +1), lower limit position Q1=0.25*(nRTT raw +1), nRTT rawis the number of measurement delay data in the period, and the interquartile range IQR = Q3-Q1. The upper threshold is d th =Q3+k Q *IQR. k Q The values can be 1.5 and 3. The former indicates a low tolerance to the upper limit, while the latter indicates a high tolerance to the upper limit. For example, if there are two communication areas and there are C channels, then d th There are C*2 (c, t) thresholds, which represent the threshold conditions of the channel during peak and non-peak time periods, and record the maximum delay d max (c,t).
[0079] Step S4: Use the actual delay data and delay threshold to construct a reward function, and build a channel switching model based on the reinforcement learning architecture, environmental conditions, and the reward function. Taking two communication areas and two channels as an example, a dynamic reward function is constructed, and models are built for peak and non-peak time periods. In this embodiment, the reward distance is set to:
[0080]
[0081] Where d(c, t) is the actual current delay in step 2. In this embodiment, if the delay is greater than or equal to 0.05, the reward is set to a constant. Otherwise, it is strongly correlated with m(c, t). The reward function x(c, t) is obtained based on the environmental condition s:
[0082] x(c,t)=10*m(c,t) when m(c,t)<0.05 (7)
[0083] x(c, t)=0.5 when m(c, t)>0.05 (8)
[0084] The environmental condition machine reward weights are shown in Table 1:
[0085] Table 1
[0086] Environmental conditions Reward weight f <![CDATA[d(1,t)≤d th (1,t);d(2,t)≤d th (2,t);C t =1]]> <![CDATA[x(1,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)≤d th (2,t);C t =2]]> <![CDATA[x(2,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)>d th (2,t);C t =1]]> <![CDATA[x(1,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)>d th (2,t);C t =2]]> <![CDATA[-x(2,t) w2 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)≤d th (2,t);C t =1]]> <![CDATA[-x(1,t) w2 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)≤d th (2,t);C t =2]]> <![CDATA[x(2,t) w2 ]]> <![CDATA[d(1,t)>>d th (1,t);d(2,t)>d th (2,t);C t =1]]> <![CDATA[-min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>>d th (1,t);d(2,t)>d th (2,t);C t =2]]> <![CDATA[min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)>>d th (2,t);C t =1]]> <![CDATA[min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)>>d th (2,t);Ct=2]]> <![CDATA[-min[x(1,t),x(2,t)] w3 ]]>
[0087] In the above table, C t Indicates the channel selected at the current time t; w1 = 0.25, w2 = 0.5, w3 = 0.25. In this embodiment, w2 is set to be larger, indicating that the current channel has a higher delay and other channels have lower delays, giving a larger channel switching reward.
[0088] According to the dynamic strategy in Table 1, the final reward function is expressed as:
[0089]
[0090] Where N represents the N time points to be reviewed with t as the timestamp, C NIt represents the number of channels switched within N time points, and tn is the value of x in different time periods.
[0091] According to the state s and reward function r(t) shown in Table 1, the value function used in this embodiment is expressed as follows:
[0092] V π (S)=E π [tr(s t , a t )+δV π (s t+1 )|s0=s] (10)
[0093] Q t+1 (s, a) = Q t (s,a)+β[r t (s, a)+δmax a′ Q t (s, a′)-Q t (s, a)] (11)
[0094] Where a represents whether to switch channels, and the subscript t represents the corresponding time. The algorithm in this embodiment is to find the best strategy to make the state s reach the best Q(s,a) value. V π (S) represents the value function with π as the strategy in reinforcement learning, E π Expresses expectation, r t is the reward value at time point t, δ represents the hyperparameter of the next state value function. a represents the action taken at t, that is, whether to switch links, a' represents the set of possible actions in state t, Q t represents the expected cumulative reward of taking action a at time t, β represents the hyperparameter of the expected cumulative reward of the Q function, s0 represents the initial state, and s represents the current state, which is set according to the link data packet delay time.
[0095] Step S5: Use the network basic delay data, application layer delay data, real delay data, peak period and non-peak period to train and optimize the channel switching model to obtain a trained channel switching model.
[0096] This embodiment builds the model input based on steps 1-2. Taking two channels as an example, the input is [RTT1, RTT2, L eff,1 , L eff,2 , d(1, t), d(2, t), C, P], where ET1 and ET2 are input into formula (5), and Table 1 is the judgment of environmental conditions; C is the currently selected channel, represented by 0 / 1; P indicates whether it is the peak time period, represented by 0 / 1. The middle layer is a hidden layer of 160, and the output layer is action a, such as Figure 3shown.
[0097] Based on the definition in step 4, the training method of this embodiment is as follows:
[0098] (1) Define state s;
[0099] (2) Calculate action a t =argmax α Q * (s t , a t );
[0100] (3) Within time period N, perform action a;
[0101] (4) t , a t , r t , s t+1 ) is stored in cache D;
[0102] (5) Enter the next state s t+1 =s t ;
[0103] (6) Get several (s) recently from cache D t , a t , r t , s t+1 );
[0104] (7) Calculate Q value: y j =r j +γargmax α Q * (s t , a t );
[0105] According to the Q value, optimize the network.
[0106] Step S6: Input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
[0107] For the reinforcement learning training model obtained in step S5, in the application stage, the delay data in the current time period t is collected according to steps 1 and 2 to determine whether it is in the peak time period or the non-peak time period. max and the delay time and environment of different channel detection [RTT1,RTT2,L eff,1 ,L eff,2 ,d(1,t),d(2,t),C,P], and perform switching action selection.
[0108] Example 2
[0109] A SD-WAN traffic distribution system based on deep reinforcement learning, such as Figure 1 As shown, the system includes:
[0110] The data collection module is used to collect network basic delay data and application layer delay data of the edge devices in the two communication areas of SD-WAN.
[0111] For edge devices in both SD-WAN zones, channel performance data is continuously collected through active probing (such as ICMP delay measurement) and passive monitoring (including but not limited to deep packet inspection and analysis of real-time traffic). This embodiment combines the network base delay provided by ICMP delay measurement with the application layer delay provided by deep packet inspection to record delay indicators and define delay thresholds.
[0112] Passive detection (such as deep packet inspection): Based on L3-L7 layer packet characteristics, identify application types and mark service priorities for use in step 2. Combined with TCP / UDP flow state tracking, record the round-trip time of the first / last packet of the session and eliminate retransmission packet interference. The specific steps are as follows:
[0113] Calculation method of the RTT difference between the first and last packets:
[0114] First packet RTT: RTT 首包 =time 收到着包 -time 发送首包 ;
[0115] Last packet RTT: RTT 末包 =time 收到末包 -time 发送末包 ;
[0116] RTT difference between the first and last packets: RTT 差值 =|RTT 首包 -RTT 末包 |;
[0117] Dynamic threshold setting: RTT threshold =0.2*RTT 首包 ;
[0118] When RTT 插值 >RTT threshold It is judged as abnormal retransmission.
[0119] Implement packet-by-packet jitter measurement on real-time audio and video streams and dynamically calculate the effective delay through jitter calculation. The specific steps are as follows:
[0120] Calculate the transmission delay difference between adjacent packets i and j
[0121] ΔD i,j =|D i -Dj | (12)
[0122] Update the effective delay estimate L using the exponentially weighted moving average (EWMA) eff , reported to the server in step 2.
[0123]
[0124] Active detection (including but not limited to ICMP delay measurement):
[0125] 1. Send a probe packet to the peer gateway every 5 seconds and calculate the delay time as the delay benchmark.
[0126] 2. The original time series of the acquired active detection data is expressed as:
[0127] RTT raw =[rtt1, rtt2, ..., rtt L ] (14)
[0128] L represents the time period for delayed system review.
[0129] 3. Use the exponentially weighted moving average to obtain the fluctuation value at the current time t:
[0130] E t =γE t-1 +(1-γ)rtt t (15)
[0131] Among them, rtt t is the actual measurement delay at time t. E0 can be taken as the average of the same time period over the past four weeks. γ can be between 0.7 and 0.9. A value of 0.7 indicates sensitivity to rapid changes, while a value of 0.9 indicates smoothness and resistance to short-term fluctuations.
[0132] d(c, t) = E t +k*σ t (16)
[0133] Where d(c, t) is the fluctuation delay value calculated by the cth channel at time t, k is the sensitivity coefficient, which can be 3 when the network environment is stable and 2 when the network environment fluctuates greatly; σ t For History t The standard deviation of .
[0134] Integrated ICMP measurement and deep packet inspection, the delay value measurement solution between channels is as follows Figure 1 shown
[0135] The time period division module is used to calculate the real delay data based on the network basic delay data and the application layer delay data and then divide the time period into peak time period and non-peak time period.
[0136] The link delay continuously measured by ICMP active detection, i.e., the round-trip time of ICMP packets, is obtained by combining a sliding window algorithm with an exponentially weighted moving average. Data is collected for M days, and the average RTT and packet loss rate are calculated. Using a peak criterion of more than 150% of the average RTT or a packet loss rate exceeding 3%, periodic high-delay periods are identified and marked as basic peak areas. The schematic diagram of the hybrid method used by the present invention to measure inter-channel delay is shown below. Figure 2 shown.
[0137] Time period division logic:
[0138] (1) If a period of time has an active measurement delay trigger greater than the peak standard for 3 or more times in the past 5 working days, it is marked as a peak period; otherwise, it is a candidate non-peak period.
[0139] (2) When deep packet inspection detects applications such as video conferencing, γ can be temporarily increased to 0.6 to improve sensitivity and re-identify candidate non-peak periods.
[0140] The delay threshold calculation module is used to calculate the delay thresholds of the peak period and the non-peak period based on the network basic delay data using quartiles.
[0141] For both peak and non-peak periods, historical delay data RTT is obtained by active measurement. raw,1 and RTT raw ,2, respectively use the quartiles to calculate the upper limit threshold of delay.
[0142] According to the quartile calculation method, the upper limit position Q3=0.75*(nRTT raw +1), lower limit position Q1=0.25*(nRTT raw +1), nRTT raw is the number of measurement delay data in the period, and the interquartile range IQR = Q3-Q1. The upper threshold is d th =Q3+k Q *IQR. k Q The values can be 1.5 and 3, the former indicating low tolerance for the upper limit and the latter indicating high tolerance for the upper limit.
[0143] Take two regions communicating as an example, there are C channels, then d th There are C*2 (c, t) thresholds, which represent the threshold conditions of the channel during peak and non-peak time periods, and record the maximum delay d max (c,t).
[0144] A model building module is used to build a reward function using the real delay data and the delay threshold, and to build a channel switching model based on the reinforcement learning architecture, environmental conditions and the reward function.
[0145] Taking two communication areas and two channels as an example, a dynamic reward function is constructed, and models are built for peak and non-peak time periods respectively. In this embodiment, the reward distance is set as:
[0146]
[0147] Where d(c,t) is the current real delay in step 2.
[0148] In this example, if the delay is greater than or equal to 0.05, the reward is set to a constant. Otherwise, it is strongly correlated with m(c, t). The reward function x(c, t) is defined as:
[0149] x(c,t)=10*m(c,t) when m(c,t)<0.05 (18)
[0150] x(c,t)=0.5 when m(c,t)>0.05 (19)
[0151] The environmental condition machine reward weights are shown in Table 2:
[0152] Table 2
[0153] Environmental conditions Reward weight f <![CDATA[d(1,t)≤d th (1,t);d(2,t)≤d th (2,t);C t =1]]> <![CDATA[x(1,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)≤d th (2,t);C t =2]]> <![CDATA[x(2,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)>d th (2,t);C t =1]]> <![CDATA[x(1,t) w1 ]]> <![CDATA[d(1,t)≤d th (1,t);d(2,t)>d th (2,t);C t =2]]> <![CDATA[-x(2,t) w2 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)≤d th (2,t);C t =1]]> <![CDATA[-x(1,t) w2 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)≤d th (2,t);C t =2]]> <![CDATA[x(2,t) w2 ]]> <![CDATA[d(1,t)>>d th (1,t);d(2,t)>d th (2,t);C t =1]]> <![CDATA[-min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>>d th (1,t);d(2,t)>d th (2,t);C t =2]]> <![CDATA[min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)>>d th (2,t);C t =1]]> <![CDATA[min[x(1,t),x(2,t)] w3 ]]> <![CDATA[d(1,t)>d th (1,t);d(2,t)>>d th (2,t);C t =2]]> <![CDATA[-min[x(1,t),x(2,t)] w3 ]]>
[0154] In the above table, C t Indicates the channel selected at the current time t; w1 = 0.25, w2 = 0.5, w3 = 0.25. In this embodiment, w2 is set to be larger, indicating that the current channel has a higher delay and other channels have lower delays, giving a larger channel switching reward.
[0155] According to the dynamic strategy in Table 2, the reward function is expressed as:
[0156]
[0157] Where N represents the N time points to be reviewed with t as the timestamp, C N Indicates the number of channels switched within N time points.
[0158] According to the state s and reward function r(t) shown in Table 2, the value function used in this embodiment is expressed as follows:
[0159] The channel switching model is specifically:
[0160] V π (S)=E π [r t(s t , a t )+δV π (s t+1 )|s0 (21)
[0161] Q t+1 (s, a) = Q t (s,a)+β[r t (s, a)+δmax a′ Q t (s, a′)-Q t (s, a)] (22)
[0162] Among them, a represents whether to switch the channel action, the subscript t represents the corresponding time, s is the current state, Q(s,a) represents the value of the Q function when the strategy is π, V π (S) represents the value function with π as the strategy in reinforcement learning, E π Expresses expectation, r t is the reward value at time point t, δ represents the hyperparameter of the next state value function, a' represents the set of possible actions in state t, Q t represents the expected cumulative reward of taking action a at time t, β represents the learning rate of the expected cumulative reward of the Q function, and s0 represents the initial state.
[0163] The algorithm of this embodiment is to find the best strategy to make the state s reach the best Q(s,a) value.
[0164] A model training module is used to train and optimize the channel switching model using network basic delay data, application layer delay data, real delay data, peak period and non-peak period to obtain a trained channel switching model;
[0165] Model construction: Taking two channels as an example, the input is:
[0166] [RTT1, RTT2, L eff,1 , L eff,2 , d(1, t), d(2, t), C, P], where ET1 and ET2 are input into Equation (5), and Table 1 is the judgment of environmental conditions; C is the currently selected channel, represented by 0 / 1; P indicates whether it is the peak time period, represented by 0 / 1. The middle layer is a hidden layer with 160 pixels, and the output layer is action a.
[0167] The training method of this embodiment is as follows:
[0168] (8) Define state s;
[0169] (9) Calculate action a t =argmax α Q *(s t , a t );
[0170] (10) During time period N, perform action a;
[0171] (11) t , a t , r t , s t+1 ) is stored in cache D;
[0172] (12) Enter the next state s t+1 =s t ;
[0173] (13) Recent acquisition of several (s) from cache D t , a t , r t , s t+1 );
[0174] (14) Calculate Q value: y j =r j +γargmax α Q * (s t , a t );
[0175] According to the Q value, optimize the network.
[0176] The traffic distribution module is used to input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
[0177] For the reinforcement learning training model obtained in step S5, in the application stage, the delay data in the current time period t is collected according to steps 1 and 2 to determine whether it is in the peak time period or the non-peak time period. max and the delay time and environment of different channel detection [RTT1,RTT2,L eff,1 ,L eff,2 ,d(1,t),d(2,t),C,P], and perform switching action selection.
[0178] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A SD-WAN traffic distribution method based on deep reinforcement learning, characterized in that: The method comprises: Step S1: Collect network basic delay data and application layer delay data of two communication area edge devices of SD-WAN, where the two communication area edge devices are CPE and uCPE; Step S2: calculating the actual delay data based on the network basic delay data and the application layer delay data and dividing the delay data into peak time periods and non-peak time periods; Step S3: Calculate the delay thresholds of the peak period and the non-peak period using quartiles based on the basic network delay data; Step S4: constructing a reward function using the real delay data and the delay threshold, and constructing a channel switching model based on the reinforcement learning architecture, environmental conditions, and the reward function; Step S5: training and optimizing the channel switching model using network basic delay data, application layer delay data, real delay data, peak time periods and non-peak time periods to obtain a trained channel switching model; Step S6: Input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
2. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 1 is characterized in that: In step S1, the two communication area edge devices are CPE and uCPE, and an active monitoring method is used to obtain network basic delay data, and a passive detection method is used to obtain application layer delay data.
3. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 2 is characterized in that: The active detection method is specifically as follows: Send a detection packet to the peer gateway at a preset sending period, calculate the delay time, and obtain the initial delay time data; obtaining an original time series based on the initial delay time data; Based on the sequence value of the original time series, the fluctuation value at time t is obtained using the exponentially weighted moving average method; A fluctuation delay value is calculated based on the fluctuation value and recorded as network basic delay data.
4. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 3 is characterized in that: The passive detection method includes: Extract the message features from layer 3 to layer 7 in TCP / IP; Calculating the difference between the first and last packets based on the message characteristics; Eliminate interference from retransmitted packets based on the calculation result of the difference between the first and last packets to obtain a corrected data packet; Calculating the effective delay of the modified data packet by jitter; Based on the effective delay result, an exponentially weighted moving average method is used to update the effective delay estimate, which is recorded as application layer delay data.
5. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 1 is characterized in that: The specific operation process of step S3 is: Based on the network basic delay data, RTT raw,1 and RTT raw,2 , using the quartile calculation method, calculate the upper delay threshold d th : is th =Q3+k Q *IQR Among them, Q3 is the upper limit position, k Q is the lower limit, and IQR is the interquartile range.
6. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 1, characterized in that: In step S4, the reward function r(t) is specifically: Where N represents the N time points to be reviewed with t as the timestamp, C N Indicates the number of channels switched within N time points, where n is the time position from the 0th time point to the Nth time point.
7. The SD-WAN traffic distribution method based on deep reinforcement learning according to claim 6 is characterized in that: The channel switching model is specifically: V π (S)=E π [r t (s t ,a t )+δV π (s t+1 )|s0=s] Q t+1 (s,a)=Q t (s,a)+β[r t (s,a)+δmax a′ Q t (s,a′)-Q t (s,a)] Among them, a represents whether to switch the channel action, the subscript t represents the corresponding time, S represents the current state, Q(s,a) represents the value of the Q function when the strategy is π, V π (S) represents the value function with π as the strategy in reinforcement learning, E π Expresses expectation, r t is the reward value at time point t, δ represents the hyperparameter of the next state value function, a' represents the set of possible actions in state t, Q t represents the expected cumulative reward of taking action a at time t, β represents the learning rate of the expected cumulative reward of the Q function, and s0 represents the initial state.
8. A SD-WAN traffic distribution system based on deep reinforcement learning, the system is used to implement the method according to any one of claims 1 to 7, characterized in that the system include: The data collection module is used to collect network basic delay data and application layer delay data of edge devices in two communication areas of SD-WAN; A time period division module, configured to calculate the real delay data based on the network basic delay data and the application layer delay data and then divide the delay data into peak time periods and non-peak time periods; a delay threshold calculation module, configured to calculate the delay thresholds of the peak period and the non-peak period using quartiles based on the basic network delay data; A model building module, configured to build a reward function using the real delay data and the delay threshold, and to build a channel switching model based on a reinforcement learning architecture, environmental conditions, and the reward function; A model training module is used to train and optimize the channel switching model using network basic delay data, application layer delay data, real delay data, peak time periods and non-peak time periods to obtain a trained channel switching model; The traffic distribution module is used to input the real-time collected delay data into the trained channel switching model to achieve SD-WAN traffic distribution.
9. The SD-WAN traffic distribution system based on deep reinforcement learning according to claim 8, characterized in that The two communication area edge devices are CPE and uCPE, and an active monitoring method is used to obtain network basic delay data, and a passive detection method is used to obtain application layer delay data.
10. The SD-WAN traffic distribution system based on deep reinforcement learning according to claim 8, characterized in that: The specific working process of the delay threshold calculation module is as follows: Based on the network basic delay data, RTT raw,1 and RTT raw,2 , using the quartile calculation method, calculate the upper delay threshold d th : is th =Q3+k Q *IQR Among them, Q3 is the upper limit position, k Q is the lower limit, and IQR is the interquartile range.