Space-air-ground network congestion control method based on deep reinforcement learning
By dynamically adjusting the sending rate of the air-space-ground network through a method based on deep reinforcement learning, the congestion control deficiencies of traditional methods in complex dynamic environments are solved, achieving higher reliability and accuracy.
Patent Information
- Application Number
- CN202411609681.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Traditional air-ground-space network congestion control methods are difficult to achieve efficient resource utilization and avoid congestion when faced with complex and highly dynamic network environments. In addition, existing methods have shortcomings in considering real-time status and optimal solutions.
A method based on deep reinforcement learning is adopted to obtain data information of the air-space-ground network, set optimization goals and conduct modeling, build a state prediction module, a reward redistribution module and an optimization strategy module, and dynamically adjust the sending rate to achieve congestion control.
It achieves higher reliability and accuracy in the air-space-ground network, can effectively handle congestion problems in complex dynamic environments, and meet the diverse needs of users.
Smart Images

Figure CN119364423B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of communication, and particularly relates to a space-air-ground network congestion control method based on deep reinforcement learning. BACKGROUND
[0002] In order to realize global seamless coverage and meet the personalized needs of users for network services, a space-air-ground integrated network (GASN) by combining a low-orbit satellite system, an aerial network and a ground communication has become a new network architecture and attracted wide attention. A network transmission layer adaptive congestion control mechanism becomes a key technology for users to guarantee transmission quality in different network scenarios by trying to seek a balance between improving network resource utilization and avoiding network congestion.
[0003] Compared with a traditional ground network link environment, the mobility of a low-orbit satellite, the unpredictability of weather changes and the diversity of user needs in a space-air-ground network system result in the characteristics of long data round-trip delay, high bit error rate and high heterogeneity of uplink and downlink; therefore, a traditional congestion control protocol algorithm applied to a ground network is difficult to achieve good performance in this scenario.
[0004] At present, traditional and commonly used congestion control methods mainly include state-independent schemes and optimization-based schemes. The state-independent method is usually simple in algorithm or protocol design, but this kind of scheme lacks careful consideration of real-time state and optimal solution, so its performance in a high time-varying network can be very poor; the optimization-based scheme must accurately predict the future values of some key parameters (such as link utilization, user usage) as input, and needs an accurate mathematical model to estimate or represent network behavior, so the comprehensive performance of this kind of scheme is relatively poor when facing a complex and highly dynamic network, and the implementation is difficult. SUMMARY
[0005] The application aims to provide a space-air-ground network congestion control method based on deep reinforcement learning with high reliability, good accuracy and good effect.
[0006] The space-air-ground network congestion control method based on deep reinforcement learning provided by the application comprises the following steps:
[0007] S1. Obtain data information of a target space-air-ground network;
[0008] S2. Set an optimization target and perform modeling according to the data information obtained in step S1;
[0009] S3. Set a state space, an action space and a reward function of a decision-making process;
[0010] S4. Constructing a state prediction module, a reward redistribution module and an optimization strategy module of the decision-making process;
[0011] S5. According to the constructed modules, carrying out congestion control of the space-air-ground network based on deep reinforcement learning.
[0012] According to the data information obtained in step S1, the optimization target is set and modeling is carried out, which specifically includes the following steps:
[0013] The set optimization target includes throughput, delay, jitter, packet loss rate and reliability of user application;
[0014] The throughput represents the amount of data successfully transmitted per unit of time, which is used to reflect the transmission capacity of the network and the data transmission efficiency of the application;
[0015] The delay represents the total time from sending to receiving data, which is used to reflect the response speed and real-time performance of the user application;
[0016] The jitter represents the variation of the time interval between consecutive data packets, which is used to reflect the playback quality of the streaming media application;
[0017] The packet loss rate represents the proportion of lost data packets in the total sent data packets during the data transmission process, which is used to reflect the reliability and integrity of the user application;
[0018] The reliability represents the uninterrupted service capability of the network system within a given time, which is used to reflect the continuity and availability of the user application;
[0019] Let F be a flow set with at least one and at most two of the optimization targets, T f (t) is the long-term throughput of the sub-flow f initiated by a single ground user measured at time t, L f (t) is the long-term packet loss rate of the sub-flow f initiated by a single ground user measured at time t, and is defined as
[0020]
[0021] where L ep is the number of periods considered; is the throughput measured in the kth period; η is the set decay coefficient; is the packet loss rate measured in the kth period;
[0022] D f (t) represents the delay of the sub-flow f measured at time t, J f (t) represents the jitter of the sub-flow f measured at time t, R f(t) represents the reliability of the sub-flow f measured at time t; the single-index utility function of the sub-flow f at time t is denoted as is set as
[0023]
[0024] wherein SI represents a single-index factor considered; T represents a throughput index; L represents a delay index; D represents a jitter index; J represents a packet loss rate index; and R represents a reliability index;
[0025] F represents the sum of the single-index utilities of all flows except the sub-flow f at time t is denoted as
[0026]
[0027] wherein f' represents any flow in F except the sub-flow f; is a single-index utility function of f' at time t;
[0028] For simplicity and without loss of generality, two double-index utility functions are adopted, denoted as
[0029]
[0030] wherein is a utility function considering the throughput and delay double index; is a throughput weight; is a delay weight; is a utility function considering the reliability and packet loss rate double index; is a reliability weight; is a packet loss rate weight; U α () is a fairness coefficient function, and is a fairness parameter;
[0031] F represents the sum of the double-index utilities of all flows except the sub-flow f at time t is denoted as
[0032]
[0033] The state space of the decision process in step S3 is set, specifically including the following steps:
[0034] The space-time-aerospace network environment information and network layer information are included in the state space, and the potential bottleneck link is dynamically learned;
[0035] The state space at time t includes the space-time-aerospace integrated network state, the network layer state and the transmission layer state, denoted as wherein, is the state of the sub-flow f at time t, and is denoted as is the space-air-ground integrated network state of subflow f at time t, is the network layer state of subflow f at time t, is the transport layer state of subflow f at time t;
[0036] and is specifically expressed as
[0037]
[0038] in the formula is the space-air link channel gain at time t; is the space-satellite link channel gain at time t; is the satellite-air link channel gain at time t; is the ratio of the used power to the maximum power of the space-air link at time t; is the ratio of the used power to the maximum power of the space-satellite link at time t; is the ratio of the used power to the maximum power of the satellite-air link at time t; is the ratio of the used frequency band to the total frequency band resource of the space-air link at time t; is the ratio of the used frequency band to the total frequency band resource of the space-satellite link at time t; is the ratio of the used frequency band to the total frequency band resource of the satellite-air link at time t; is the number of data packets accumulated in the LEO satellite; is the number of data packets accumulated in the UAV; is the number of data packets accumulated in the satellite earth station; is the export link channel resource utilization rate of the LEO satellite; is the export link channel resource utilization rate of the UAV; is the export link channel resource utilization rate of the satellite earth station; sr t f is the sending rate of subflow f at time t; is the throughput of subflow f at time t; is the average RTT of subflow f at time t; is the RTT average deviation of subflow f at time t; is the congestion window size of subflow f at time t; is the congestion window change value of subflow f between time t-1 and time t; is the throughput change value of subflow f between time t-1 and time t; is the ratio of the minimum RTT of subflow f to the RTT at time t; lr t f is the loss rate of subflow f at time t; is the RTT of sub-flow f at time t; is the interval value of the arrival time of ACK returned between time t-1 and t of sub-flow f.
[0039] The action space of the decision process in step S3 is set, specifically including the following steps:
[0040] The action expression based on the sending rate control is set, denoted as
[0041]
[0042] wherein is the action space expression of sub-flow f at time t; Δsr t f is the sending rate conversion value, and Δsr t f ≥ 0 corresponds to the sending rate increment value, Δsr t f < 0 corresponds to the sending rate decrement value; is the repetition factor of the data packet launched on sub-flow f at time t;
[0043] According to the action expression and the current state, the sending rate update action is converted into the congestion window update action, denoted as
[0044]
[0045] wherein L MTU is the length of the maximum transmission unit of the transport layer;
[0046] Therefore, the global action space a t is denoted as
[0047] The reward function of the decision process in step S3 is set, specifically including the following steps:
[0048] The reward function of sub-flow f is set as
[0049]
[0050] wherein r t f is the cumulative reward value; is the trade-off coefficient used for balancing the benefits between sub-flow f and other flows; denotes the utility function of sub-flow f considering single-index factors or multi-index factors;
[0051] The whole state-action sequence of one time is divided into several same and length L SASshort sequences, and estimate the cumulative reward of each short sequence; for the sequence from t time to t+L SAS -1 time end, the corresponding cumulative reward is expressed as The cumulative reward is clipped to ensure the value range is [-1, 1].
[0052] The state prediction module of the decision process in step S4 specifically includes the following steps:
[0053] Obtain the space-air-ground integrated network state, network layer state and transmission layer state provided by the current environment, and predict the latest space-air-ground integrated network state and network layer state through the constructed state prediction module;
[0054] The ground user can implement and obtain the space-air-ground integrated system network state and network layer state, and transmit the obtained space-air-ground integrated network system state and network layer state together with the free field domain in the ACK data packet protocol header to the sender of the response data packet;
[0055] In each cycle, each sender obtains a number of space-air-ground integrated network state and network layer state sequences with timestamps for predicting the latest space-air-ground integrated network state and network layer state;
[0056] The indicators for predicting the space-air-ground integrated network state and network layer state include channel gain, ratio of used power to maximum power, ratio of used frequency band to total frequency band resource, data packet backlog and bottleneck node's egress link utilization;
[0057] Five GRU-based neural network models with the same structure are used to predict the channel gain, ratio of used power to maximum power, ratio of used frequency band to total frequency band resource, data packet backlog and bottleneck node's egress link utilization respectively and individually; meanwhile, the time interval of the state is taken as a feature for prediction; the true value of the space-air-ground integrated network state and network layer state is obtained by the LEO satellite, unmanned aerial vehicle, satellite ground station base station or ground user at each time slot.
[0058] The reward redistribution module of the decision process in step S4 specifically includes the following steps:
[0059] The congestion control problem with dynamic delay, balance exploration and fairness guarantee is regarded as a sequence Markov decision process SDP with delay reward, which is represented by a five-tuple (S, A, R, P, γ), wherein S is the global state space, A is the joint action space of all individuals, R is the reward function set of all individuals, P is the state transition probability matrix, and γ is the discount factor;
[0060] Based on the second-order Markov reward redistribution framework, a sequence Markov decision process M without cumulative rewards is obtained from a sequence Markov decision process M with cumulative rewards and M and have the same state space, action space, state transition probability and optimal policy;
[0061] Set as the Q-value function of M with delayed rewards, the redistribution reward of M with delayed rewards is denoted as , the second-order Markov reward redistribution is denoted as
[0062]
[0063] Set the reward-prediction function g f , which is used to predict the expected cumulative reward of M at the end of a given state-action sequence at time t; the output of the reward-prediction function g f is Since contains the information of , set the difference function Δ f and use it to calculate the information carried in ; Δ f is defined as the numerical difference between and , and the state-action sequence processed based on Δ f is denoted as
[0064]
[0065] At the same time, g f must ensure
[0066] Thus, we have: (1) is calculated by and is used as an estimate of ; (2)
[0067]
[0068] Convert g f (Δ 0:T ) and g f (Δ 0:t ) into where is the contribution of the state-action sequence to the predicted expected reward, and is calculated by ;
[0069] Set Then And There is an error between
[0070] Because And There is an error between And
[0071]
[0072] Finally get
[0073] The calculation process of the reward redistribution module specifically includes the following steps:
[0074] For a state-action sequence Indicates the expected feedback before execution, Indicates the expected feedback after execution, And The numerical difference between Contribution of As input, As a label, and Redistribute to each period, eventually forming a sequence Markov decision process with delayed rewards that can be solved using DRL methods
[0075]
[0076] Specifically includes the following steps:
[0077] Input: state-action sequence pair Feedback prediction function gf, difference function Δf, 0≤t≤T, f∈F;
[0078] Output: state-action sequence Contribution to the predicted expected reward
[0079] A. Calculate
[0080] B. Calculate
[0081] C. Calculate Get
[0082] D. Calculate Get
[0083] E. Return the result of the calculation and
[0084] The calculation process of the reward redistribution module is corrected, specifically including the following steps:
[0085] The expected value of is expressed by the real value, obtaining Each round is divided into several short sequences with length L SAS The cumulative reward of each short sequence is returned to each action; the redistribution correction method is used to ensure that holds;
[0086] Specifically including the following steps:
[0087] Input: feedback prediction function Cumulative reward
[0088] Output: redistributed reward
[0089] a. Calculate
[0090] b. Calculate the uncorrected redistribution reward is where t≠0, and
[0091]
[0092] c. Calculate the average error
[0093] d. Calculate the corrected redistribution reward is
[0094] e. Return
[0095] The optimization strategy module for constructing the decision-making process in step S4, specifically including the following steps:
[0096] The feature extraction network first outputs a feature vector Since at most L rp rounds are involved, at most L rp feature vectors are generated; the global state sequence is represented as Then, is input into another representation network to generate a global feature vector containing global information; finally, each agent generates an action using the global feature vector; a GRU-based neural network is used as the representation network;
[0097] Each agent shares a global feature vector; the actor network parameters The action probability of the output sub-flow f based on the current state, the critic network parameters Used to update and feed back to the actor network by predicting future rewards; based on the feedback behavior, the actor network can continuously adjust the probability of future action selection, and repeat the feedback process in each learning iteration to improve the strategy;
[0098] In order to prevent fluctuations in the training process from exceeding the set range and to ensure the stability of the learning process, the following formula is used to update the neural network parameters:
[0099]
[0100] In the formula is the clipping loss function, which is used to limit the amplitude of the strategy update; represents the expected value of all rounds; is the ratio between the old and new strategies, and For the new strategy in state Take action The probability of For the old policy in state Take action probability; is the advantage function of the action at time t, which is used to measure the advantage of the action at time t relative to the average action; is an intermediate function, and ε is a set positive number, which is used to clip the ratio in the objective function To avoid policy updates being too large or too small;
[0101] Calculated using the generalized advantage estimator Expressed as
[0102]
[0103] Where γ is the discount factor used to reduce the weight of future returns, λ is the parameter that controls the trade-off between bias and variance, T is the total number of periods, is the Bellman residual term and V() is an approximate value function, and V() is usually estimated by a neural network;
[0104] The total loss is calculated as
[0105]
[0106] In the formula For the value function loss, which is used to measure the approximation of the value function and the expected return at time t; H is the entropy function that the policy promotes exploration by hindering the deterministic policy at each time t; c1 is the first weight coefficient; c2 is the second weight coefficient; R represents the GRU-based neural network;
[0107] During the optimization process, all the sequences of log probabilities are used to output the action sequence to ensure fairness; the single log associated with each action selected before is modified to a set value, which is used to prevent the softmax function from generating the selected operation for the log; the gradient of the log of the invalid operation is set to zero; the action is output according to the modified log using the softmax function, denoted as
[0108] a f,j =softmax(L f )
[0109]
[0110] In the formula, a f,j is the jth action of the subflow f; L f is the modified sequence of log probabilities of the subflow f, and L f ={l f,1 ,...,l f,j ,...,l f,J}, J is the number of actions that can be selected by the subflow f; σ is a set negative reward;
[0111] The subflow f must select one action from each of the following two types of action spaces respectively:
[0112]
[0113] In the formula, Δsr is the transmission rate increment of the subflow f at time t; is a negative value indicating a transmission rate reduction of the subflow f, is a positive value indicating a transmission rate increase of the subflow f, is 0 indicating the same transmission rate of the subflow f; N is the maximum number of repetitions in a data packet; Δsr is the change degree of the transmission rate; is the data packet repetition factor of the subflow f at time t;
[0114] Therefore, the action a is converted to a The action space expression for controlling the transmission rate is as follows:
[0115]
[0116] By the reward redistribution mechanism, it can be assumed that the action feedback in each sequence will be returned at the moment L SAS .
[0117] The space-ground-ground network congestion control method based on deep reinforcement learning, wherein the training process comprises the following steps:
[0118] Input: maximum training round E, training period length T, algorithm weight λ, discount factor γ, batch size N batch ;
[0119] Output: actor network parameters critic network parameters representation network parameters
[0120] (1) Random parameter initialization of actor network parameters critic network parameters and representation network parameters
[0121] (2) Initialize the replay buffer B f ;
[0122] (3) Initialize an Ornstein-Uhlenbeck random process for exploratory action;
[0123] (4) At each time t, collect the space-ground-ground integrated network state, network layer state and transport layer state of all flows;
[0124] (5) Use the state prediction module to predict the latest chapter, get the global state s t ;
[0125] (6) Get the feature vector h rp from the latest L t segment global state by GRU neural network R
[0126] (7) Get an action a according to the policy function π()
[0127] (8) Based on the action a , generate an action a using a random process
[0128] (9) Execute the action a and observe the current utility function;
[0129] (10) Store the global state s t , the feature vector h t , the action a and the corresponding utility function value;
[0130] (11) Calculate
[0131] (12) Calculate the cumulative reward of each sequence according to the utility function value;
[0132] (13) Redistribute the cumulative reward through the calculation process of the reward redistribution module to obtain
[0133] (14) Store the transition process to the buffer B f ;
[0134] (15) Extract N batch samples from the replay buffer and divide them into K batches;
[0135] (16) In each mini-batch, sequentially execute the following steps (17) to (20);
[0136] (17) Obtain the feature vector h k for
[0137] (18) Determine the action according to the policy function π()
[0138] (19) Calculate and
[0139]
[0140] (20) Update the actor network parameters , critic network parameters and representation network parameters according to the total loss
[0141]
[0142] (21) Return the finally obtained actor network parameters , critic network parameters and representation network parameters
[0143] The space-time network congestion control method based on deep reinforcement learning provided by the application, by obtaining data information of multiple target space-time network data information, and based on the setting of state space, action space and reward function, as well as the realization of state prediction, reward redistribution and optimization strategy, not only realizes the congestion control of space-time network based on deep reinforcement learning, but also has higher reliability, better accuracy and better effect. BRIEF DESCRIPTION OF DRAWINGS
[0144] Figure 1 A method flowchart of the method of the application.
[0145] Figure 2 A typical end-to-end data transmission scenario diagram of the space-air-ground integrated network of the method of the application.
[0146] Figure 3 A whole structure design diagram of the scheme based on the heterogeneous multi-agent framework of the method of the application.
[0147] Figure 4 A feature extraction network architecture diagram of the method of the application.
[0148] Figure 5 A model training convergence process diagram of the embodiment of the method of the application.
[0149] Figure 6 A satisfaction situation diagram of the quality of service requirement in the model training process of the embodiment of the method of the application. DETAILED DESCRIPTION
[0150] As Figure 1 shown is a method flowchart of the method of the application: the space-air-ground network congestion control method based on deep reinforcement learning disclosed in the application comprises the following steps:
[0151] S1. acquiring data information of a target space-air-ground network;
[0152] As Figure 2A typical end-to-end data transmission scenario in the integrated space-air-ground network is shown. Among them, a group of ground users can access remote application servers through an integrated space-air-ground network composed of wireless access points, satellite earth stations, unmanned aerial vehicles as mobile base stations and low-orbit satellites. In addition, the network architecture can also transmit the data collected by the sensing nodes to the remote data center, wherein the aggregation node is responsible for establishing a connection with the remote storage device to ensure the quality of data transmission. Of course, the isolated sensing node can also be individually responsible for the remote end-to-end transmission of its own sensing data. Due to the burstiness of the sensing data upload task, the aggregation node needs to deploy a congestion control scheme to avoid congestion of the data upload path. On the contrary, for the ground users, the data size of the request packet is usually small, and the data size of the response packet is usually large. Therefore, the application server also needs to deploy a congestion control scheme to avoid congestion on the data download path. It is known that in the integrated space-air-ground network, the channel quality variation of the space-air link and the space-ground link has high dynamics, which can be considered as a potential bottleneck of the overall network performance. According to the general principle of congestion control, when congestion occurs in the network environment due to various complex factors (such as a large number of unacknowledged data packets, poor link quality, etc.), the transmission layer sending rate should be reduced. The purpose of the present application is to propose a dynamic adaptive congestion control method for the above-mentioned potential bottleneck link of the network, to solve the problem that the existing congestion control method cannot be well applied to the highly dynamic and heterogeneous space-air-ground network.
[0153] S2. Set the optimization target according to the data information obtained in step S1, and model; Specifically, it includes the following steps:
[0154] The present application defines the congestion control as a decision problem considering dynamic delay, seeking balance and ensuring fairness. The dynamic delay characteristic of the problem means that the feedback reward of each action will not return at a fixed time, and the feedback reward may be generated by several consecutive actions. Therefore, the concept of reward redistribution should be used to deal with this problem. The balance seeking feature of the problem means that a balance needs to be found between several mutually exclusive QoS performance targets, such as achieving good throughput under the RTT acceptable to users. The fairness characteristic of the problem requires that when a group of end-to-end flows share limited bandwidth, the fairness of all sub-flows needs to be maintained;
[0155] The set optimization target includes the throughput, delay, jitter, packet loss rate and reliability of the user application;
[0156] Throughput represents the amount of data successfully transmitted per unit time, which is used to reflect the transmission capacity of the network and the data transmission efficiency of the application; For large data applications such as video and file transmission, higher throughput is required;
[0157] The total time of delay represents the data from sending to receiving, which is used to reflect the response speed and real-time of user application; for example, the voice and video call application requires low delay for high real-time requirement;
[0158] The jitter represents the variation of the time interval of continuous data packet, which is used to reflect the play quality of streaming media application and affect the play quality of streaming media application; for example, the streaming media application requires as small jitter as possible;
[0159] The packet loss rate represents the proportion of lost data packet in the total sent data packet in the data transmission process, which is used to reflect the reliability and integrity of user application; it directly affects the reliability and integrity of application; for example, the file transmission application requires as low packet loss rate as possible for high reliability requirement;
[0160] The reliability represents the uninterrupted service ability of network system within a given time, which is used to reflect the continuity and availability of user application; for example, the key business application requires high reliability guarantee;
[0161] In fact, most of the current application programs may focus on one or two of the above five QoS indicators; while the application programs with high requirements for three or more indicators are rare; on this basis, the above congestion control problem should be applied to all combinations of two indicators with practical significance, and the lower bound of other QoS indicators is taken as a constraint condition, instead of only limiting to throughput and delay, so as to more flexibly adapt to the QoS requirements of most application programs;
[0162] Let F be a flow set with at least one and at most two optimization objectives, T f (t) is the long-term throughput of the sub-flow f initiated by a single ground user measured at t, L f (t) is the long-term packet loss rate of the sub-flow f initiated by a single ground user measured at t, and is defined as
[0163]
[0164] In the formula, L ep is the number of periods considered; is the throughput measured in the kth period; η is the decay coefficient set, and the value range is (0, 1]; is the packet loss rate measured in the kth period;
[0165] D f (t) represents the delay of the sub-flow f measured at t, J f (t) represents the jitter of the sub-flow f measured at t, R f (t) represents the reliability of the sub-flow f measured at t; the single-index utility function at t is is set as
[0166]
[0167] wherein SI denotes a single-index factor under consideration; T denotes a throughput index; L denotes a latency index; D denotes a jitter index; J denotes a packet loss rate index; R denotes a reliability index;
[0168] F is a single-index utility of all flows except the sub-flow f at time t and is expressed as
[0169]
[0170] wherein f' denotes any flow in F except the sub-flow f; is a single-index utility function of f' at time t;
[0171] For simplicity without loss of generality, two double-index utility functions are adopted, expressed as
[0172]
[0173] wherein is a utility function considering throughput and latency double indexes; is a throughput weight; is a latency weight; is a utility function considering reliability and packet loss rate double indexes; is a reliability weight; is a packet loss rate weight; U α is a fairness coefficient function, and is a fairness parameter;
[0174] F is a double-index utility of all flows except the sub-flow f at time t and is expressed as
[0175]
[0176] The present application proposes an improved semi-independent near-end policy optimization (ISIPPO) model method for solving the problem of optimizing the above utility function;
[0177] S3. Setting the state space, action space and reward function of the decision-making process; specifically comprising the following steps:
[0178] Multiple response data streams for the same ground user can come from different application servers, so for the congestion control scheme deployed at the sender, the multi-agent model is more adaptive and scalable. In addition, the ground user can start a request with different QoS requirements at the same time, resulting in response streams with different QoS requirements. Therefore, it is necessary to design a special action space and reward function for each data stream with different QoS requirements to guarantee its specific QoS performance requirements. At the same time, these streams need to share the same state space because they all face the same bottleneck network environment to compete for resources. In the present application, a multi-heterogeneous agent DRL model is used to design a customized congestion control scheme to adapt to different QoS requirements.
[0179] By adopting semi-independent training, the state and action information of neighbor agents are introduced, the observation ability and decision performance of the agent are enhanced, a coordination reward item based on the state and action of neighbor agents is added, and the coordination reward encourages the agent to make more coordinated decisions, improves the stability of the system, and thus exhibits better performance in complex multi-agent reinforcement learning tasks.
[0180] The state space of the decision-making process is set, specifically including the following steps:
[0181] The space-time-ground network environment information and network layer information are included in the state space, and the potential bottleneck link is dynamically learned;
[0182] The state space at time t includes space-time-ground integrated network state, network layer state and transmission layer state, and is expressed as Wherein, is the state of the sub-flow f at time t, and is expressed as is the space-time-ground integrated network state of the sub-flow f at time t, is the network layer state of the sub-flow f at time t, is the transmission layer state of the sub-flow f at time t;
[0183] And Specifically expressed as
[0184]
[0185] In the formula is the air-ground link channel gain at time t; is the air-space link channel gain at time t; is the space-ground link channel gain at time t; is the ratio of the used power to the maximum power of the air-ground link at time t; is the ratio of the used power to the maximum power of the air-space link at time t; is the ratio of the used power of the space-ground link to the maximum power at time t; is the ratio of the frequency band used by the air-ground link to the total frequency band resources at time t; is the ratio of the frequency band used by the air-space link to the total frequency band resources at time t; is the ratio of the frequency band used by the space-ground link to the total frequency band resources at time t; is the number of data packets backlogged in the LEO satellite; is the number of data packets backlogged in the UAV; is the number of data packets backlogged at the satellite earth station; is the LEO satellite’s export link channel resource utilization rate; is the channel resource utilization rate of the UAV’s export link; is the channel resource utilization rate of the satellite earth station’s export link; sr t f is the sending rate of sub-flow f at time t; is the throughput of subflow f at time t; is the average RTT of subflow f at time t; is the average RTT deviation of subflow f at time t; is the congestion window size of subflow f at time t; is the congestion window change between time t-1 and time t of subflow f; is the throughput change between time t-1 and time t of subflow f; is the ratio of the minimum RTT of subflow f to the RTT at time t; lr t f is the loss rate of subflow f at time t; is the RTT of sub-flow f at time t; is the arrival time interval of the ACKs returned between time t-1 and time t of sub-flow f;
[0186] The data backlog and egress link utilization of LEO satellites refer specifically to the values when they send data to satellite earth stations or drone relay stations. Similarly, the data backlog and egress link utilization of satellite earth stations or drone relay stations refer to the values when they send data to wireless access points.
[0187] Setting the action space of the decision-making process includes the following steps:
[0188] Compared with the existing congestion control method by updating the size of congestion window, adjusting the sending rate in time can deal with the congestion event more directly and effectively; however, in the actual system, the transmission operation of the transport layer is carried out in the granularity of data packets; therefore, it is an effective method to convert the sending rate limit into the number of data packets of the determined congestion window size; on this basis, considering different congestion control objectives, based on the same sending rate limit, the determined congestion control window size should be different;
[0189] Therefore, the action expression based on the sending rate control is set as
[0190]
[0191] In the formula, a is the action space expression of the sub-flow f at time t; Δsr t f is the sending rate conversion value, and Δsr t f ≥ 0 corresponds to the sending rate increment value, and Δsr t f < 0 corresponds to the sending rate decrement value; is the repetition factor of the launched data packet on the sub-flow f at time t;
[0192] According to the action expression and the current state, the sending rate update action is converted into the congestion window update action, which is expressed as
[0193]
[0194] In the formula, L MTU is the length of the maximum transmission unit of the transport layer;
[0195] Therefore, the global action space a t at time t is expressed as
[0196] The reward function of the decision-making process is set, which specifically includes the following steps:
[0197] The reward function should be designed according to the actual needs of the upper application program; since there are data streams with different congestion control objectives, each data stream with a specific congestion control objective should have its own unique reward function; in addition, in order to avoid affecting the interests of other flows due to greed, the overall benefits of the system must be considered when designing the individual reward function;
[0198] Therefore, the reward function of the sub-flow f is set as
[0199]
[0200] In the formula, r tf is the cumulative reward value; is the trade-off coefficient used to weigh the benefits between sub-flow f and other flows; Indicates the utility function of the sub-flow considering single indicator factors or multiple indicator factors;
[0201] Divide the entire state-action sequence of a time into several equal and length L SAS short sequence, and estimate the cumulative reward of each short sequence; for the time from t to t+L SAS -1 The sequence ends at the moment, the corresponding cumulative reward Expressed as The cumulative reward is clipped to ensure that the value range is [-1, 1];
[0202] S4. Construct the state prediction module, reward redistribution module, and optimization strategy module of the decision-making process; specifically, the following steps are included:
[0203] Each agent receives three types of state information from the environment: the air-ground integrated network state (GASN), the network layer state (NL), and the transport layer state (TL). These state information are first predicted by the GASN-NL state prediction module to obtain the latest GASN and NL states;
[0204] The predicted GASN-NL state and the TL state of all flows are then combined as the global state and fed into the representation network to extract the feature vector. During action selection, each agent uses a classic actor-critic (AC) architecture to optimize its utility function by receiving information about the GASN, NL, and TL states associated with its responsible flows and feature vectors, thereby determining the sending rate of the specific flow it controls.
[0205] During the specific training process, the reward redistribution module redistributes the cumulative rewards provided by the environment in a specific period to each state-action pair, ultimately forming an experience sample that can be directly learned (i.e., state, action, redistributed reward, next state); the overall framework is shown in the attached Figure 3 As shown;
[0206] Constructing the state prediction module of the decision-making process includes the following steps:
[0207] Obtain the air-space-ground integrated network status, network layer status, and transport layer status provided by the current environment, and use the built status prediction module to predict and obtain the latest air-space-ground integrated network status and network layer status;
[0208] The ground user can obtain the space-air-ground integrated network state and the network layer state, and transmit the obtained space-air-ground integrated network state and the network layer state to the sender of the response data packet together with the idle field domain of the ACK data packet protocol header;
[0209] In each cycle, each sender obtains a plurality of space-air-ground integrated network state and network layer state sequences with timestamps for predicting the latest space-air-ground integrated network state and network layer state;
[0210] The indicators of the predicted space-air-ground integrated network state and network layer state include channel gain, ratio of used power to maximum power, ratio of used frequency band to total frequency band resource, data packet backlog, and exit link utilization of bottleneck node;
[0211] Five GRU-based neural network models with the same structure are used to respectively and individually predict the channel gain, the ratio of used power to maximum power, the ratio of used frequency band to total frequency band resource, the data packet backlog, and the exit link utilization of bottleneck node; meanwhile, the time interval of the state is taken as a feature for prediction; the true value of the space-air-ground integrated network state and the network layer state is obtained by the LEO satellite, the unmanned aerial vehicle, the satellite ground station base station, or the ground user at each time slot;
[0212] The reward redistribution module of the decision-making process is constructed, specifically including the following steps:
[0213] The congestion control problem with dynamic delay, balance exploration, and fairness guarantee is taken as a sequence Markov decision process with delayed reward, which is represented by a five-tuple (S, A, R, P, γ), wherein S is the global state space, A is the joint action space of all individuals, R is the reward function set of all individuals, P is the state transition probability matrix, and γ is the discount factor;
[0214] Based on the second-order Markov reward redistribution framework, a sequence Markov decision process without cumulative reward is obtained from the sequence Markov decision process M and M and have the same state space, action space, state transition probability, and optimal strategy;
[0215] It is set that is the Q value function of M in the sequence Markov decision process with delayed reward M, the redistributed reward of M in the sequence Markov decision process with delayed reward M, then the second-order Markov reward redistribution is represented as
[0216]
[0217] The reward-prediction function g is setf , for predicting the expected cumulative reward of a given state-action sequence at time t; the return-prediction function g f The output of g Since g contains information about , the difference function Δ f is defined and used to calculate the information carried in ; Δ f is defined as the numerical difference between and , and based on Δ f processing, the state-action sequence is represented as
[0218]
[0219] At the same time, g f must ensure that
[0220] Thus, we have: (1) is calculated by and serves as an estimate of ; (2)
[0221]
[0222] Convert g f (Δ 0:T ) and g f (Δ 0:t ) to where is the contribution of the state-action sequence to the predicted expected reward, and is calculated by ;
[0223] Set then and there is
[0224] Since there is an error between and , set the additional re-allocated reward and
[0225]
[0226] Finally, we get
[0227] where the calculation process of the reward re-allocation module specifically includes the following steps:
[0228] For a state-action sequence representing the expected feedback before execution, representing the expected feedback after execution, and the numerical difference between them as the contribution of ; the network based on GRU is used to predict g f , and in the training process, g is taken as input, as label, and g is redistributed to each period, finally forming a sequence Markov decision process with delayed rewards which can be solved by DRL method
[0229]
[0230] Specifically, the following steps are included:
[0231] Input: state-action sequence pair feedback prediction function g f , difference function Δ f , 0≤t≤T, f∈F;
[0232] Output: state-action sequence contribution to the expected reward of prediction
[0233] A. Calculate
[0234] B. Calculate
[0235] C. Calculate get
[0236] D. Calculate get
[0237] E. Return the calculation result and
[0238] At the same time, the calculation process of the reward redistribution module is corrected, which specifically includes the following steps:
[0239] The expected value of g is expressed by the true value, and g is obtained Divide each round into several short sequences with length L SAS , and return the cumulative reward of each short sequence respectively, which will be redistributed to each action; the redistribution correction method is used to ensure that ;
[0240] Specifically, the following steps are included:
[0241] Input: Feedback prediction function Cumulative reward
[0242] Output: Redistributed reward
[0243] a. Calculate
[0244] b. Calculate uncorrected redistributed reward For Where t≠0, and
[0245]
[0246] c. Calculate average error
[0247] d. Calculate corrected redistributed reward For
[0248] e. Return
[0249] The optimization strategy module of the decision-making process is constructed, specifically including the following steps:
[0250] The feature extraction network first outputs a feature vector by taking the state of all active flows in each cycle as input Since at most L rp rounds are involved, at most L rp feature vectors are generated; the global state sequence is represented as Then, the is input into another representation network to generate a global feature vector containing global information; finally, each agent generates an action using the global feature vector; a GRU-based neural network is used as the representation network, as shown in Figure 4 ;
[0251] Each agent shares a global feature vector; the actor network parameter is used to output the action probability of the sub-flow f based on the current state, and the critic network parameter is used to update and feedback to the actor network by predicting future rewards; based on the feedback behavior, the actor network can continuously adjust the probability of selecting future actions, and the feedback process is repeatedly repeated in each learning iteration to improve the strategy;
[0252] In order to avoid the fluctuation of the training process being greater than the set range and ensure the stability of the learning process, the neural network parameters are updated as follows:
[0253]
[0254] wherein is a clipping loss function used to limit the magnitude of policy update; denotes the expected value of all rounds; is the ratio between the new and old policy, and is the probability of the new policy taking action in state , is the probability of the old policy taking action in state ; is the advantage function of action at time t, which measures the advantage of action at time t relative to the average action; is an intermediate function, and ε is a positive number set to clip the ratio to avoid the policy update being too large or too small;
[0255] is calculated using the generalized advantage estimator denotes
[0256]
[0257] wherein γ is a discount factor used to reduce the weight of future returns, λ is a parameter that controls the trade-off between bias and variance, T is the total number of periods, is the Bellman residual term and V() is the value function and is usually estimated by a neural network;
[0258] The total loss is calculated as
[0259]
[0260] wherein is the value function loss, which measures the approximation of the value function to the expected return at time t; H is the entropy function that the policy promotes exploration by hindering the deterministic policy at each time t; c1 is the first weight coefficient; c2 is the second weight coefficient; R represents the GRU-based neural network;
[0261] In the optimization process, all the sequences of logarithmic probabilities are used to output the sequence of actions to ensure fairness; the single logarithm associated with each action selected before is modified to a set value, which is used to prevent the softmax function from generating the selected operation for the logarithm; the gradient of the logarithm of the invalid operation is set to zero; the action is output using the softmax function according to the modified logarithm, denoted as
[0262] a f,j = softmax(L f)
[0263]
[0264] where a f,j is the jth action of subflow f; L f is the modified log probability sequence of subflow f, and L f ={l f,1 ,...,l f,j ,...,l f,J}, J is the number of selectable actions of subflow f; σ is a negative reward set;
[0265] Subflow f must select one action from each of the following two types of action spaces respectively:
[0266]
[0267] where is the transmission rate increment of subflow f at time t; is a negative value indicating a transmission rate decrease of subflow f, is a positive value indicating a transmission rate increase of subflow f, is 0 indicating a same transmission rate of subflow f; N is the maximum number of repetitions in a data packet; Δsr is the degree of change in transmission rate; is the repetition factor of subflow f at time t;
[0268] Therefore, the action is converted to by the following formula:
[0269]
[0270] Through the reward redistribution mechanism, it can be assumed that the action feedback in each sequence will be returned at L SAS time after the end of the sequence;
[0271] S5. According to the constructed module, the congestion control of space-air-ground network based on deep reinforcement learning is carried out.
[0272] In specific implementation, in the congestion control method of space-air-ground network based on deep reinforcement learning, the training process includes the following steps:
[0273] Input: maximum training round E, training period length T, algorithm weight λ, discount factor γ, batch size N batch ;
[0274] Output: actor network parameters critic network parameters representation network parameters
[0275] (1) Initialize actor network parameters with random parameters critic network parameters and representation network parameters
[0276] (2) Initialize replay buffer B f ;
[0277] (3) Initialize an Ornstein-Uhlenbeck stochastic process for exploratory actions
[0278] (4) At each time t, collect the space-air-ground integrated network state, network layer state and transport layer state of all flows
[0279] (5) Use the state prediction module to predict the latest chapter station to obtain the global state s t ;
[0280] (6) Obtain the feature vector h t from the latest L rp global states by GRU neural network R
[0281] (7) Obtain an action a
[0282] (8) Based on the action a , generate an action a using a stochastic process
[0283] (9) Execute the action a and observe the current utility function
[0284] (10) Store the global state s t , the feature vector h t , the action a and the corresponding utility function value
[0285] (11) Calculate
[0286] (12) Calculate the cumulative reward of each sequence according to the utility function value
[0287] (13) Redistribute the cumulative reward to obtain
[0288] (14) Store the transition process to the buffer B f ;
[0289] (15) Extract N batchsamples and divided into K batches;
[0290] (16) In each small batch, the following steps (17) to (20) are performed sequentially;
[0291] (17) Obtain the feature vector h through the representation network R k for
[0292] (18) Determine the action based on the policy function π() for
[0293] (19) Calculation and
[0294]
[0295] (20) Based on total loss Update actor network parameters Critic network parameters and characterize network parameters
[0296]
[0297] (21) Return the final actor network parameters Critic network parameters and characterize network parameters
[0298] The method of the present invention is further described below with reference to an embodiment:
[0299] Taking the end-to-end data transmission scenario in a typical air-ground integrated network and the downlink network parameters of the StarLink low-orbit satellite Internet constellation as the main reference, the simulation simulates the data transmission scenario between ground users and 22 satellites in the 550km upper orbit plane through four drones that communicate at a fixed speed for partial area coverage and a satellite earth station. The specific simulation parameters are detailed in Table 1.
[0300] Table 1 Simulation parameter diagram
[0301] Parameter Value LEO low earth orbit altitude 550km LEO low earth orbit number 22 Satellite flight speed 7.59km / s Satellite operating frequency band Ku Single carrier bandwidth 50-400MHz Antenna gain 30-40dBi Equivalent isotropically radiated power 20-50dBW Free space loss 150-200dB System noise temperature 300K Other losses 0-30dB Sensing node number 200 Sink node number 4 Drone flight height 100m Drone moving speed 30m / s Data packet segment size 1KB ACK packet size 40B
[0302] During the implementation, users initiated session flows with various QoS indicators through their PCs, including dual-indicator session flows (latency and bandwidth, jitter and packet loss rate) and single-indicator session flows (high reliability, low latency, and high bandwidth). As the responding end, the data center network is responsible for transmitting a large number of data packets, which take a long time and are data-intensive. Therefore, this example only considers performing congestion control operations on the responding end.
[0303] In the present application, first, according to different application requirements (such as video conference, file transmission, etc.), the priority of each QoS indicator is set. Ensure that the key indicators (such as delay and reliability) are given priority to guarantee in network congestion. Subsequently, a monitoring module is deployed at the response end to collect network state information in real time, including delay, bandwidth, jitter, etc. By using a sliding window update mechanism, the current network condition is dynamically evaluated. In each time slot, first, according to the state information collected from all sub-flows, the trained GRU prediction model is used to predict the latest global link state, and then the observation state of each agent is obtained. Input these states into the present application scheme for training iteration, and finally obtain reasonable network parameters, and then realize the substantial improvement of bandwidth resource utilization on the basis of acceptable transmission delay. The specific training parameters are shown in Table 2.
[0304] Table 2: Deep Reinforcement Learning Simulation Parameter Table
[0305] Parameter Value Discount factor γ 0.99 Learning rate ξ 0.0005 Replay buffer size N 5000 Batch size 32 GAE hyperparameter λ 0.95 Total training rounds 30000 [weight coefficients c1 and c2 in the reward function] 0.8,0.2 Penalty term ε in reward function 0.001 Number of neurons in hidden layer 128 Hidden layer activation function ReLU Output layer activation function Identity Number of GRU layers 1 Number of neurons in GRU hidden layer 256 Number of GRU training rounds 10000
[0306] The present application uses python+pytorch tools to realize the above scheme, and obtains the results shown in Figure 5 and Figure 6 .
[0307] Figure 5 The present application method and the model using only ISIPPO module are shown. It can be seen that the present application method increases the reward redistribution strategy and state prediction module, so that the system performance and stability of the model are significantly improved. Specifically, these improvements help the model to allocate resources more effectively and adapt to dynamic network environments, thereby enhancing the overall response capability. From Figure 5 , it can be observed that compared with not using the reward redistribution strategy, after using the strategy, the model can reach the convergence state more quickly. This accelerated convergence effect shows that the algorithm is more accurate in adjusting the strategy, reducing the uncertainty in the training process. In addition, the fluctuation and spike in the later stage of the graph are significantly reduced, reflecting the improvement of the stability of the model in long-term operation. This stability is crucial for practical applications, as it ensures the reliability and sustained performance of the system under various network conditions.
[0308] In order to further illustrate the beneficial effects of the present application, the above space-air-ground network hybrid path is abstracted into a single special link, whose delay and packet loss rate are the delay and packet loss rate of the total path, and the bandwidth is the bottleneck link bandwidth, and based on this link, the classic congestion control scheme Cubic algorithm, the optimization-based scheme BBR, and the artificial intelligence-based scheme DRL_CC are obtained as Figure 6The results are shown. As can be seen from the figure, first, in terms of delay and bandwidth utilization, the method of the application can effectively dynamically adjust its sending rate, significantly reducing the frequency of network congestion. In contrast, the Cubic algorithm performs poorly under high delay and rapidly changing network conditions, easily leading to packet loss and retransmission, thereby reducing overall throughput. Although BBR has improved in bandwidth utilization, its sending strategy may cause excessive delay fluctuations when the network load is heavy. Second, the method of the application can monitor network status in real time and make adaptive adjustments by introducing a reward redistribution mechanism and a state prediction module. This mechanism enables the method of the application to take preventive measures before congestion occurs, thereby effectively reducing the possibility of packet loss and delay increase. Although DRL_CC optimizes the congestion control strategy using deep reinforcement learning, its training and inference process is complex and has certain limitations in real-time performance. Finally, by comparing the stability of different algorithms under long-term operation, the method of the application shows better stability and consistency. This shows that the method of the application not only improves instantaneous performance, but also maintains efficient network resource utilization in the long term. In summary, the method of the application has shown significant advantages in the field of congestion control, with higher adaptability and stability, suitable for complex and variable network environments.
[0309] The application proposes a new congestion control method based on heterogeneous multi-agent deep reinforcement learning (MHADRL) from the perspective of cross-layer, in addition to transport layer (TL) information, GASN and network layer (NL) information are also included. This helps each agent to directly understand the information of potential bottlenecks and take measures to prevent congestion in GASN. At the same time, the application also proposes a method to predict the latest GASN-NL information from the delayed GASN-NL information.
[0310] The application uses the proposed model to make congestion control decisions for each active flow, and uses a gated recurrent unit (GRU) based model to learn the representation of all active flows with variable number and different quality of service (QoS) characteristics. At the same time, the GRU-based method is used to extract the GASN-NL and TL information of all active flows into a fixed-size feature vector. This feature vector is concatenated with the personalized information of each target flow as the input of each agent.
[0311] The performance of the application in key performance indicators such as throughput and latency has been greatly improved compared to the baseline. The effectiveness and robustness of the method of the application in handling congestion control of multiple data streams with heterogeneous optimization targets have been proven.
Claims
1. A method for air-ground network congestion control based on deep reinforcement learning, comprising the following steps: S1. Obtain data information of the target air-space-ground network; S2. Based on the data information obtained in step S1, set the optimization goal and perform modeling; During implementation, optimization objectives include user application throughput, latency, jitter, packet loss rate, and reliability. Congestion control is defined as a decision-making problem that considers dynamic latency, seeks equilibrium, and ensures fairness. S3. Set the state space, action space and reward function of the decision process; In the specific implementation, the state space includes the state of the space-air-ground integrated network (GASN), the network layer (NL), and the transport layer (TL). Based on the action expression of the sending rate control and the current state, the sending rate update action is converted into a congestion window update action. The reward function of the subflow f is set, and the entire state-action sequence of a time is divided into several short sequences of the same length LSAS, and the cumulative reward of each short sequence is estimated. S4. Construct the state prediction module, reward redistribution module and optimization strategy module of the decision-making process; In specific implementation, the state prediction module predicts the latest GASN and NL states, merges the predicted GASN and NL states with the TL states of all flows as the global state, and inputs it into the representation network to extract the feature vector; During action selection, each agent uses a classic actor-critic architecture to optimize its utility function by receiving GASN, NL, and TL state information related to the stream and feature vector it is responsible for, thereby determining the sending rate of the specific stream it controls. During training, the reward redistribution module redistributes the accumulated rewards provided by the environment at a specific time period to each state-action pair to form experience samples that can be directly learned. A GRU-based neural network is used as the representation network. S5. Based on the constructed modules, perform congestion control of the air-space-ground network based on deep reinforcement learning.
2. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 1 is characterized in that Step S2, based on the data information obtained in step S1, sets the optimization target and performs modeling, specifically including the following steps: The optimization goals set include user application throughput, latency, jitter, packet loss rate, and reliability; Throughput refers to the amount of data successfully transmitted per unit time, and is used to reflect the transmission capacity of the network and the data transmission efficiency of the application; Latency indicates the total time from sending to receiving data, and is used to reflect the response speed and real-time performance of user applications; Jitter indicates the variation in the time interval between arrival of consecutive data packets and is used to reflect the playback quality of streaming media applications. The packet loss rate indicates the ratio of data packets lost during data transmission to the total number of data packets sent, which is used to reflect the reliability and integrity of user applications. Reliability refers to the ability of a network system to provide uninterrupted services within a given period of time, and is used to reflect the continuity and availability of user applications. Let F be the set of flows with at least one and at most two optimization objectives, T f (t) is the long-term throughput of subflow f started by a single terrestrial user measured at time t, L f (t) is the long-term packet loss rate of subflow f initiated by a single terrestrial user measured at time t and is defined as Where L ep is the number of cycles considered; is the throughput measured in the kth cycle; η is the set attenuation coefficient; The packet loss rate measured for the kth period; Using D f (t) represents the delay of sub-flow f measured at time t, J f (t) represents the jitter of sub-flow f measured at time t, R f (t) represents the reliability of sub-flow f measured at time t; the single-index utility function at time t is Set to Where SI represents the single indicator factor considered; T represents the throughput indicator; L represents the packet loss rate indicator; D represents the delay indicator; J represents the jitter indicator; R represents the reliability indicator; The sum of the single-index utilities of all flows of F except subflow f at time t Expressed as Where f' represents any flow in F except subflow f; is the single-index utility function of f' at time t; Two dual-index utility functions are used, expressed as In the formula It is a utility function that considers both throughput and latency. is the throughput weight; is the delay weight; It is a utility function that considers both reliability and packet loss rate; is the reliability weight; is the packet loss rate weight; U α () is the fairness coefficient function, and α is the fairness parameter; The double index utility of all flows of F except subflow f at time t and Expressed as 3. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 2 is characterized in that The state space of the decision-making process is set in step S3, which specifically includes the following steps: Incorporate air-space-ground network environment information and network layer information into the state space to dynamically learn potential bottleneck links; The state space at time t includes the air-space-ground integrated network state, the network layer state, and the transport layer state, which is expressed as in, is the state of subflow f at time t, and is expressed as is the space-ground integrated network state of subflow f at time t, is the network layer state of subflow f at time t, is the transport layer state of sub-flow f at time t; and Specifically expressed as In the formula is the air-ground link channel gain at time t; is the channel gain of the air-space link at time t; is the space-ground link channel gain at time t; is the ratio of the used power of the air-ground link to the maximum power at time t; is the ratio of the used power of the air-space link to the maximum power at time t; is the ratio of the used power of the space-ground link to the maximum power at time t; is the ratio of the frequency band used by the air-ground link to the total frequency band resources at time t; is the ratio of the frequency band used by the air-space link to the total frequency band resources at time t; is the ratio of the frequency band used by the space-ground link to the total frequency band resources at time t; is the number of data packets backlogged in the LEO satellite; is the number of data packets backlogged in the UAV; is the number of data packets backlogged at the satellite earth station; is the LEO satellite’s export link channel resource utilization rate; is the channel resource utilization rate of the UAV’s export link; is the egress link channel resource utilization rate of the satellite earth station; is the sending rate of sub-flow f at time t; is the throughput of subflow f at time t; is the average RTT of subflow f at time t; is the average RTT deviation of subflow f at time t; is the congestion window size of subflow f at time t; is the congestion window change between time t-1 and time t of subflow f; is the throughput change between time t-1 and time t of subflow f; is the ratio of the minimum RTT of subflow f to the RTT at time t; is the loss rate of subflow f at time t; is the RTT of sub-flow f at time t; is the arrival time interval of the ACKs returned between time t-1 and time t of sub-flow f.
4. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 3 is characterized in that The action space of the decision-making process is set in step S3, which specifically includes the following steps: Set the action expression based on the sending rate control, expressed as In the formula is the action space expression taken by subflow f at time t; is the sending rate change value, and Corresponding to the increasing value of the sending rate, Corresponding to the sending rate decrement value; is the repetition factor of the data packet launched on sub-flow f at time t; According to the action expression and the current state, the sending rate update action is converted into the congestion window update action, which is expressed as Where L MTU is the length of the maximum transmission unit of the transport layer; Therefore, the global action space a at time t t Expressed as The reward function of the decision-making process in step S3 is set, which specifically includes the following steps: Set the reward function of subflow f to In the formula is the cumulative reward value; is the trade-off coefficient used to weigh the benefits between sub-flow f and other flows; represents the utility function of subflow f considering single indicator factors or multiple indicator factors; Divide the entire state-action sequence of a time into several equal and length L SAS short sequence, and estimate the cumulative reward of each short sequence; for the time from t to t+L SAS -1 The sequence ends at the moment, the corresponding cumulative reward Expressed as The cumulative reward is clipped to ensure that the value range is [-1,1].
5. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 4 is characterized in that The state prediction module of the decision-making process in step S4 specifically includes the following steps: Obtain the air-space-ground integrated network status, network layer status, and transport layer status provided by the current environment, and use the built status prediction module to predict and obtain the latest air-space-ground integrated network status and network layer status; It is assumed that the ground user can obtain the integrated air-space-ground network status and network layer status, and transmit the obtained integrated air-space-ground network status and network layer status to the sender of the response data packet together with the idle field of the ACK data packet protocol header; In each cycle, each sender obtains several time-stamped sequences of air-space-ground integrated network states and network layer states to predict the latest air-space-ground integrated network states and network layer states; Indicators for predicting the status of the air-space-ground integrated network and the network layer include channel gain, the ratio of used power to maximum power, the ratio of used frequency band to total frequency band resources, the amount of packet backlog, and the utilization rate of the bottleneck node's egress link; Five GRU-based neural network models with the same structure are used to separately predict channel gain, the ratio of used power to maximum power, the ratio of used frequency band to total frequency band resources, the amount of packet backlog, and the bottleneck node's egress link utilization. The time interval between states is also incorporated into the prediction as a feature. The true values of the air-ground integrated network status and network layer status are obtained by LEO satellites, UAVs, satellite ground station base stations or ground users in each time slot.
6. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 5 is characterized in that The reward redistribution module of the decision-making process in step S4 specifically includes the following steps: The congestion control problem with dynamic delay, balance exploration, and fairness guarantee is formulated as a sequential Markov decision process, represented by a five-tuple (S, A, R, P, γ), where S is the global state space, A is the joint action space of all individuals, R is the reward function set of all individuals, P is the state transition probability matrix, and γ is the discount factor. Based on the second-order Markov reward redistribution framework, a sequential Markov decision process with delayed rewards is obtained from the sequential Markov decision process M. And M and Have the same state space, action space, state transition probability and optimal strategy; set up is a sequential Markov decision process with delayed rewards middle The Q-value function, Represents a sequential Markov decision process with delayed rewards middle The redistribution reward of , then the second-order Markov reward redistribution is expressed as Set the reward-prediction function g f , used to predict the expected cumulative reward of M at the end of a given state-action sequence at time t; the reward-prediction function g f The output is because Contains Therefore, the difference function Δ f and used to calculate The information carried in Δ f Defined as and The numerical difference between f The state-action sequence of the process is represented as At the same time, g f Must ensure From this we can get: (1) Depend on Calculated and used as Estimates; (2) Will g f (Δ 0:T ) and g f (Δ 0:t ), converted to in, State-Action Sequence The contribution to the expected reward of the prediction, and is given by Calculated; set up but and There exists because and There is an error between the two, so set additional redistribution rewards and Finally got 7. The method for controlling air-ground-space network congestion based on deep reinforcement learning according to claim 6, characterized in that The calculation process of the reward redistribution module includes the following steps: For a state-action sequence Indicates the expected feedback before execution. Indicates the expected feedback after execution, and The numerical difference between Contribution of GRU-based network to g f To make predictions, during the training process, As input, As a label, Redistributed to each cycle, it eventually forms a sequential Markov decision process with delayed rewards that can be solved by the DRL method. The specific steps include: Input: state-action sequence pairs Feedback prediction function g f , the difference function Δ f , 0≤t≤T,f∈F; Output: state-action sequence Contribution to the predicted expected reward A. Calculation B. Calculation C. Calculation get D. Calculation get E. Return calculation results and 8. The method for controlling air-ground-space network congestion based on deep reinforcement learning according to claim 7, characterized in that Correct the calculation process of the reward redistribution module, including the following steps: Will The expected value of is expressed as the true value, and we get Divide each round into several blocks of length L SAS The cumulative reward of each short sequence will be redistributed to each action; the redistribution correction method is used to ensure Established; The specific steps include: Input: Feedback prediction function Accumulated Rewards Output: Redistribution reward a. Calculation b. Calculate the uncorrected redistribution reward Where t≠0, and c. Calculate the average error d. Calculate the revised redistribution reward for e.Return 9. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 8 is characterized in that The optimization strategy module for constructing the decision-making process in step S4 specifically includes the following steps: The feature extraction network first takes the state of all activity flows in each cycle as input and outputs a feature vector Since most of the L rp rounds, so at most L rp feature vectors; the global state sequence is expressed as Then, The input is fed into another representation network to generate a global feature vector containing global information. Finally, each agent uses the global feature vector to generate actions. A GRU-based neural network is used as the representation network. Each agent shares a global feature vector; the actor network parameters The action probability of the output sub-flow f based on the current state, the critic network parameters Used to update and feed back to the actor network by predicting future rewards; Based on feedback behavior, the actor network can continuously adjust the probability of future action selection, and continuously repeat the feedback process in each learning iteration to improve the strategy; In order to prevent fluctuations in the training process from exceeding the set range and to ensure the stability of the learning process, the following formula is used to update the neural network parameters: In the formula is the clipping loss function, which is used to limit the amplitude of the strategy update; represents the expected value of all rounds; is the ratio between the old and new strategies, and For the new strategy in state Take action The probability of For the old policy in state Take action probability; is the advantage function of the action at time t, which is used to measure the advantage of the action at time t relative to the average action is an intermediate function, and ε is a set positive number; Calculated using the generalized advantage estimator Expressed as Where γ is the discount factor used to reduce the weight of future returns, λ is the parameter that controls the trade-off between bias and variance, T is the total number of periods, is the Bellman residual term and V() is the approximate value function; The total loss is calculated as In the formula is the value function loss, which is used to measure the approximation between the value function and the expected return at time t; H is the entropy function of the strategy at each time t that promotes exploration by hindering the deterministic strategy; c1 is the first weight coefficient; c2 is the second weight coefficient; R represents the GRU-based neural network; During the optimization process, all logarithmic probability sequences are used to output action sequences to ensure fairness; the single logarithm associated with each previously selected action is modified to a set value to prevent the softmax function from generating selected actions for the logarithm; the gradient of the logarithm of invalid operations is set to zero; the softmax function is used to output the action based on the modified logarithm, which is expressed as a f,j =softmax(L f ) Where a f,j is the j-th action of subflow f; L f is the modified logarithmic probability sequence of sub-flow f, and L f ={l f,1 ,...,l f,j ,...,l f,J }, J is the number of actions that can be selected by subflow f; σ is the set negative reward; Subflow f must select an action from each of the following two types of action spaces: In the formula is the sending rate increment of sub-flow f at time t; A negative value indicates that the sending rate of sub-flow f is reduced. A positive value of indicates that the sending rate of sub-flow f is increased. A value of 0 indicates that the sending rate of the sub-flow f is the same; N is the maximum number of repetitions in a data packet; Δsr is the degree of change in the sending rate; is the packet repetition factor of subflow f at time t; Therefore, the action Converted from the following formula The action space expression for controlling the sending rate is constructed as follows: Through the reward redistribution mechanism, it is possible to assume that the action feedback in each sequence will be L after the end of the sequence. SAS Return at any time.
10. The air-ground-space network congestion control method based on deep reinforcement learning according to claim 9 is characterized in that The training process of the air-ground-space network congestion control method based on deep reinforcement learning includes the following steps: Input: maximum training round E, training cycle length T, algorithm weight λ, discount factor γ, batch size N batch ; Output: Actor network parameters Critic network parameters Characterizing network parameters (1) Initialize the actor network parameters using random parameters Critic network parameters and characterize network parameters (2) Initialize replay buffer B f ; (3) Initialize an Ornstein-Uhlenbeck random process for exploratory actions; (4) At each time t, collect the space-space integrated network state, network layer state and transport layer state of all flows; (5) Use the state prediction module to predict the latest chapter and obtain the global state s t ; (6) Through the GRU neural network R from the nearest L rp The global state of the segment is obtained by the feature vector h t ; (7) Get an action based on the strategy function π() (8) Action-based Generate actions using a random process (9) Execute action And observe the current utility function; (10) Store the global state s t , eigenvector h t ,action and the corresponding utility function value; (11) Calculation (12) Calculate the cumulative reward of each sequence based on the utility function value; (13) Through the calculation process of the reward redistribution module, the accumulated rewards are redistributed to obtain (14) Transfer process Store to buffer B f ; (15) Extract N from the replay buffer batch samples and divided into K batches; (16) In each small batch, the following steps (17) to (20) are performed sequentially; (17) Obtain the feature vector h through the representation network R k for (18) Determine the action based on the policy function π() for (19) Calculation and (20) Based on total loss Update actor network parameters Critic network parameters and characterize network parameters (21) Return the final actor network parameters Critic network parameters and characterize network parameters
Citation Information
Patent Citations
Packet transmission method and system based on reinforcement learning and stream coding driving
CN112822718A
Reinforcement learning with optimization-based policy
US20230289612A1