A Congestion Control Method for Space-Air-Ground Networks Based on Deep Reinforcement Learning

Through the deep reinforcement learning method of aerospace and earth network congestion control, the multi-agent Markov decision-making process and composite reward function are used, combined with the centralized training-distributed execution framework, the problems of slow response and low resource utilization of congestion control in the satellite network are solved, and efficient network resource optimization and rapid response are achieved.

CN120223635BActive Publication Date: 2025-07-22NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510664097.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Traditional network congestion control methods face problems in satellite networks with slow response, low resource utilization and difficulty in node collaboration. Especially in the aerospace and earth network environment, it is difficult to achieve efficient global optimization and distributed execution.

Method used

The aerospace and earth network congestion control method based on deep reinforcement learning is adopted, and the composite reward function is designed through multi-agent Markov decision-making process modeling, and a centralized training-distributed execution framework is adopted to realize collaborative optimization queue management between nodes, dynamically adjust service rates, and adapt to link capacity fluctuations in combination with the MPLS traffic engineering protocol.

Benefits of technology

It realizes fast dynamic response and global resource optimization, significantly improves network throughput, reduces transmission delay and packet loss rate, and adapts to the high dynamic topology and resource-constrained characteristics of satellite networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223635B_ABST
    Figure CN120223635B_ABST
Patent Text Reader

Abstract

The present invention discloses a congestion control method for air-space-ground networks based on deep reinforcement learning, which realizes fast dynamic response and global resource optimization through distributed node cooperation and dynamic service rate adjustment. This method is modeled based on the multi-node agent Markov decision process, designs a composite reward function, and uses it to drive the training of the policy optimization algorithm. The framework adopts a "centralized training - distributed execution" framework. During the centralized training stage, the parameters of the policy network are optimized, and during the distributed execution stage, each node independently adjusts the service rate of each node based on local observations without global communication. In view of the high-dynamic topology and resource-limited characteristics of satellite networks, the MPLS traffic engineering protocol is used to construct label-switched paths, which can dynamically adapt to link capacity fluctuations and satellite orbit movements. The present invention realizes accurate queue control for congestion control of air-space-ground nodes and significantly improves the throughput of the network, providing a reliable solution for efficient congestion control of air-space-ground networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network communication technologies, and in particular, to a space-air-ground network congestion control method based on deep reinforcement learning, aiming to achieve network congestion control and performance optimization by dynamically adjusting the local queue length and cooperation strategy of each node in the space-air-ground network. Background Art

[0002] With the acceleration of the global digitalization process, emerging services such as instant communication, high-definition video streaming, Internet of Things (IoT), and industrial Internet have put forward higher requirements for network bandwidth, latency, and reliability. Traditional terrestrial communication networks are limited by high fiber deployment costs and limited coverage, making it difficult to meet the high-reliability connection requirements in remote areas (such as oceans, deserts, and mountains) and emergency communication scenarios. Against this background, satellite communication networks, with their advantages of wide-area coverage, flexible deployment, and strong anti-destruction capabilities, have become the core components of building space-air-ground integrated networks. In particular, the rapid development of low-earth orbit (LEO) satellite constellations has made global seamless coverage and low-latency communication possible. However, the high-dynamic topology, long propagation delay, and limited on-board resources of satellite networks also pose great challenges to network congestion control.

[0003] In traditional terrestrial networks, congestion control mainly relies on the Transmission Control Protocol (TCP) and its variants (such as TCP Reno, CBBIC), and dynamically adjusts the sending rate through end-to-end feedback mechanisms (such as round-trip delay RTT, packet loss detection). However, such methods face significant limitations in satellite networks: First, the propagation delay of satellite links is much higher than that of terrestrial networks, resulting in a serious lag in congestion feedback; second, the high-speed dynamic movement of satellite nodes makes the network topology continuously change dynamically, and traditional congestion control algorithms based on static routing are difficult to adapt; moreover, the on-board computing and storage resources are limited, and complex control algorithms may exceed the capabilities of on-board processors.

[0004] To address these challenges, the academic community has proposed various improvement schemes. At the protocol level, Explicit Congestion Notification (ECN) and fast retransmission mechanisms have been introduced to reduce unnecessary waiting time; at the architecture level, Software-Defined Network (SDN) technology provides a global view through a centralized controller, but its dependence on a ground control center may fail due to long delays in satellite networks. In addition, bursty services (such as disaster emergency communication) may lead to a sudden increase in local traffic. These factors make it difficult for traditional congestion detection methods based on fixed thresholds or linear models to accurately reflect the actual network state.

[0005] In recent years, machine learning methods have gradually been applied to the field of network resource management due to their strong environmental adaptability and potential for decision optimization. Among them, the deep reinforcement learning algorithm, as an important branch of reinforcement learning, solves the problem of unstable training of traditional policy gradient algorithms through a policy update and pruning mechanism. Its continuous action space characteristic is particularly suitable for the fine adjustment of network rate control. However, existing research mainly focuses on single-agent scenarios. How to extend it to large-scale distributed systems such as satellite networks and resolve the contradiction between local node observations and global optimization remains an open question. Especially in the space-air-ground network environment, each satellite node can only obtain limited neighbor information but needs to make decisions beneficial to the global performance, which poses higher requirements for algorithm design. Therefore, there is an urgent need for a new intelligent congestion control mechanism that can balance global optimization and distributed execution service rates in an environment with dynamic topology, long delays, and limited resources. Summary of the Invention

[0006] The object of the present invention is to overcome the deficiencies of existing network congestion control technologies in dealing with space-air-ground networks, such as slow response, low resource utilization, and difficult node cooperation. A network-internal congestion control method is provided to optimize queue management through intelligent cooperation of nodes in the space-air-ground network, achieve efficient network congestion prevention and mitigation, and improve network resource utilization and communication quality.

[0007] A space-air-ground network congestion control method based on deep reinforcement learning according to the present invention includes the following steps:

[0008] Step 1: Modeling of the space-air-ground network and definition of the global state space. Establish a dynamic behavior model of the space-air-ground network and define the global state space and local observations.

[0009] Step 2: Design of a multi-dimensional composite reward function. Guide agents to cooperate and optimize congestion control through a multi-dimensional reward mechanism.

[0010] Step 3: Construction of a centralized training-distributed execution framework. Optimize policy parameters through centralized training and perform distributed real-time decision-making.

[0011] Step 4: Dynamic adjustment of service rate and node queue management. Respond to traffic changes in real time and balance the load between nodes.

[0012] Step 1 provides a state input basis for the reward function design in Step 2 and the policy training in Step 3.

[0013] The composite reward in Step 2 drives the centralized policy optimization in Step 3 through multi-objective coordination.

[0014] Step 3 transfers the global optimization policy to the distributed execution in Step 4 through parameter sharing.

[0015] The dynamic rate adjustment and topology adaptation in step 4 directly depend on the output of the policy network in step 3, and enhance the network robustness through the MPLS traffic protocol, forming a closed-loop control process from modeling, training, execution to dynamic adaptation.

[0016] Furthermore, it includes

[0017] (1) Network modeling and state space construction: Model the space-air-ground network as a multi-agent Markov decision process (MDP), with each node as an independent agent, and construct the global state space , including the normalized queue length of each node , link service rate and traffic fluctuation ;

[0018] (2) Reward function design: Design a composite reward function including directional reward , smoothness reward and queue efficiency reward, and the formula is:

[0019]

[0020] where , is a hyperparameter used as the baseline for service rate selection; if the remaining queue capacity of node is more than that of its neighbor node , selecting a service rate lower than can relieve the pressure on downstream neighbors;

[0021] ; used to evaluate the rationality of the port service rate, is used to measure the relationship between the service rate of the upstream router and the change in the queue length of the downstream router ; indicates that the smaller the difference in queue length brought by the selected action at time , the greater the reward; indicates that keeping the queue length as small as possible can also increase the reward;

[0022] (3) Policy optimization training: Adopt a centralized training - distributed execution framework, and optimize the policy network parameters by maximizing the cumulative reward function , and ensure the stability of policy updates by clipping the policy update ratio and the advantage function , and the formula is:

[0023]

[0024] wherein is the clipping range. To ensure the stability of policy updates, the clipping range is ;

[0025] (4)Dynamic adjustment of service rate: During the distributed execution phase, each node independently adjusts the service rate of the output port of the node based on local observations , and the formula is:

[0026]

[0027] wherein is determined by the output of the policy network.

[0028] Furthermore, step 1 includes

[0029] First, model the space-air-ground network as a multi-agent Markov decision process MDP, and regard each node as an independent agent; among them, set the satellite node to periodically call method to simulate the link switching caused by orbital movement; distinguish the queue capacity and link bandwidth of ground nodes and satellite nodes, optimize resource allocation, and finally construct the global state space , including the normalized queue length of each node ; the normalized service rate of each link wherein is the link capacity; at the same time, construct the local observation space for each node, which includes its own queue length , the historical traffic of neighbor nodes , providing an input basis for subsequent reward function design and policy optimization.

[0030] Furthermore, step 2 includes

[0031] Based on the global state and local observations defined in step 1, design a composite reward function to guide the agents to cooperate and optimize congestion control; the composite reward function contains three parts:

[0032] Directional reward Adjust the service rate direction according to the queue difference of adjacent nodes. If the service rate is reduced when the upstream node queue is longer, a positive reward is obtained;

[0033] Smoothing reward Encourage the minimization of the difference in queue lengths of adjacent nodes by evaluating the smoothness of the queue differences of adjacent nodes to avoid local congestion;

[0034] Queue efficiency reward is to encourage the router to reduce queue backlogs and reduce transmission delays;

[0035] ​The reward function drives the agent to alleviate congestion and improve network service quality by weighted fusion of the above sub-rewards.

[0036] Further, step 3 includes

[0037] Collecting global state data including queue length, service rate, and traffic fluctuation during the centralized training phase, and maximizing the cumulative reward based on the input basis of step 1 and the algorithm designed in step 2 , by clipping the policy update ratio and the advantage function to ensure training stability, and introducing an experience replay pool to store historical data and improve sample utilization;

[0038] During the distributed execution phase, each node independently adjusts the service rate based on local observations without global communication, achieving decentralized real-time decision-making.

[0039] Further, step 4 includes

[0040] The policy network trained based on step 3 outputs an adjustment value , dynamically updating the service rate of the node output port. The service rate update formula is ; At the same time, through the queue dynamic model calculate the change in queue length in real time. The queue change is jointly determined by the input traffic fluctuation and the service rate adjustment based on the policy network in step 3;

[0041] For the high-dynamic characteristics of satellite networks, dynamically adapt to satellite link capacity fluctuations through the MPLS traffic engineering protocol, periodically update the topology to cope with orbital movement, and when a link interruption is detected, trigger a fast recovery mechanism to reallocate service rates and traffic paths, optimize resource allocation and load balancing, and enhance the robustness of the network.

[0042] Further, step 4 includes smoothing the pressure on downstream nodes by buffering burst traffic. The queue model satisfies:

[0043]

[0044] where is the sampling interval, and the change in queue length is jointly determined by traffic fluctuation and service rate adjustment,

[0045] Deploy a dynamic queue model in each node, and describe the change in queue length through the following formula:

[0046]

[0047] where is the traffic splitting ratio, and are the arrival rate change of the input port and the service rate change of the output port respectively; the node adjusts the service rate of the output port in real time according to the current queue length and the queue difference of adjacent nodes, and adopts a deep reinforcement learning algorithm to balance the traffic pressure between adjacent nodes.

[0048] Furthermore, step 3 includes a deep reinforcement training process:

[0049] The experience replay mechanism stores historical data ; among which, the centralized Critic network calculates the Q value based on the global state and the joint action to guide the parameter update of the distributed Actor network, and the loss function is:

[0050]

[0051] where , is the discount factor, are the parameters of the target policy network;

[0052] In the centralized training-distributed execution framework, the centralized Critic network evaluates the global state-action value function, and the distributed Actor network outputs the service rate adjustment strategy to achieve decentralized execution.

[0053] Furthermore, in the distributed execution stage, each node adjusts the service rate only relying on local observations without global communication, and the action space is a continuous value .

[0054] Furthermore, the directional reward in the reward function measures whether the service rate adjustment direction is correct, and judges the service rate adjustment direction through the benchmark value. If > and < , the reward is positive, otherwise it is punished.

[0055] In summary, the present invention discloses a method for congestion control in a space-air-ground network based on deep reinforcement learning, which realizes fast dynamic response and global resource optimization through distributed node cooperation and dynamic service rate adjustment. The method is modeled based on the multi-node agent Markov decision process (MDP), and a composite reward function is designed, including: a directional reward (balancing the traffic pressure of adjacent nodes), a smoothness reward (minimizing queue differences), and a queue efficiency reward (reducing transmission delays), and uses this to drive the training of the policy optimization algorithm; the framework adopts a "centralized training - distributed execution" framework, optimizing the parameters of the policy network in the centralized training stage, and in the distributed execution stage, each node independently adjusts the service rate of each node based on local observations (queue length, neighbor traffic) without global communication. Aiming at the high-dynamic topology and resource-limited characteristics of satellite networks, the MPLS traffic engineering protocol is used to construct label-switched paths (LSPs), which can dynamically adapt to link capacity fluctuations and satellite orbit movements. The present invention realizes precise queue control for congestion control of space-air-ground nodes and significantly improves the throughput of the network, providing a reliable solution for efficient congestion control of space-air-ground networks.

[0056] The method for congestion control in a space-air-ground network based on deep reinforcement learning proposed by the present invention has significant technical advantages and good application effects, which are specifically reflected in the following aspects:

[0057] With the help of the distributed cooperation mechanism and real-time policy adjustment, the present invention realizes a fast dynamic response to sudden network traffic conditions. In the complex and changeable space-air-ground network environment, when sudden situations such as a sudden surge in traffic occur, the system can quickly perceive and respond, and timely adjust the service strategies of each node, thereby effectively reducing the end-to-end transmission delay.

[0058] Experiments show that the present invention can delay the occurrence time of congestion by up to 70 seconds, a 17% improvement compared to the traditional method (Baseline); the data packet loss is reduced by 5% - 23%, and the PDR effect is improved, superior to the control group baseline algorithm; through distributed cooperation and real-time policy adjustment, it can effectively cope with traffic bursts and reduce the end-to-end delay; balance the queue load between nodes, and the link utilization rate is increased to more than 25%; there is no need to modify the existing network protocol, and it is compatible with the SDN / NFV architecture.

[0059] It has significant advantages compared with traditional congestion control algorithms that only rely on local observations during execution. This reinforcement learning algorithm for policy optimization based on global information effectively improves the congestion control ability of the network and the ability to cope with the high-speed mobility of satellites, and can more accurately handle various complex network congestion scenarios.

[0060] The deep reinforcement learning algorithm adopted by the present invention can make full use of the queue length information of adjacent nodes during the training phase, enabling each node agent to learn a better cooperation strategy. Introducing this algorithm into in-network congestion control innovatively realizes the queue control cooperation of nodes in the space-air-ground network. Without relying on frequent communication between nodes, this cooperation mode can still achieve efficient congestion control, has significant advantages compared with traditional congestion control algorithms during execution, effectively improves the network congestion control ability, and reduces the communication cost and system complexity.

[0061] The present invention designs a composite reward function that combines queue equilibrium and low-latency objectives. While considering congestion mitigation, this reward function also takes into account the improvement of service quality, enabling the system to better meet users' requirements for low-latency and high-reliability communication during the process of controlling congestion. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a simulation topology design diagram of the space-air-ground network according to an embodiment of the present invention;

[0063] Figure 2 It is a schematic diagram of a queue model based on multi-protocol label switching according to an embodiment of the present invention;

[0064] Figure 3 It is a reward curve graph of deep reinforcement learning optimized queue congestion control according to an embodiment of the present invention;

[0065] Figure 4 It is a comparison graph of the verification algorithm in terms of throughput performance according to an embodiment of the present invention;

[0066] Figure 5 It is a comparison graph of the verification algorithm in terms of the maximum queue length performance of the queue according to an embodiment of the present invention;

[0067] Figure 6 It is a comparison graph of the verification algorithm in terms of packet loss performance according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] An embodiment of the present invention proposes an in-network congestion collaborative control method for a space-air-ground network based on deep reinforcement learning, which realizes fast dynamic response and global resource optimization through the collaborative learning of distributed nodes and dynamic service rate adjustment. The specific technical solutions are as follows:

[0070] (1) Network Modeling and State Space Construction: Model the space-air-ground network as a multi-agent Markov decision process (MDP), with each node as an independent agent, and construct the global state space , which includes the normalized queue length of each node , the link service rate and traffic fluctuations ;

[0071] (2) Reward Function Design: Design a composite reward function that includes a directional reward , a smoothness reward and a queue efficiency reward. The formula is:

[0072]

[0073] where , is a hyperparameter used as a baseline for service rate selection; if the remaining queue capacity of node is more than that of its neighbor node , selecting a service rate lower than can relieve the pressure on downstream neighbors;

[0074] ; used to evaluate the rationality of the port service rate, is used to measure the relationship between the service rate of the upstream router and the change in the queue length of the downstream router ; indicates that the smaller the difference in queue length brought by the selected action at time , the greater the reward; indicates that keeping the queue length as small as possible can also increase the reward;

[0075] (3) Policy Optimization Training: Adopt a centralized training - distributed execution framework to optimize the policy network parameters by maximizing the cumulative reward function . Ensure the stability of policy updates by clipping the policy update ratio and the advantage function . The formula is:

[0076]

[0077] where is the clipping range. To ensure the stability of policy updates, the clipping range is ;

[0078] (4) Dynamic Adjustment of Service Rate: In the distributed execution stage, each node is based on local observations Independent adjustment of the service rate of the node output port , the formula is:

[0079]

[0080] where is determined by the output of the policy network.

[0081] In the embodiment of the present invention, a deep reinforcement learning algorithm is adopted to construct an intelligent congestion control framework of "centralized training - distributed execution". In the centralized training stage, the system collects the queue lengths of each node , service rates , traffic fluctuations and other network state data to construct a global state space , and maximizes the cumulative reward function through a policy algorithm to optimize the policy network parameters . By trimming the policy update ratio and the advantage function , the stability of policy updates is ensured. The formula is:

[0082]

[0083] where =0.2 is the trimming range. In the distributed execution stage, each node only depends on local observations such as its own queue length and the load of neighboring nodes) to independently adjust the service rate, reducing communication overhead. Based on the continuous action value output by the policy network, the node traffic is dynamically regulated to achieve real-time response.

[0084] In this algorithm, each space-air-ground network node is regarded as an independent agent, which has the functions of local state perception, dynamic service rate adjustment, and local queue management. The node continuously collects the arrival rate of the input port , the service rate of the output port , the queue length , and the historical traffic data of adjacent nodes . Based on the continuous action value output by the policy network, the service rate of each output port is dynamically adjusted , the formula is:

[0085]

[0086] where is the adjustment value output by the policy network. By caching the burst traffic to smooth the pressure on downstream nodes, the upper limit of the satellite node queue capacity is set , and the queue update formula is:

[0087]

[0088] wherein is the sampling interval, which is jointly determined by traffic fluctuations and rate adjustments.

[0089] Deploy a dynamic queue model at each node, and describe the change of queue length through the following formula:

[0090]

[0091] wherein is the traffic splitting ratio, and are the change of arrival rate at the input port and the change of service rate at the output port respectively. The node adjusts the service rate of the output port in real time according to the current queue length and the queue difference of adjacent nodes, and adopts a deep reinforcement learning algorithm to balance the traffic pressure between adjacent nodes.

[0092] In the embodiment of the present invention, each node is defined as an agent, and its observation space is (including its own queue length and traffic statistics of neighbor nodes), and the action space is (adjustment of service rate of output port). The observation state of each node includes the local queue length and the historical packet forwarding volume of adjacent nodes ; the action space is the continuous adjustment value of the service rate of the output port. Based on the above parameters, a reward function is designed, which includes three parts and takes into account congestion control and communication quality: directional reward measures whether the adjustment direction of the service rate is correct, and the formula is:

[0093]

[0094] is a hyperparameter used as the baseline for service rate selection. If the remaining capacity of the queue of node is more than that of neighbor node more, select a service rate lower than to relieve the pressure on the downstream neighbor node and obtain a positive reward, otherwise it will be punished. When < is longer than longer than , an action to release packets faster should be selected. Smoothness reward is used to evaluate the rationality of the port service rate, and the formula is:

[0095]

[0096] is used to measure the upstream router The relationship between the service rate and the change in the queue length of the downstream router is as follows. It means that the smaller the difference in queue length brought by the selected action at time, the greater the reward. This is to encourage the minimization of the difference in queue length between adjacent nodes and avoid local congestion. The queue efficiency reward indicates that keeping the queue length as small as possible can also increase the reward, so as to motivate the router to reduce queue backlogs and lower transmission delays.

[0097] In the embodiment of the present invention, the centralized Critic network evaluates the global state-action value function, and the distributed Actor network outputs continuous actions (service rates). By introducing experience replay and the target network to stabilize the training process, the non-stationary problem in the multi-agent environment is solved. The centralized Critic network calculates the Q value based on the global state and the joint action to guide the parameter update of the distributed Actor network. The loss function is:

[0098]

[0099] where is the discount factor, is the parameter of the target policy network. The distributed Actor network adjusts the policy according to the local observation to output the service rate, realizing decentralized execution. By storing historical data in the experience pool and sampling regularly to update the network parameters, the training process is stabilized, supporting dynamic traffic simulation under complex network topologies (such as ring and regular graph).

[0100] In summary, the present invention discloses a space-air-ground network congestion control method based on deep reinforcement learning. Through distributed node cooperation and dynamic service rate adjustment, fast dynamic response and global resource optimization are achieved. This method is modeled based on the multi-node agent Markov decision process (MDP), and a composite reward function is designed, including: a fusion directional reward (balancing the traffic pressure of adjacent nodes), a smoothness reward (minimizing queue differences), and a queue efficiency reward (reducing transmission delays), and the policy optimization algorithm training is driven by this; in terms of the framework, a "centralized training - distributed execution" framework is adopted. In the centralized training stage, the parameters of the policy network are optimized, and in the distributed execution stage, each node independently adjusts the service rate of each node based on local observations (queue length, neighbor traffic), without global communication. For the high-dynamic topology and resource-constrained characteristics of satellite networks, the MPLS traffic engineering protocol is used to construct label-switched paths (LSPs), which can dynamically adapt to link capacity fluctuations and satellite orbit movements. The present invention realizes precise queue control for space-air-ground node congestion control and significantly improves the network throughput, providing a reliable solution for efficient congestion control of space-air-ground networks.

[0101] The following is described in conjunction with the accompanying drawings:

[0102] Figure 1 It is a topological design diagram of the space-air-ground network. Among them, a space-air-ground integrated network model is constructed, which includes two types of heterogeneous nodes: ground nodes and low-earth orbit (LEO) satellite nodes. The ground nodes form a grid or ring topology through fixed links, and the satellite nodes simulate the dynamic connection changes caused by high-speed orbital movement through a periodic link update mechanism. Different local queue capacities are configured for the ground nodes and satellite nodes, and the transmission capacities of the ground links and satellite links are distinguished. The traffic generation module injects random traffic obeying a normal distribution into each node, and introduces a dynamic service rate fluctuation mechanism to simulate the uncertainties in the actual network through noise perturbation, ensuring that the link capacity is adaptively adjusted within a reasonable range.

[0103] Figure 2 It is a schematic diagram of a queue model based on multi-protocol label switching, depicting the collaborative mechanism of label forwarding and queue management in the MPLS network. The queue state increment indicates that the current data packet arrival rate exceeds the service rate, and the queue length is accumulating. It shows that the queue resources are released, the data packet arrival rate is less than the service rate, and more traffic can be carried. The convergence steady state dynamically adjusts the LSP path and queue scheduling strategy. It constructs the network layer based on multi-protocol label switching traffic engineering (MPLS-TE), specifies the label switching path (LSP) tunnel through the forwarding equivalence class (FEC), and realizes the labeled forwarding of data packets (r1-r2-r3-r4). Each node-neighbor pair is mapped to an independent LSP tunnel, and the service rate is dynamically regulated by the local policy of the node. During the data forwarding process, the node takes out data packets from the queue according to the current service rate and sends them, and ensures the transmission stability through dynamic link capacity constraints. The observed state includes the local queue length and the historical traffic data of adjacent nodes. The action space is the continuous adjustment value of the service rate, supporting three operation modes: deceleration, hold, and acceleration.

[0104] The present invention constructs a "centralized training - distributed execution" framework using a deep reinforcement learning algorithm. In the centralized training stage, the queue length, service rate, and traffic fluctuation data of each node are collected to construct a global state space, and the policy network parameters are optimized by maximizing the cumulative reward function. The reward function includes a directional reward, a smoothness reward, and a queue efficiency reward:

[0105] Directional reward : Adjust the service rate direction according to the queue difference of adjacent nodes. If the upstream queue is longer, the service rate is reduced to relieve the downstream pressure;

[0106] Smoothness reward Avoid local congestion by minimizing the ratio of the lengths of adjacent node queues;

[0107] Queue efficiency reward Motivate to reduce queue backlog and lower transmission latency.

[0108] During the training process, optimize the policy network through historical data caching and periodic parameter updates. The target network stabilizes the training process through a soft update mechanism to ensure the convergence of the algorithm. Figure 3 It is a schematic diagram of the reward curve for algorithm training.

[0109] In the distributed execution stage, each node independently adjusts the service rate based on local observations (its own queue length and adjacent node traffic statistics) without global communication. The formula for adjusting the service rate is:

[0110]

[0111] where is output by the policy optimization network. Adapt to satellite orbital motion through periodic link state updates, and dynamically reconstruct the LSP tunnel to cope with topological changes. When a link interruption is detected, trigger a fast recovery mechanism to reallocate the service rate and adjust the traffic path to ensure network robustness.

[0112] Figure 4 , 5, 6 are schematic diagrams of the comparison results of congestion metrics. After deployment, monitor the network congestion metrics (throughput, queue length, packet loss) in real time to verify the adaptability of the nodes under traffic mutations and topological changes. Statistically analyze the performance differences between traditional methods and the deep reinforcement learning algorithm through comparative experiments. Through the above steps, the method proposed in the present invention realizes adaptive congestion control in a dynamic space-air-ground network environment, significantly reduces the risk of queue overflow, and improves the network communication quality.

[0113] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.

[0114] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method. In another embodiment provided in the present application, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute any one of the above-described space-air-ground network congestion control methods based on deep reinforcement learning.

[0115] It is understandable that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention. For the explanations, examples, and beneficial effects of the relevant content, reference can be made to the corresponding parts in the above methods. In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0116] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element. Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A congestion control method for space-air-ground network based on deep reinforcement learning, characterized in that, The method includes: Step 1: Modeling of the space-air-ground network and definition of the global state space, establishing a dynamic behavior model of the space-air-ground network, and defining the global state space and local observations; Step 2: Design of a multi-dimensional composite reward function, guiding the cooperation of agents to optimize congestion control through a multi-dimensional reward mechanism; Step 3: Construction of a centralized training - distributed execution framework, optimizing policy parameters through centralized training and making distributed real-time decisions; Step 4: Dynamic adjustment of service rate and node queue management, responding to traffic changes in real time and balancing the load between nodes; Step 1 provides the state input basis for the reward function design in Step 2 and the policy training in Step 3; The composite reward in Step 2 drives the centralized policy optimization in Step 3 through multi-objective coordination; Step 3 transfers the global optimization policy to the distributed execution in Step 4 through parameter sharing; The dynamic rate adjustment and topology adaptation in Step 4 directly depend on the output of the policy network in Step 3, and enhance the network robustness through the MPLS traffic protocol, forming a closed-loop control process from modeling, training, execution to dynamic adaptation; Specifically, it includes the following steps: (1) Network Modeling and State Space Construction: Model the space-air-ground network as a multi-agent Markov decision process (MDP), with each node as an independent agent, and construct the global state space , which includes the normalized queue length of each node , link service rate and traffic fluctuation ; (2)Reward function design: Design a composite reward function Including directional rewards , smoothness rewards and queue efficiency rewards, with the formula: Among them , is a hyperparameter used as a baseline for service rate selection; if the remaining capacity of the queue of node is more than that of its neighbor node , selecting a service rate lower than can relieve the pressure on downstream neighbors; ; used to evaluate the reasonableness of the port service rate, for measuring the upstream router 's service rate and the downstream router 's relationship between the changes in queue length; indicating that the smaller the difference in queue length brought by the action selected at the moment, the greater the reward; indicating that keeping the queue length small can also increase the reward; (3) Policy optimization training: Adopt a centralized training - distributed execution framework to optimize the policy network parameters by maximizing the cumulative reward function , and ensure the stability of policy updates by clipping the policy update ratio and the advantage function . The formula is as follows: ​ Among them is the clipping range. To ensure the stability of policy update, the clipping range is ; (4)Dynamic adjustment of service rate: During the distributed execution phase, each node independently adjusts the service rate of the output port of the node based on local observations , and the formula is as follows: , the formula is: Among them It is determined by the output of the policy network.

2. The method for congestion control of the space-air-ground network based on deep reinforcement learning according to claim 1, wherein Step 1 includes, First, model the space-air-ground network as a multi-agent Markov decision process (MDP), with each node regarded as an independent agent. Among them, set the satellite nodes to periodically call the method to simulate the link switching caused by orbital movement; distinguish the queue capacities and link bandwidths of ground nodes and satellite nodes, optimize resource allocation, and finally construct the global state space , including the normalized queue lengths of each node ; the normalized service rates of each link , where is the link capacity; at the same time, construct a local observation space for each node , which includes its own queue length , the historical traffic of neighbor nodes , providing an input basis for subsequent reward function design and policy optimization.

3. The congestion control method for the space-air-ground network based on deep reinforcement learning according to claim 2, wherein Step 2 includes, Based on the global state and local observations defined in Step 1, designing a composite reward function to guide the cooperation of agents to optimize congestion control; the composite reward function includes three parts: Directional Reward Adjust the service rate direction according to the difference in the adjacent node queues. A positive reward is obtained if the service rate is decreased when the upstream node queue is long; Smoothness Reward By evaluating the smoothness of the difference in the queues of adjacent nodes to encourage the minimization of the difference in the queue lengths of adjacent nodes and avoid local congestion; Queue efficiency reward It is to encourage the router to reduce queue backlog and lower transmission latency; The reward function drives the agent to relieve congestion and improve network service quality by weighted fusion of the above sub-rewards.

4. The congestion control method for the space-air-ground network based on deep reinforcement learning according to claim 3, characterized in that, Step 3 includes, Collect global state data during the centralized training phase, including queue length, service rate, and traffic fluctuations, and maximize the cumulative reward based on the input basis in Step 1 and the algorithm designed in Step 2. , update the ratio through the clipping strategy and the advantage function Ensure training stability, and introduce an experience replay pool to store historical data to improve the sample utilization rate; During the distributed execution phase, each node makes independent adjustments to the service rate based on local observations without the need for global communication, achieving decentralized real-time decision-making. ​ 5. The method for air-space-ground network congestion control based on deep reinforcement learning according to claim 4, characterized in that, Step 4 includes, The adjustment value output by the policy network trained based on Step 3 , dynamically updates the service rate of the node output port, and the service rate update formula is ; At the same time, through the queue dynamic model calculates the change in queue length in real time, and the queue change is jointly determined by the input traffic fluctuation and the service rate adjustment based on the policy network in Step 3; Aiming at the high dynamic characteristics of the satellite network, dynamically adapting to the satellite link capacity fluctuation through the MPLS traffic engineering protocol, periodically updating the topology to cope with orbital movement, and triggering a fast recovery mechanism to reallocate service rates and traffic paths when a link interruption is detected, optimizing resource allocation and load balancing, and enhancing the network robustness.

6. The method for air-space-ground network congestion control based on deep reinforcement learning according to claim 5, wherein Step 4 includes smoothing the downstream node pressure by buffering burst traffic, and the queue model satisfies: Among them is the sampling interval, and the change in queue length is jointly determined by traffic fluctuations and service rate adjustments Deploying a dynamic queue model in each node and describing the change of queue length by the following formula: wherein is the traffic splitting ratio, and are the arrival rate change of the input port and the service rate change of the output port respectively; the node adjusts the service rate of the output port in real time by using a deep reinforcement learning algorithm according to the current queue length and the queue difference of adjacent nodes to balance the traffic pressure between adjacent nodes.

7. The congestion control method for the space-air-ground network based on deep reinforcement learning according to claim 6, wherein Step 3 includes a deep reinforcement training process: The experience replay mechanism stores historical data ; among which the centralized Critic network is based on the global state and the joint action to calculate the Q value and guide the parameter update of the distributed Actor network. The loss function is as follows: wherein , is the discount factor, are the target policy network parameters; In the centralized training - distributed execution framework, the centralized Critic network evaluates the global state-action value function, and the distributed Actor network outputs the service rate adjustment strategy to achieve decentralized execution.

8. The method for congestion control of the space-air-ground network based on deep reinforcement learning according to claim 4, characterized in that In the distributed execution stage, each node only relies on local observations Adjust the service rate without global communication, and the action space is a continuous value .

9. The method for air-space-ground network congestion control based on deep reinforcement learning according to claim 3, wherein, The directional reward in the reward function measures whether the direction of service rate adjustment is correct. By using a reference value to judge the direction of service rate adjustment. If > and < , the reward is positive; otherwise, it is a penalty.

Citation Information

Patent Citations

  • Air-space-ground network congestion control method based on deep reinforcement learning

    CN119364423A

  • Leo satellite congestion control routing method

    US20240305364A1