Giant LEO constellation deterministic routing method based on deep reinforcement learning

By employing deep reinforcement learning methods and utilizing time-slotted network topology and a dynamic Q-network decision framework, the challenges of deterministic transmission in multi-layered mega-LEO constellations were addressed, enabling real-time response to time-sensitive services and strict quality of service assurance.

CN121887273APending Publication Date: 2026-04-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-01-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing satellite networks struggle to achieve end-to-end deterministic transmission in low-Earth orbit constellations, especially in multi-layered mega-LEO constellations, where they face challenges such as high computational complexity due to high dynamism and rapid topology changes, improper resource allocation, and insufficient adaptability to multi-service concurrent scenarios.

Method used

We employ a deep reinforcement learning-based approach to predict link lifetimes using a time-slotted network topology sequence. By combining spatial and temporal attention calculations to generate multi-scale spatiotemporal feature vectors, we construct a dynamic Q-network decision framework to optimize path selection and meet deterministic constraints.

Benefits of technology

It enables real-time response to time-sensitive services and strict quality of service assurance, reduces the impact of network dynamics on routing planning, and provides an efficient end-to-end deterministic routing method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887273A_ABST
    Figure CN121887273A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of satellite deterministic communication, and particularly relates to a giant LEO constellation deterministic routing method based on deep reinforcement learning, which comprises the following steps of: performing time slot modeling on a satellite network, generating a continuous time slot topological sequence, predicting path survival time, and defining deterministic service quality constraints of time-sensitive services; extracting multi-scale spatio-temporal features representing dynamic changes of the network; designing a hop-by-hop dynamic Q network routing decision framework based on the DQN, and selecting a next hop node for a data packet in real time; a high-performance candidate path set of each time slot is generated by using a dynamic Q network, and an end-to-end deterministic routing path with both transmission performance and continuity is finally dynamically output for the time-sensitive service in combination with a deterministic path selection mechanism across time slots. The method can effectively adapt to high dynamic characteristics of constellations, and provides strict time delay, jitter and packet loss rate guarantee for time-sensitive services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite deterministic communication, specifically relating to a deterministic routing method for giant LEO constellations based on deep reinforcement learning. Background Technology

[0002] With the rapid development of modern information warfare and military intelligence systems, time-sensitive operations such as intelligence reconnaissance, battlefield situational awareness, and real-time combat command place higher demands on the determinism of end-to-end data transmission. These operations typically have strict constraints on transmission latency and jitter, and their deterministic transmission requires data to arrive reliably within a specific time frame with an extremely low packet loss rate; otherwise, the normal operation of the entire system will be affected.

[0003] In recent years, Low Earth Orbit (LEO) satellite constellations have been widely used in aviation, marine, energy, and military fields due to their advantages such as wide coverage, low latency, large bandwidth, and flexible networking. Among them, emerging mega-constellations such as Starlink and Telesat have become important representatives of space information networks. Compared with single-layer small-scale networks, multi-layer mega-LEO satellite networks have more significant advantages in terms of coverage, network redundancy, and transmission latency, and are particularly suitable for meeting the stringent requirements of increasingly time-sensitive services for deterministic transmission. However, the high-speed relative motion between LEO satellites causes the inter-satellite link status to change drastically over time, resulting in highly dynamic network topology. Therefore, designing efficient and stable routing methods has become the core challenge for achieving end-to-end deterministic transmission.

[0004] Currently, research on deterministic transmission in satellite networks mainly focuses on end-to-end latency guarantees, network resource scheduling, and dynamic topology adaptation. Although existing work has improved transmission performance to some extent, it still has the following main limitations: First, most studies are based on relatively lenient quality of service constraints and fail to strictly address the hard indicators of latency, jitter, and packet loss rate for time-sensitive services; second, existing heuristic optimization algorithms often have high computational complexity, making it difficult to respond in real time when the topology changes rapidly; third, existing methods mostly rely on prior network information or static models, which have limited adaptability to highly dynamic link environments and diverse service requirements. In recent years, intelligent methods such as reinforcement learning have been introduced into satellite network routing planning to enhance its online decision-making and adaptive optimization capabilities. However, existing research still fails to systematically incorporate key indicators such as jitter and packet loss rate into the optimization objectives, and also lacks in-depth consideration of dynamic resource allocation in scenarios with multiple concurrent services.

[0005] Unlike traditional "best-effort" transmission modes, deterministic transmission emphasizes the controllability and predictability of the transmission process, requiring strictly defined delay, jitter, and reliability guarantees for time-sensitive services. In mega-LEO constellations, the instability of inter-satellite links and the time-varying nature of transmission paths are the main challenges facing inter-satellite deterministic routing planning. With the rapid expansion of network scale, increasingly complex topologies, and more frequent link changes, the multidimensional time-varying nature of node and path resources, as well as the link resource allocation problem under multi-service concurrency, collectively constitute the key bottlenecks for satellite networks to provide deterministic services. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a deterministic routing method for giant LEO constellations based on deep reinforcement learning, comprising:

[0007] S1: Generate a time-slotted network topology sequence based on satellite orbit parameters, predict the link lifetime of the end-to-end path, and obtain the transmission requests of time-sensitive services and their deterministic constraints on end-to-end latency, jitter, and packet loss rate;

[0008] S2: For the current and historical time slot topology, perform spatial attention calculation and temporal attention calculation to obtain spatial correlation weight and temporal evolution weight respectively, and perform feature fusion to obtain multi-scale spatiotemporal feature vectors;

[0009] The spatial attention calculation involves: for each time slot's topology, introducing prior information reflecting the global network topology, and combining node characteristics to calculate the spatial association weights between nodes.

[0010] The time attention calculation involves calculating the time evolution weight of the state features of each network node across different time slots.

[0011] The feature fusion: the spatial correlation weights and temporal evolution weights are fused to generate a multi-scale spatiotemporal feature vector representing the network dynamics;

[0012] S3: Construct a dynamic Q-network decision framework based on deep Q-network. In each decision time slot, construct the state with the spatiotemporal feature vector, network state and service constraints. Select the next hop node for the data packet and perform forwarding through the framework. Calculate the reward based on the service performance after the action is executed. Optimize network parameters by using experience replay and time difference learning.

[0013] S4: In each time slot topology, the feasible path from the source node to the destination node is evaluated using the dynamic Q network, and several paths are selected to form a candidate path set based on the cumulative expected reward of the path.

[0014] S5: Using the transmission path of the previous time slot as a reference, select the path with the highest link repetition rate with the reference path from the candidate path set as the forwarding path for the current time slot; if there are multiple paths with the same repetition rate, select the path with the highest cumulative historical transmission performance reward value; finally, output the selected path as the end-to-end routing path that satisfies the deterministic constraint.

[0015] The beneficial effects of this invention are:

[0016] This invention addresses the highly dynamic characteristics of giant LEO constellations. First, it constructs a basic framework for network evolution through time-slotted discrete modeling. Then, it proposes a dynamic spatiotemporal network model that adaptively captures the evolutionary patterns of topology and node characteristics, effectively reducing the impact of network dynamics on routing planning. Finally, by integrating multi-dimensional deterministic indicators such as latency, jitter, and packet loss rate, it constructs a dynamic Q-network decision framework based on deep reinforcement learning. Combined with a cross-time-slot deterministic path selection mechanism, this forms an end-to-end deterministic routing method capable of responding to network changes in real time and providing strict quality-of-service guarantees for time-sensitive services. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of a giant LEO constellation provided according to an embodiment of the present invention;

[0018] Figure 2 This invention provides a framework for a deterministic routing method based on dynamic Q networks.

[0019] Figure 3 This is a schematic diagram of the dynamic spatiotemporal network model of the present invention;

[0020] Figure 4 This is a schematic diagram of the time-independent attention block design of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Build as Figure 1 The giant LEO satellite network shown consists of Very Low Earth Orbit (VLEO), Medium Low Earth Orbit (MLEO), and High Low Earth Orbit (HLEO). Based on... Figure 1The network scenario shown illustrates a deterministic routing method for giant LEO constellations based on deep reinforcement learning, such as... Figure 2 As shown, this is achieved through the following sub-steps:

[0023] S1: Generate a time-slotted network topology sequence based on satellite orbit parameters, predict the link lifetime of the end-to-end path, and obtain the transmission requests of time-sensitive services and their deterministic constraints on end-to-end latency, jitter, and packet loss rate;

[0024] S2: For the current and historical time slot topology, perform spatial attention calculation and temporal attention calculation to obtain spatial correlation weight and temporal evolution weight respectively, and perform feature fusion to obtain multi-scale spatiotemporal feature vectors;

[0025] The spatial attention calculation involves: for each time slot's topology, introducing prior information reflecting the global network topology, and combining node characteristics to calculate the spatial association weights between nodes.

[0026] The time attention calculation involves calculating the time evolution weight of the state features of each network node across different time slots.

[0027] The feature fusion: the spatial correlation weights and temporal evolution weights are fused to generate a multi-scale spatiotemporal feature vector representing the network dynamics;

[0028] S3: Construct a dynamic Q-network decision framework based on deep Q-network. In each decision time slot, construct the state with the spatiotemporal feature vector, network state and service constraints. Select the next hop node for the data packet and perform forwarding through the framework. Calculate the reward based on the service performance after the action is executed. Optimize network parameters by using experience replay and time difference learning.

[0029] S4: In each time slot topology, the feasible path from the source node to the destination node is evaluated using the dynamic Q network, and several paths are selected to form a candidate path set based on the cumulative expected reward of the path.

[0030] S5: Using the transmission path of the previous time slot as a reference, select the path with the highest link repetition rate with the reference path from the candidate path set as the forwarding path for the current time slot; if there are multiple paths with the same repetition rate, select the path with the highest cumulative historical transmission performance reward value; finally, output the selected path as the end-to-end routing path that satisfies the deterministic constraint.

[0031] S11: The purpose of this step is to transform the continuous dynamic motion of the satellite network into a series of discrete, static topological snapshots, providing a structured time frame for subsequent analysis. This is achieved through the following sub-steps:

[0032] Input the set of orbital parameters for the target giant LEO satellite constellation. This set contains the orbital six elements for each satellite (or each standard orbital plane), including: semi-major axis a, eccentricity e, orbital inclination i, and right ascension of the ascending node. The perigee argument ω, and the mean perigee angle M0 or the true perigee angle θ.

[0033] Given a system time base and a time slot length ∆t, for each discrete time slot start time t k (k=0,1,2,…), based on the orbital parameters obtained above, the three-dimensional position vector of each satellite in the geocentric inertial coordinate system at that moment is calculated using orbital mechanics formulas. A specific calculation formula is as follows:

[0034]

[0035] Where h is the satellite's angular momentum and μ is the Earth's gravitational constant.

[0036] For time slot t k Based on the position vectors of all satellites Determine the v of two satellites i and v j Does an available inter-satellite link exist between them? The following conditions must be met simultaneously: 1) The Euclidean distance R between them... ij The maximum communication distance D of the inter-satellite link max 2) The line connecting the two is not obstructed by the Earth, meaning the angle subtended by the Earth's center to the line is less than the maximum visible elevation angle θ. LOS All links that meet the conditions ij This is recorded, forming a static topology snapshot G for that time slot. k =(V,E k ), where V is the set of satellite nodes, E k This is the set of links that are valid for this time slot.

[0037] Arrange all time-slot topology snapshots {G0, G1, ..., G} in chronological order. m-1 This yields a time-slotted network topology sequence that describes the continuous dynamic evolution of the network.

[0038] S12: The purpose of this step is to establish a series of key quantitative indicators for evaluating network dynamics and path stability. This is achieved through the following sub-steps:

[0039] For time slot t k Any valid inter-satellite link Its link lifetime Defined from time t kThe maximum estimated duration for which the link will remain connected (i.e., continuously meet the distance and unobstructed conditions) is calculated as follows: First, obtain satellite v i and v j In t k Position vector at time , With velocity vector , Secondly, calculate the relative position vector between the two. With relative velocity vector Finally, the maximum communication limit was reached by solving for the relative motion of the satellite. time And the time of link interruption due to Earth's obstruction. Ultimately and The minimum value is used as The approximate value, that is:

[0040]

[0041] in, For the equation The minimum positive solution; For the equation The minimum positive solution, For two satellites The angle subtended by time relative to the Earth's center. This is the maximum visible geocentric angle.

[0042] For an end-to-end communication path P={l1,l2,…,l...} consisting of several links in sequence, n Define the path lifetime. The minimum lifetime of all links on this path:

[0043]

[0044] To quantify the dynamic stability of a link, link l is defined. ij ranging rate This is the rate of change of the distance between two satellites, used to predict the jitter trend of propagation delay. The calculation method is as follows:

[0045]

[0046] in, It is the dot product of the relative position and the relative velocity vector. This indicates that the two stars are moving away from each other, and the distance between them is increasing; This indicates that the distance is decreasing as the link approaches. This value is recorded as link l. ij A dynamic feature attribute.

[0047] S13: The purpose of this step is to receive external communication tasks and extract the necessary service features and constraints for routing decisions in a structured manner. This is achieved through the following sub-steps:

[0048] The system first receives time-sensitive service f k The transmission request, and extract the source node n from it. src,k and the destination node n dst,k Bandwidth required for the service (c) k The size of the service data packets sent per cycle is pf k Sending period c p Business duration [t] start , t end Simultaneously, extract the threshold ζ with end-to-end delay. k jitter threshold j k With packet loss rate threshold pl k The threshold is defined as a deterministic boundary constraint that must be satisfied in the routing decision.

[0049] To ensure the accuracy of routing decisions, formal formulas for calculating key performance indicators are defined. These formulas will be used in subsequent steps for accurate evaluation and reward calculation.

[0050] Business f k End-to-end total delay along the path The sum of the delays experienced by its data packets, including: 1) Transmission delay: Where R is the satellite transmission rate, f k 1) Size of a single data packet; 2) Propagation delay: , where v c At the speed of light, For link l ij Length; 3) Processing delay: , where q i For node v i Maximum cache capacity 4) Queuing delay: (This refers to the data packet processing rate;) ,in The time slot length, For node v i The number of data packets at the beginning of time slot t.

[0051] Therefore, each time-sensitive task f k The formula for calculating the total end-to-end delay is: .in, for f k The sum of all data packets, f kFrom source node n src,k to n dst,k The set of links that need to be traversed.

[0052] Business f k shaking End-to-end latency for all its data packets The fluctuation range is calculated using the following formula: .

[0053] Business f k packet loss rate To start from the total number of data packets N sent send N of the total number of successfully received data packets receive The percentage deviation is calculated using the following formula: .

[0054] Furthermore, the purpose of step S2 is to extract features characterizing the dynamic evolution of the network from the continuous time-slot topology. For example... Figure 3 The dynamic spatiotemporal network model shown achieves accurate representation of dynamic networks, specifically through the following three sub-steps:

[0055] S21: The purpose of this step is to extract structural features from the network topology of a single time slot that reflect the global spatial location and correlation of nodes. This is achieved through the following sub-steps:

[0056] For the current time slot topology G i Each node v in i Calculate its proximity centrality cc i This is used to quantify the global reachability of a node in the network. Specifically, it calculates the reciprocal of the average distance to all its neighboring nodes, using the following formula:

[0057]

[0058] Wherein, N(v) i ) represents node v i The set of neighboring nodes, R ij This represents the distance between the two nodes.

[0059] The original features of each node (May include information such as position, velocity, and load) and the calculated proximity centrality code cc i These are combined to form the initial feature representation of the node.

[0060] The initial feature representations of all nodes are input into a self-attention block consisting of L stacked graph attention layers (GL). Each graph attention layer (GL) performs the following operations sequentially: a) Layer normalization (LN) on the input features; b) Input the normalized features into a multi-head self-attention layer (MHA) to calculate the global spatial association weights between nodes; c) Residually connect the output of the MHA layer to its input; d) Perform layer normalization again; e) Perform a nonlinear transformation through a feedforward neural network (FFN); f) Residually connect the output of the FFN to its input to obtain the output of that layer. After processing by L layers, the output of each node in the time slot topology G is obtained. i Structural features .

[0061] For time-slot topology G i Each node in the algorithm is assigned a learnable position encoding vector p. i The obtained structural features With the corresponding position code p i Add them together to inject the absolute or relative position information of the nodes.

[0062] A linear transformation is performed on the features that incorporate positional encoding to generate the final spatial structure feature vector for each node. The transformation formula is:

[0063]

[0064] Among them, W e b is a learnable linear transformation weight matrix. e This is the bias term. This vector This is a feature representation that captures the global spatial correlation of nodes.

[0065] S22: The purpose of this step is to extract regular features reflecting the evolution of a node's state over time from a feature sequence spanning multiple consecutive time slots. A schematic diagram of the time-dependent attention block design is shown below. Figure 4 As shown, this is achieved through the following sub-steps:

[0066] For the current node v to be processed i The spatial structure feature vectors generated in step S21 in the most recent T consecutive time slots are obtained and arranged in chronological order to form the temporal input sequence of the node.

[0067] Each feature vector in the aforementioned time series is mapped in parallel to its corresponding query vector Q using multiple sets of learnable linear transformation matrices. i Key vector K i Sum vector Vi For the i-th attention head, its mapping process is implemented using the following formula:

[0068]

[0069] in, , , These are the query, key, and value projection weight matrices corresponding to the i-th attention head.

[0070] For the i-th attention head, its query vector Q is calculated. i With K i The dot product of the transpose of the expression yields the unnormalized attention score. This score is then scaled (divided by the square root of the key vector dimension). Then, the Softmax function is applied to obtain the normalized temporal attention weight matrix a. i The calculation process is as follows:

[0071]

[0072] The weight matrix a i It represents the importance of node characteristics at different historical moments for the prediction of the current moment.

[0073] Using the obtained time attention weight matrix a i For the corresponding value vector sequence V i Perform weighted summation to generate the output feature h of the i-th attention head. i,head :

[0074]

[0075] Output features of all N attention heads The concatenation is performed along the feature dimension using a learnable projection weight matrix W. O Perform a linear transformation on the concatenated features to obtain node v i Final time evolution feature vector :

[0076]

[0077] S23: The purpose of this step is to effectively fuse the spatial structure features and temporal evolution features extracted from the nodes to generate a unified and information-rich dynamic spatiotemporal feature representation of the nodes for use by the subsequent decision-making module. This is achieved through the following sub-steps:

[0078] node v i Spatial structural feature vector Evolutionary eigenvectors over time By adding elements one by one, we can obtain the initial characteristics of fusion. .

[0079] Features after fusion Layer normalization is performed, followed by nonlinear transformation through a feed-forward network to further enhance its feature representation capabilities.

[0080] The output of the feedforward neural network is node v i The final multi-scale spatiotemporal feature vector at the current moment The set of feature vectors of all nodes. This will serve as the state input for the subsequent reinforcement learning agent to perceive the network environment.

[0081] Furthermore, the purpose of step S3 is to make real-time decisions about the next-hop node for each time-sensitive service based on a deep reinforcement learning agent in each decision time slot, aiming to optimize long-term network performance while meeting deterministic transmission requirements. Specifically, the following operations are performed on the topology in each decision time slot:

[0082] S31: The purpose of this step is to construct a multi-dimensional vector that can comprehensively represent the current network environment and service status, serving as the perceptual basis for the reinforcement learning agent to make routing decisions. This is achieved through the following sub-steps:

[0083] The aggregation is done by step S23 for each node v in the network i Generated spatiotemporal feature vectors This forms a set of node-level feature representations.

[0084] Obtain and incorporate parameters describing the overall network size, including the total number of active time-sensitive services K, the total number of satellite nodes N, and the v of each node. i Number of neighbor links M i .

[0085] Collect and integrate real-time changing dynamic information, specifically including: a) the position vector of each node. and velocity vector b) Track height H of each node i c) The vector formed by the remaining bandwidth of each output link of each node. d) Remaining cache capacity of each node e) All time-sensitive services f k Real-time status information (such as source node, destination node, remaining data volume, etc.).

[0086] The information from the above three aspects (node ​​spatiotemporal features, global parameters, and real-time dynamic information) is concatenated and fused to construct a unified multidimensional state vector. ,in , , , , , , , which serves as the input to the deep reinforcement learning agent in decision time slot t.

[0087] S32: The purpose of this step is to define a set of optional forwarding actions for each time-sensitive service, and to actually execute the next-hop forwarding of data packets based on the agent's decision output and exploration strategy. This is specifically achieved through the following sub-steps:

[0088] For each time-sensitive service f to be routed k When a data packet is located at a certain satellite node, the set of all communicable neighbor nodes of that node is defined as the discrete action space A of the service at the current moment. k The task of the intelligent agent is to obtain information from A. k Choose a node as the next hop.

[0089] The state vector S constructed in step S31 t Input a Deep Q-Network (DQN). The network outputs a corresponding Q-value for each possible action in the current task (i.e., for each neighboring node). A final decision is made using an ε-greedy strategy: actions are selected from the action space A with probability ε. k Randomly select an action a t (Exploration), select the action with the highest Q value with probability 1-ε (Exploitation).

[0090] Based on the action selected by the agent (i.e., the specified next-hop node), the control system will perform time-sensitive services. The current data packet is forwarded from the current node to the next hop node.

[0091] S33: The purpose of this step is to define a comprehensive, multi-dimensional reward function based on the network feedback after the selected action. This function calculates the immediate evaluation score for the selected next-hop action, guiding the agent to learn a routing strategy that meets deterministic requirements. This is achieved through the following sub-steps:

[0092] Based on the ranging rate of the selected next-hop link Define dynamic quality rewards for links The reward value and Negative correlation, at the same time Under the same circumstances, for Additional rewards are given to negative states to reduce transmission latency, thereby encouraging the selection of links that can reduce propagation latency jitter.

[0093] Based on the lifetime of the selected link ,definition The reward value and Positive correlation, to encourage the selection of links with longer lifespans and greater stability.

[0094] Check the remaining cache of the next hop node. Is it not less than the size pf of the data packet to be forwarded? k If so, then define the cache satisfaction reward C. space =1; otherwise C space =-1.

[0095] Check the remaining bandwidth of the selected link. Is it not less than the bandwidth c required by the business? k If so, then define a bandwidth satisfaction reward l. band =1, otherwise l band =-1.

[0096] Define jump reward Its value It is negatively correlated with the current accumulated path hop count m to encourage the selection of paths with fewer hops.

[0097] If the orbital height of the next hop node is different from that of the current node (H) i ≠H j If the off-track reward is defined, then the off-track reward is defined. ;otherwise .

[0098] Estimate the end-to-end latency (including transmission, propagation, processing, and queuing latency) of the service along the current path and compare it with the latency threshold ζ. k Comparison. Define latency compliance rewards. Its value is higher when the estimated delay is close to but does not exceed the threshold, and decreases significantly when it exceeds the threshold.

[0099] Monitor the cumulative latency, jitter, and packet loss rate of the monitoring service. If the cumulative latency exceeds the threshold ζ... k Then a penalty item will be generated. If the cumulative jitter exceeds the threshold j k Then a penalty item will be generated. If the packet loss rate exceeds the threshold pl k Then a penalty item will be generated. .

[0100] If the time-sensitive business successfully reaches the destination node while satisfying all deterministic constraints, a substantial positive completion reward will be generated. ;otherwise .

[0101] The sub-reward items and penalty items defined above are weighted and summed according to preset weights to calculate the execution action a. t The final total instant reward value r t The calculation formula is:

[0102]

[0103] Wherein, β1 to β7 are the weights of each sub-reward item, and α1 to α4 are the weights of the penalty item and the completion reward.

[0104] S34: The purpose of this step is to continuously optimize the agent's policy network parameters using historical decision data through experience replay and Deep Q-Network (DQN) training mechanisms, in order to learn a deterministic routing policy that maximizes long-term cumulative rewards. This is achieved through the following sub-steps:

[0105] The quadruple experience (S) generated at each decision step t ,a t ,r t ,S t+1 Stored in a fixed-size experience playback buffer.

[0106] Periodically sample a small batch (Minibatch) of historical experience data randomly from the buffer.

[0107] Using the sampled data, the error (loss) between the current Q-network's predicted value and the target value is calculated using the temporal difference (TD) algorithm. This loss is then minimized using gradient descent (e.g., the Adam optimizer), thereby updating the parameters of the deep Q-network. This process aims to teach the agent to maximize future cumulative discount rewards. The strategy is as follows, where γ is the decay factor and K is the cumulative reward calculation duration parameter.

[0108] Furthermore, the purpose of step S4 is to integrate pre-network modeling, feature extraction, and reinforcement learning decision-making to form a complete and automated routing method process, dynamically generating and outputting end-to-end candidate routing paths for each time slot within the duration of each time-sensitive service. This method is based on a dynamic Q-network framework, and its specific implementation is achieved through the following sub-steps:

[0109] S41: The purpose of this step is to initialize all the core parameters, network model, and data structures required by the routing method agent, laying the foundation for subsequent iterative training and real-time routing decisions. This is achieved through the following sub-steps:

[0110] Randomly initialize the main Q-network parameters θ and initialize the target Q-network parameters θ.target And let θ target =0.

[0111] Initialize the experience playback memory unit D, which is used to store historical experience data of the agent's interaction with the environment.

[0112] Define the key hyperparameters for reinforcement learning: learning rate α, discount factor γ, and exploration rate ε.

[0113] Set the update frequency C of the main network (i.e., synchronize the target network parameters once every C training steps) and the size of the minibatch of data randomly sampled from the experience replay buffer during each training session.

[0114] S42: The purpose of this step is to initiate an independent path planning and agent training process for each time-sensitive service, ensuring that the method can learn and generate routing strategies that are customized for each task and adapt to its entire lifecycle. This is achieved through the following sub-steps:

[0115] For the set F of communication tasks that need to be planned k Each time-sensitive task f in k Initiate a separate training and planning episode. For the current task f k First, reset the environment state to the initial state S0, and then set the instant reward r for the current round. t Reset to zero. Status S0 should include the satellite network topology, link status, and service requirements at the start of the mission.

[0116] S43: The purpose of this step is to systematically process the network topology of each discrete time slot that a time-sensitive service traverses within its complete transmission cycle, in chronological order, and to generate real-time routing decisions for each time slot. This is achieved through the following sub-steps:

[0117] According to the time-sensitive service f k start time t start and end time t end To determine the continuous time slot sequence covered by its transmission process.

[0118] Following the chronological order, iterate through each time slot topology G experienced by this service. t The subsequent steps (S44 to S48) are executed sequentially to complete the state perception, agent decision-making, path determination and other operations within the time slot.

[0119] Complete the current time slot G t After all processing is complete, the circular index will be pointed to the next time slot topology G. t+1 Repeat step 3 until all time slots covered by the service have been processed.

[0120] S44: The purpose of this step is to perceive the network environment at the start of the time slot and integrate multi-source information to construct a structured state vector that can comprehensively represent the current network and service status, serving as the input for the agent's routing decision. This is achieved through the following sub-steps:

[0121] In time slot G t At the start time, obtain the current topology snapshot generated by step S1, and read the real-time dynamic parameters of all satellite nodes in the network, including position, velocity, orbital altitude, remaining bandwidth of each link, and remaining cache capacity of the nodes.

[0122] Call the processing result of step S23 to obtain the result for each node v. i Generated spatiotemporal feature vectors .

[0123] Incorporate time-sensitive business processes currently awaiting decision-making. k The real-time status (such as source node, destination node, remaining data volume) and the deterministic quality of service constraints (latency, jitter, packet loss rate threshold) extracted from step S13.

[0124] The acquired real-time network information, node spatiotemporal characteristics, service status, and constraints are concatenated and encoded according to a preset format to construct a unified multidimensional state vector S. t The vector S t This is the complete representation of the environmental state upon which the reinforcement learning agent makes routing decisions at the current moment.

[0125] S45: The purpose of this step is to enable the agent to balance exploring unknown actions with utilizing existing experience based on the current environmental state, and ultimately decide on the next hop forwarding node. This is achieved through the following sub-steps:

[0126] The state vector S constructed in step S44 t The input is fed into the deep Q-network (main network) during training.

[0127] The Q network processes the input state, providing each candidate neighbor node of the node containing the current data packet (i.e., each possible action a). t Calculate and output a corresponding Q value Q(S) t ,a;θ), forming an action value vector.

[0128] Randomly select one of the all valid neighbor nodes as the next hop action a with a preset probability ε. t .

[0129] Choose the neighbor node with the highest Q value as the next hop action a with probability 1-ε. t ,Right now .

[0130] The neighbor node identifier selected according to the above strategy will be used as the final routing action a. t Output and pass it to the execution module.

[0131] S46: The purpose of this step is to implement the routing action decided by the agent, calculate the immediate reward based on the network feedback after the action is executed, and drive the environmental state update to prepare for the next round of decision-making. This is achieved through the following sub-steps:

[0132] Based on the routing action a determined in step S45 t (i.e., the specified next-hop node), the control system will handle time-sensitive services f k The current data packet is forwarded from the current satellite node to the next-hop node.

[0133] After the action is executed, based on the new network state (such as updated link load, node queue length, etc.), the multi-dimensional reward function defined in step S33 is invoked to calculate the instantaneous reward value r obtained from executing the action. t This reward value quantifies the immediate effect of the selected action on meeting deterministic transmission requirements and optimizing transmission performance.

[0134] The environment (or simulator) drives the entire system state from the current state S based on the actions already performed and the dynamic model of the satellite network (such as orbital motion and flow changes). t Transition to state S at the next decision moment t+1 The new state S t+1 This will serve as input for the next round of decision-making.

[0135] S47: The purpose of this step is to store the agent's experience of interacting with the environment and to use this historical data to periodically train the deep Q network to optimize its parameters, thereby improving the decision-making performance of the routing strategy.

[0136] Encapsulate a complete interaction experience generated in step S46 into an experience tuple e. t =(S t ,a t ,r t ,S t+1 ), and store it in a fixed-capacity experience playback buffer D.

[0137] Once the amount of data in the experience replay buffer reaches a preset threshold, a small batch of historical experience data is randomly and evenly sampled from it for use in this round of network training.

[0138] Based on the sampled mini-batch data, the loss function L(θ) of the current master Q network is calculated. The loss function adopts the form of time difference error, and the formula is as follows:

[0139]

[0140] Where N is the batch size, γ is the discount factor, and Q(S) t ,a;θ) are the predicted values ​​of the main network. It is calculated by the target network and serves as a stable estimate of future returns.

[0141] Using a gradient descent algorithm (such as the Adam optimizer), the gradient is calculated based on the loss function L(θ), and the parameters θ of the main Q network are updated along the negative gradient direction.

[0142] After every C updates of the main network parameters (where C is the preset update frequency), the parameters θ of the main network are copied to the target network, i.e., θ is set to... target =θ, to stabilize the training process.

[0143] S48: The purpose of this step is to integrate the policies and path smoothing mechanisms learned by the agent to determine a final forwarding path for the current time slot that satisfies both optimal instantaneous performance and cross-time slot continuity. This is achieved through the following sub-steps:

[0144] Using the currently trained deep Q-network (main network), evaluate the performance in time slot G. t Find all feasible paths from the source node to the destination node. Calculate or estimate the cumulative expected reward for each path, and select the top N paths with the highest reward values ​​to form the candidate path set {p1, p2, ..., p...} for the current time slot. N}

[0145] Furthermore, the purpose of step S5 is to select a new path for time-sensitive services that can maximize transmission continuity and stability when the network topology is updated, so as to reduce performance fluctuations or packet loss caused by path switching. This is specifically achieved through the following sub-steps:

[0146] When the system detects that the network has entered a new time slot topology G t This triggers the deterministic path selection mechanism. At this point, the input information includes: a) the transmission path p used in the previous topology slot. prev and its link set L prev b) In the current new topology G t In step S4, the N candidate paths obtained are {p1, p2, ..., p...} N}, and their corresponding link sets are L1, L2, ..., L N The reward function is Reward(p).

[0147] For each candidate path p i (i=1,2,…,N), calculate its relationship with the previous topological path p. prev Link repetition rate The repetition rate is calculated using the Jaccard similarity coefficient, as shown in the following formula:

[0148]

[0149] in, This indicates the number of duplicate links between two paths. This represents the total number of links after deduplication.

[0150] Compare the link duplication rates of all candidate paths and select the subset p of paths with the highest link duplication rate. max :

[0151]

[0152] If |p max If |=1, meaning there is only one path with the highest repetition rate, then that path p i The forwarding path directly determined as the current time slot If |p max If |>1, meaning there are multiple candidate paths with the same repetition rate, a secondary selection needs to be made based on their transmission performance reward. The reward function is Reward(p i Defined by a deep reinforcement learning framework (as described in step S3), its value comprehensively reflects the path's performance on deterministic metrics such as latency, jitter, and packet loss rate. The path with the highest reward value is selected as the optimal path.

[0153]

[0154] in, For the Kth step in the future on path p i The instant reward obtained above, where γ is the discount factor. The step size for calculating cumulative rewards.

[0155] The final determined path As the current time slot topology G t The new deterministic forwarding path for time-sensitive services outputs a deterministic routing path within each time slot based on the transmission duration of the time-sensitive task, and delivers it to the routing execution module for packet forwarding.

[0156] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A deep reinforcement learning based deterministic routing method for mega LEO constellation, characterized in that, include: S1: Generate a time-slotted network topology sequence based on satellite orbit parameters, predict the link lifetime of the end-to-end path, and obtain the transmission requests of time-sensitive services and their deterministic constraints on end-to-end latency, jitter, and packet loss rate; S2: For the current and historical time slot topology, perform spatial attention calculation and temporal attention calculation to obtain spatial correlation weight and temporal evolution weight respectively, and perform feature fusion to obtain multi-scale spatiotemporal feature vectors; The spatial attention calculation involves: for each time slot's topology, introducing prior information reflecting the global network topology, and combining node characteristics to calculate the spatial association weights between nodes. The time attention calculation involves calculating the time evolution weight of the state features of each network node across different time slots. The feature fusion: the spatial correlation weights and temporal evolution weights are fused to generate a multi-scale spatiotemporal feature vector representing the network dynamics; S3: Construct a dynamic Q-network decision framework based on deep Q-network. In each decision time slot, construct the state with the spatiotemporal feature vector, network state and service constraints. Select the next hop node for the data packet and perform forwarding through the framework. Calculate the reward based on the service performance after the action is executed. Optimize network parameters by using experience replay and time difference learning. S4: In each time slot topology, the feasible path from the source node to the destination node is evaluated using the dynamic Q network, and several paths are selected to form a candidate path set based on the cumulative expected reward of the path. S5: Using the transmission path of the previous time slot as a reference, select the path with the highest link repetition rate with the reference path from the candidate path set as the forwarding path for the current time slot; if there are multiple paths with the same repetition rate, select the path with the highest cumulative historical transmission performance reward value; finally, output the selected path as the end-to-end routing path that satisfies the deterministic constraint.

2. The deterministic routing method for mega-LEO constellation based on deep reinforcement learning according to claim 1, characterized in that, Based on satellite orbital parameters, a time-slotted network topology sequence is generated, and the link lifetime of the end-to-end path is predicted, including: Obtain satellite and exist Position vector at time , With velocity vector , Calculate the relative position vector between the two. With relative velocity vector And solve for the satellite's relative motion causing its distance to reach the maximum communication limit. time And the time of link interruption due to Earth's obstruction Ultimately and The minimum value is used as link l ij Survival time The approximate value, that is: ,in, For the equation The minimum positive solution, For the equation The minimum positive solution, For two satellites The angle subtended by time relative to the Earth's center. This is the maximum visible geocentric angle; For paths whose path lifetime is determined by: .

3. The deterministic routing method for mega-LEO constellation based on deep reinforcement learning according to claim 1, characterized in that, Obtain the deterministic constraints on the transmission requests of time-sensitive services and their end-to-end latency, jitter, and packet loss rate, including: Business f k End-to-end total delay along the path The sum of all delays experienced by its data packets, including: transmission delay. Propagation delay Processing delay Queuing delay Therefore, business f k The formula for calculating the total end-to-end delay is: Where R is the satellite transmission rate, For business f k The size of a single data packet; v c For the speed of light, R ij For link l ij Length; q i For node v i Maximum cache capacity, Dp rate This refers to the data packet processing rate. The time slot length, For node v i The number of data packets at the beginning of time slot t; For business f k The sum of all data packets, Indicates business f k From source node n src,k to destination node n dst,k The set of links that need to be traversed; Business f k shaking End-to-end latency for all its data packets fluctuation range ; Business f k Packet loss rate P sd To start from the total number of data packets N sent send N of the total number of successfully received data packets receive percentage of deviation .

4. The deterministic routing method for mega-LEO constellation based on deep reinforcement learning according to claim 1, characterized in that, The spatial attention calculation includes: For each node v i in the current time slot topology G i , compute its closeness centrality encoding cc i : where N(v i ) denotes the set of neighbor nodes of node v i , and R ij is the distance between two nodes. The original features of the nodes are combined with the proximity centrality encoding cc i to form the initial feature representation of the nodes; inputting the initial feature representation into a structural self-attention module to output structural features of each node in a time slot topology G i ;​ The module contains multiple sequentially connected graph attention layers. Each graph attention layer processes all node features through a multi-head self-attention mechanism to capture global spatial correlations. A learnable position encoding vector is assigned to each node in a slot topology G i The structural features output by the structural self-attention module are added to the learnable position encoding of the node After linear transformation, the spatial structure feature vector of the node is generated ​​ where W e is a learnable linear transformation weight matrix, b e is a bias term.

5. The deep reinforcement learning based deterministic routing method for mega-LEO constellation according to claim 1, wherein, The time attention calculation includes: For node v i , its spatial structure feature vector in the last T consecutive time slots is obtained to form a time sequence input sequence , where is the spatial structure feature vector of time slot t N . Each feature vector in the sequence of timing inputs is mapped to a corresponding query vector Q by a set of learnable linear transformation matrices i , key vector K i , and value vector V i : wherein, , , are the query, key, value projection weight matrices corresponding to the i-th attention head, respectively. Based on the query vector Q i and bond vector K i Calculate the time attention weight matrix and for the value vector Perform weighted summation to generate preliminary time evolution features. ; where d k Let T be the dimension of the key vector, and T be the transpose operation. concatenating and linearly transforming the outputs of all the attention heads to generate a time evolution feature vector of the node : where W O is a learnable projection weight matrix, is the time-evolving feature vector of the nth node.

6. The deep reinforcement learning based deterministic routing method for mega-LEO constellation according to claim 1, wherein, The feature fusion includes: The spatial structure feature vector With the time evolution feature vector By performing element-by-element addition, preliminary fusion characteristics are obtained. After performing layer normalization on the preliminary fused features, a feedforward neural network is used for nonlinear transformation to generate the final node multi-scale spatiotemporal feature vector. .

7. The deep reinforcement learning based deterministic routing method for mega-LEO constellation according to claim 1, wherein, Rewards are calculated based on business performance after the action is executed, including: Based on the ranging rate of the selected next-hop link Set up dynamic quality rewards for links To encourage the selection of links that reduce propagation delay jitter, where the ranging rate... The rate of change of the distance between the two satellites is used to predict the jitter trend of propagation delay. The dot product of relative position and relative velocity vectors. The distance between the two nodes; According to the survival time of the selected link , define The reward value With Positive correlation, to encourage the selection of links with longer and more stable survival time; checking the remaining bandwidth of the selected link whether not less than the bandwidth c required by the service k if yes, then defining a bandwidth satisfaction reward l band = 1, otherwise l band = -1; Defining hop rewards whose value is negatively related to the currently accumulated path hop count m, to encourage the selection of paths with fewer hops; If the next hop node is at a different altitude than the current node , define the off-track reward ; otherwise ; Estimate the end-to-end latency of the service along the current path and compare it with the latency threshold ζ. k Compare and define latency compliance rewards Its value is higher when the estimated delay is close to but does not exceed the threshold, and decreases significantly when it exceeds the threshold; among them, For transmission delay, To delay the transmission time, To delay queuing, To handle latency; The cumulative latency, jitter, and packet loss rate of the monitoring service are monitored. If the cumulative latency exceeds the threshold ζ, the monitoring results will be recorded. k Then a penalty item will be generated. If the cumulative jitter exceeds the threshold j k Then a penalty item will be generated. If the packet loss rate exceeds the threshold pl k Then a penalty item will be generated. ; If the time-sensitive business successfully reaches the destination node while satisfying all deterministic constraints, a substantial positive completion reward will be generated. ;otherwise ; The sub-reward item defined above is weighted and summed with the penalty item according to a preset weight to calculate a comprehensive immediate reward value r after performing the action a t t :​ Wherein, β1 to β7 are the weights corresponding to the sub-reward items, and α1 to α4 are the weights corresponding to the penalty items and the completion reward, respectively.

8. The deep reinforcement learning based deterministic routing method for mega-LEO constellation according to claim 1, wherein, Several paths are selected from each time slot topology, including: selecting a next hop action a randomly from all legal neighbor nodes with a preset probability ε t ; selecting with probability 1 - epsilon the neighbor node with the highest Q value as the next hop action ; After the action is performed, an immediate reward r is calculated for performing the action in accordance with the new network state t ; The top N paths with the highest cumulative expected reward values ​​are selected to form a set of feasible candidate paths {p1, p2, ..., p} for the current time slot topology. N }, where p N Let N be the Nth feasible candidate path, where N is an integer greater than 1, and let R be the cumulative expected reward value. t The calculation formula is: wherein γ is a decay factor, K is a cumulative computation reward duration parameter, to perform an action the integrated immediate reward value after 9. A deterministic routing method for giant LEO constellations based on deep reinforcement learning according to claim 1, characterized in that, The link repetition rate includes: in, For two paths p prev With p i Link duplication rate between links This indicates the number of duplicate links between two paths. p represents the total number of links after deduplication. prev p is the transmission path used by the previous topology time slot. i Candidate paths, L is the set of links of the transmission path used by the previous topology time slot. i For candidate path p i The set of links.

10. A deterministic routing method for giant LEO constellations based on deep reinforcement learning according to claim 1, characterized in that, The final output selected path, as the end-to-end routing path that satisfies the deterministic constraints, includes: Compare the link duplication rates of all candidate paths and select the subset p of paths with the highest link duplication rate. max : Where, p prev p is the transmission path used by the previous topology time slot. i Candidate paths, For two paths p prev With p i Link duplication rate; If |p max If |=1, meaning there is only one path with the highest repetition rate, then that path p i The forwarding path directly determined as the current time slot If |p max If |>1, meaning there are multiple candidate paths with the same repetition rate, a secondary selection needs to be made based on their transmission performance bonus, choosing the path with the highest bonus value as the optimal path. in, For the Kth step in the future on path p i The instant reward obtained above, where γ is the discount factor. The step size for calculating cumulative rewards, where N is the number of paths; The final determined path It serves as a new deterministic forwarding path for time-sensitive services in the current time-slot topology and is delivered to the routing execution module for packet forwarding.