Internet constellation dynamic routing optimization method and system

Through deep reinforcement learning and multi-agent collaborative decision-making, the routing optimization problem under the dynamic topology changes of low-orbit satellite networks was solved, resource utilization and communication efficiency were improved, signaling overhead was reduced, and efficient and reliable routing decisions were achieved.

CN120658306APending Publication Date: 2025-09-16SUN YAT SEN UNIVERSITY SHENZHEN +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510976187.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

When faced with dynamic topology changes, existing low-orbit satellite network routing technology has problems such as high path failure rate, low resource utilization, large signaling overhead and decision lag, and is unable to effectively cope with the dynamic changes of low-orbit satellite networks.

Method used

It adopts deep reinforcement learning algorithm, combined with multi-agent collaborative decision-making mechanism, through dynamic network perception, dynamic reward design of multi-objective optimization, and uses actor-critic network and MADDPG framework to achieve dynamic routing optimization.

Benefits of technology

It significantly improves routing performance, resource utilization efficiency and task adaptability, reduces signaling overhead and decision conflict rate, and achieves efficient and reliable communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658306A_ABST
    Figure CN120658306A_ABST
Patent Text Reader

Abstract

The invention discloses an internet constellation dynamic routing optimization method and system, and the method comprises the steps: setting orbit plane parameters and satellite parameters, and carrying out the modeling of an orbit and a constellation; performing link visualization and time-varying graph generation according to ephemeris data and the constructed constellation model; defining a related space, a network and a function based on a deep reinforcement learning algorithm; and based on an MADDPG framework, introducing a multi-agent collaborative decision-making mechanism, and outputting action selection and execution results after collaborative optimization. The system comprises a model construction module, an embedding module, a reinforcement learning design module and a collaborative optimization module. By using the method, the dynamic change of the low-orbit satellite network can be effectively dealt with, satellite resources are fully utilized, and efficient and reliable communication is realized. The method can be widely applied to the field of wireless communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communications, and in particular to a method and system for dynamic routing optimization of an Internet constellation. Background Art

[0002] Low-Earth Orbit (LEO) internet constellations deploy a large number of satellites to build a global communications network, offering the advantages of low latency, high bandwidth, and wide-area connectivity. However, the high-speed movement of satellites leads to continuous and dynamic changes in network topology, frequent handoffs between inter-satellite links, and time-varying communication link quality (such as latency, bandwidth, and bit error rate). In this context, routing algorithms must adapt to changes in topology and link status in real time to ensure efficient and reliable data transmission.

[0003] Existing routing technologies are mainly divided into two categories: static routing and dynamic routing. However, they face significant limitations in low-orbit constellation scenarios: static routing algorithms rely on offline topology prediction to pre-configure routing tables and cannot respond to real-time status changes of inter-satellite links, resulting in high path failure rate and low resource utilization; although dynamic routing algorithms can perceive network changes, their convergence speed lags behind the speed of satellite network topology changes, resulting in frequent routing oscillations and global flooding-style link status updates. In ultra-large-scale constellations, this will generate non-negligible signaling overhead, exacerbating the defect of link resource competition. Summary of the Invention

[0004] In view of this, in order to solve the technical problems of the inability of existing routing technologies to effectively cope with the dynamic changes of low-orbit satellite networks and insufficient resource utilization, in a first aspect, the present invention proposes a method for dynamic routing optimization of Internet constellations, which includes the following steps:

[0005] Set orbital plane parameters and satellite parameters to model orbits and constellations;

[0006] Perform link visualization and time-varying graph generation based on ephemeris data and the constructed constellation model;

[0007] Based on deep reinforcement learning algorithms, define relevant spaces, networks, and functions;

[0008] The link quality parameters (delay, bandwidth, bit error rate, etc.) calculated in the time-varying graph are used to define the state space and action reward function. The action space is defined based on visibility calculation and the time-varying graph. The Actor-Critic network is arranged and a multi-agent collaborative decision-making mechanism is designed based on the current network state composed of link quality parameters and the network topology defined by visibility and the time-varying graph.

[0009] Based on the MADDPG framework, a multi-agent collaborative decision-making mechanism is introduced to output collaboratively optimized action selection and execution results.

[0010] In some embodiments, further comprising:

[0011] Execution logging, feedback collection, distributed model parameter updates, and simulation analysis.

[0012] In a second aspect, the present invention proposes an Internet constellation dynamic routing optimization system, comprising

[0013] Model building module, setting orbital plane parameters and satellite parameters, and modeling orbits and constellations;

[0014] Embedding module, which performs link visualization and time-varying graph generation based on ephemeris data and the constructed constellation model;

[0015] Reinforcement learning design module, based on deep reinforcement learning algorithm, defines relevant space, network and function;

[0016] The collaborative optimization module, based on the MADDPG framework, introduces a multi-agent collaborative decision-making mechanism and inputs the action selection and execution results after collaborative optimization.

[0017] Based on the above scheme, the present invention provides an Internet constellation dynamic routing optimization method and system, which significantly improves routing performance, resource utilization efficiency and task adaptability through dynamic network perception, dynamic reward design of multi-objective optimization and an efficient hybrid collaboration mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of the steps of a method for optimizing dynamic routing of an Internet constellation according to the present invention;

[0019] Figure 2 is a schematic diagram of constellation modeling according to a specific embodiment of the present invention;

[0020] Figure 3 This is a data flow diagram of a method for optimizing dynamic routing of an Internet constellation according to a specific embodiment of the present invention;

[0021] Figure 4 It is a structural block diagram of an Internet constellation dynamic routing optimization system of the present invention. DETAILED DESCRIPTION

[0022] The deep reinforcement learning routing scheme proposed in recent years has demonstrated the potential for dynamic optimization by online learning of the mapping relationship between network state and action. However, there are still bottlenecks such as single-agent decision-making bias towards local optimality, insufficient generalization, insufficient multi-agent collaboration efficiency, and suboptimal dynamic reward design.

[0023] In summary, current LEO constellation routing technology urgently needs to address issues such as efficiently characterizing the time-varying link states of high-speed mobile satellite networks and associating them with mission requirements; balancing latency minimization, bandwidth maximization, and real-time decision-making to enhance link stability; and enabling lightweight information exchange and strategic coordination between intelligent agents under the constraints of onboard computing resources. Therefore, an intelligent routing algorithm is urgently needed that can effectively cope with the dynamic changes of LEO satellite networks, fully utilize satellite resources, and achieve efficient and reliable communications.

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] It should be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of this application can be combined with each other.

[0026] It should be understood that the terms "system," "device," and / or "module" used in this application are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, other terms may be used to replace the terms if they can achieve the same purpose.

[0027] As used in this application and the claims, unless the context clearly indicates an exception, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular and may include the plural, unless the context clearly indicates otherwise. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements. The phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus that includes the elements.

[0028] In the description of the embodiments of this application, "plurality" refers to two or more than two. The terms "first" and "second" below are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0029] In addition, flow charts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0030] Reference Figure 1 , which is a flow chart of an optional example of the method for optimizing dynamic routing of an Internet constellation proposed by the present invention. This method can be applied to computer devices. The optimization method proposed in this embodiment may include but is not limited to the following steps:

[0031] Step S1, experimental parameter setting: setting orbital plane parameters and satellite parameters; constellation modeling and topology generation: building a constellation model;

[0032] Step S2, dynamic network status perception: based on the ephemeris data and the constellation model, calculate the visibility and link quality between each pair of satellites and generate a time-varying graph;

[0033] Step S3, DRL routing algorithm design: Based on the dynamic network state information generated in step S2 (including time-varying graph, state vector and real-time action space), define the state space, action space, actor-critic network and action reward function;

[0034] Step S4, multi-agent collaborative decision-making: Based on the constellation model, state space, action space, actor-critic network and action reward function, a multi-agent collaborative decision-making mechanism is introduced for simulation, and the action selection and execution results after collaborative optimization are output.

[0035] Step S5, recording and feedback: recording the path and performance indicators, and updating the experience pool;

[0036] Step S6, model parameter update: The model parameters include the parameters of the Actor network and the Critic network, and the local Critic network is globally averaged and then transmitted back;

[0037] Step S7: simulation analysis and verification.

[0038] In some feasible embodiments, step S1 specifically includes:

[0039] For the low-orbit Internet constellation, the Walker constellation with a polar orbit constellation configuration is constructed based on the Iridium constellation network. The parameters are set to 6 orbital planes, 11 satellites in each orbital plane, and a total of 66 satellites.

[0040] At the same time, intersatellite links are constructed, which are mainly divided into intra-orbit inter-satellite links and inter-orbit inter-satellite links. Intra-orbit inter-satellite links are constantly visible, while inter-orbit inter-satellite links are affected by dynamic topology changes. In addition, in low-orbit satellite constellations, due to the opposite movement directions of adjacent orbital planes, a situation called a reverse gap will occur. This situation makes it impossible for satellites on both sides of the reverse gap to establish inter-satellite links for communication. Therefore, satellites on the orbital planes on both sides of the reverse gap can only establish a maximum of 3 inter-satellite links. Figure 2 As shown in Figure 2, during the modeling process, the set of neighbor nodes N(t) for each satellite is calculated in real time, providing a topological foundation for subsequent routing decisions. The initial constellation topology is generated based on satellite orbit parameters (such as semi-major axis, inclination, and right ascension of the ascending node) and real-time ephemeris data. When outputting the initial topology, the available action space A (i.e., the selectable next-hop nodes) for each node is also generated simultaneously.

[0041] In some feasible embodiments, step S2 specifically includes:

[0042] Input the real-time ephemeris data of satellite nodes, including satellite position, speed, inter-satellite link status (delay L delay , remaining bandwidth B avl , bit error rate BER), ground mission requirements and other parameters (delay tolerance D tol , Bandwidth requirement B req ).

[0043] Based on ephemeris data and orbital mechanics models, the visibility and link quality between each pair of satellites are calculated to generate a time-varying graph (TVG). Nodes represent satellites, edges represent inter-satellite links, and weights are current link state parameters. This constructs a time-varying topology model.

[0044] Calculate the relative delay L by satellite distance delay :

[0045]

[0046] According to the total link capacity B total And the currently occupied bandwidth B used , calculate the remaining bandwidth:

[0047] B avl =B total -B used #(2)

[0048] Estimate the bit error rate BER by using the signal-to-noise ratio SNR and the Gaussian tail function Q

[0049]

[0050] The dynamic parameter update period is set to Δt = 0.2s to adapt to the link changes caused by the high-speed movement of the satellite.

[0051] For satellite nodes i Constructing the state vector

[0052]

[0053] Where Φ(Task) is a floating point number, which is based on the priority encoding of the task type, and is expressed by the task requirement vector [D tol , B req ]Mapping is generated.

[0054]

[0055] In the state perception phase, the state vector s is broadcast in real time. i It is given to neighboring nodes as input for routing decisions. According to the current TVG, the action space A of each satellite is dynamically adjusted to ensure that the decision is based on the latest network status.

[0056] Finally, the system outputs the dynamically updated TVG and state vector set {s1,s2,...,s n}, as well as the real-time action space, which will serve as the input for the subsequent DRL routing algorithm design and multi-agent collaborative decision-making steps.

[0057] In this preferred embodiment, based on time-varying graph modeling and combined with real-time ephemeris data, network topology changes are captured with a 0.2-second update cycle, and an efficient dynamic network perception method is proposed to provide real-time and accurate basic data for routing decisions.

[0058] A multi-dimensional state vector is constructed to integrate link status (delay, bandwidth, bit error rate) and task requirements to achieve comprehensive network characterization; the task requirement Φ(Task) is innovatively embedded in the state vector to achieve adaptive optimization of routing strategies for diverse tasks.

[0059] In some feasible embodiments, step S3 specifically includes:

[0060] S3.1, State space design: According to the link state vector s i (t), task requirement vector r i (t), integrate link status and task requirements, construct a state tensor, and output a high-dimensional state representation S with a dimension of 5, which contains 3 link parameters and 2 task parameters. i (t).

[0061] r i (t)=[D tol , B req ]#(6)

[0062] S i (t)=[si (t),r i (t)]#(7)

[0063] S3.2, Action Space Design: According to the current satellite visible neighbor node set N(t), the next hop selection is mapped into a discrete action space A = {a1, a2, ..., a n}, each action a j Correspondingly select neighbor node N j Then, according to the real-time time-varying topology, if a neighbor node is unavailable (link is disconnected), the corresponding action is removed from A. Output action set A = {a1, a2, ..., a n} and its probability distribution π(a|s i ).

[0064] S3.3. Construct an Actor-Critic network, where the Actor network uses a three-layer fully connected neural network, including an input layer, two hidden layers, and an output layer, outputting π(a|s i ). The critic network uses a dual critic network. Critic1 and Critic2 have the same structure, both containing an input layer, two hidden layers, and an output layer, but their parameters are independent of each other and are recorded as θ1 and θ2 respectively. When the algorithm is running, each time a data packet is received, the current state S is immediately i (t) Input the Actor Network, perform real-time inference, and output the next-hop selection probability. The Actor Network generates a policy π(a'|s') for the next state s' and selects action a'.

[0065] The two Critic networks calculate the loss based on the current state s and action a, as well as the target Q value y:

[0066] L critic1 =E[(y-Q1(s,a)) 2 ]#(8)

[0067] L critic2 =E[(y-Q2(s,a)) 2 ]#(9)

[0068] Critic1 and Critic2 independently update parameters θ1 and θ2 using gradient descent through their respective loss functions.

[0069] The Actor network updates its policy using a policy gradient method based on the Critic's feedback. To achieve the delay prediction of the subsequent N-hop path, the target Q value is expanded into a multi-step time difference (TD) form, considering the rewards of the next N steps and the value estimation of the N+1 step. Assuming N = 3 (3-hop prediction), the rewards of the next 3 steps and the value of the 4th step are considered. According to the immediate reward r of the kth step{t+k} (k=0,1,2), state s after step 3 {t+3} , Actor network in s {t+3} Action a selected on {t+3} and the discount factor γ modifies the target Q value y to:

[0070] y=r t +γr {t+1} +γ 2 r {t+2} +γ 3 min(Q1(s {t+3} ,a {t+3} ),Q2(s {t+3} ,a {t+3} ))#(10)

[0071] The agent needs to store the experience data of the last N steps, including the state s t to s {t+N} 、Action a t to a {t+N-1} and reward r t to r {t+N-1} Extract N steps of data from the experience buffer and use the Actor network in s {t+3} Generate action a {t+3} , calculate the output of Critic1 and Critic2, take the smaller value as the value estimate of the N+1th step, combine the N-step rewards, and calculate the target Q value y.

[0072] S3.4. Generate dynamic rewards based on the current network state S(t), execution action a, and task type T, and design a real-time reward function:

[0073]

[0074] where w d ,w b ,w s are the weights of delay, bandwidth and stability respectively, and their initial values ​​are all 0.33. task It is the task priority penalty item, if the task is not satisfied, negative rewards will be added. delay >D tol or B avl req , then I task =-1, otherwise I task = 0. η is the penalty coefficient, which is set to 0.5 by default. d ,w b ,w s Dynamically adjust according to the current network load, when the network is congested (bandwidth utilization ≥ 80%, increase w b Weight, adjusted to w b ​=0.5,w d =0.3,w s =0.2 Prioritize bandwidth. Increase weight w when triggering low-latency tasks d , suppress other weights and adjust to w d =0.7,w b =0.2,w s = 0.1. After each decision is executed, the reward value R(t) is immediately calculated based on the actual forwarding results (latency, bandwidth, bit error rate, etc.), and R(t) is fed back to the Actor-Critic network to adjust the next-hop selection strategy in real time and output the real-time calculated reward value R(t).

[0075] Through this embodiment, a multi-objective dynamic reward function is designed, which can adaptively adjust the weight w according to the network status (such as bandwidth utilization) and task type (low latency / high bandwidth). d ,w b ,w s For example, when the network congestion exceeds 80%, the bandwidth weight w is increased. b To 0.5; introduce task priority penalty I task , ensuring that the routing strategy meets the task requirements; using nonlinear reward mapping to effectively compensate for the long-tail delay distribution and improve the optimization ability of delay-sensitive tasks.

[0076] This paper proposes a dynamic weight adjustment mechanism based on network load and task requirements, which achieves a flexible balance between latency, bandwidth and stability, and significantly improves the end-to-end latency compliance rate of bursty low-latency tasks; it innovatively applies nonlinear reward mapping to routing optimization, avoiding the strategy convergence deviation caused by linear reward functions.

[0077] In addition, the Critic network introduces a two-layer network structure to separate action evaluation and environment modeling, improving the accuracy of delay prediction; the action space is dynamically adjusted according to the real-time time-varying topology to ensure the adaptability of the algorithm.

[0078] In some feasible embodiments, a proximal strategy is used to optimize the Actor-Critic network, which specifically includes:

[0079] Using the Proximal Policy Optimization (PPO) algorithm and parameters Advantage function Update the strategy to prevent strategy mutation:

[0080]

[0081] This paper uses the PPO algorithm to optimize the Actor-Critic network, limiting the policy update amplitude by clipping the loss function to reduce policy oscillation. A routing optimization method combining PPO with a two-layer critic network is proposed, achieving both improved policy stability and value estimation accuracy, suitable for highly dynamic satellite networks.

[0082] In some feasible embodiments, step S4 specifically includes:

[0083] S4.1, Agent Design: Each satellite acts as an independent agent, running an Actor-Critic network, responsible for making routing decisions based on local information and neighbor status. Through local decision-making and limited information sharing, each satellite agent broadcasts a local state summary (compressed into a 32-bit hash value) through the inter-satellite link and receives the neighboring nodes N (S i ) status summary.

[0084] S4.2. Conflict Avoidance Strategy: Designing a Joint Action Counterfactual Baseline (JACB): When the agent decides on the next hop, it predicts the link congestion probability P based on local observations and neighbor state summaries. cong , through a small neural network (input is s i and neighbor summary, the output is 0-1 probability) calculation. If the predicted value exceeds the threshold β = 0.7, the action is resampled. During the training phase, the agent records the local state s at each routing decision i , neighbor status summary and actual link congestion.,Generate binary classification labels based on performance,metrics.

[0085] According to the actual congestion label y, the predicted probability P cong Construct the loss function and use binary cross entropy loss:

[0086] L=-[y*log(P cong )+(1-y)*log(1-P cong )]#(13)

[0087] The Adam optimizer was used with a learning rate of 0.001. Batch training was performed based on data from the experience buffer, with network parameters updated every 100 decisions. This enabled the small neural network to accurately predict link congestion probabilities and assist the agent in selecting non-congested links.

[0088] S4.3, Decision-making stage: The agent changes its own state s i And the neighbor state summary is input into a small neural network to calculate P cong If P congIf the value is greater than β (threshold β = 0.7), the link is considered to be congested and the agent resamples actions, such as selecting a suboptimal action or randomly exploring other paths. In the centralized training phase, a global loss function is used to penalize conflicting actions of adjacent nodes.

[0089] in is the clipping loss of the PPO algorithm, and λ is the conflict penalty coefficient (default 0.1).

[0090]

[0091] During collaborative decision-making, each satellite selects the next hop based on its local state and the state of its neighbors. If a conflict is detected (e.g., multiple satellites selecting the same congested link), a new decision is triggered and a new action is executed. The collaboratively optimized action selection and execution results are ultimately output.

[0092] In this embodiment, a lightweight information sharing method is adopted, and each satellite node only broadcasts a 32-bit link state hash summary, which reduces the communication overhead compared with global state sharing. A Joint Action Counterfactual Baseline (JACB) mechanism is designed to predict the link congestion probability P through a small neural network. cong (threshold β = 0.7), dynamically adjust action selection to avoid decision conflicts; introduce a global loss function L in the centralized training stage global , add conflict penalty items to optimize individual strategies and group collaboration.

[0093] This paper proposes an efficient hybrid multi-agent collaboration mechanism that combines the advantages of distributed decision-making and limited information sharing, significantly reducing signaling overhead and routing conflict rate; innovatively applies the JACB mechanism to satellite routing decision-making, and improves collaboration efficiency through real-time congestion prediction.

[0094] In some feasible embodiments, step S5 specifically includes:

[0095] S5.1, Path selection and forwarding: According to the data packet to be forwarded and the current target node, the agent then distributes the action probability π(a|s i ) Select the next hop node and use an ε-greedy strategy to balance exploration and exploitation (initial ε = 0.3, decaying with training, decaying by 0.01 every 1000 decisions to ε = 0.05). Repeat the decision-making and forwarding process until the packet reaches the destination node, recording the complete path and performance indicators.

[0096] S5.2. Link status feedback: After the data packet is forwarded, the actual link transmission indicator is collected: delay offset ΔL delay , actual bandwidth utilization B used , bit error rate change ΔBER.

[0097] S5.3. Dynamic database update: Update local experience replay pool D i , prioritize high reward samples (PER prioritizes experience replay). Calculate sample priority based on γ = 0.9:

[0098] TD error =R+γ·max' a Q(s',a')-Q(s,a)#(15)

[0099] p i =|TD error |+ε(ε=0.01)#(16)

[0100] If a link failure occurs, the next hop is recalculated. If a routing loop occurs, the link is marked as unavailable and a new path is selected. Finally, the complete routing path from the source node to the destination node and the updated experience pool are output.

[0101] In some feasible embodiments, step S6 specifically includes:

[0102] S6.1, using the parameter sharing mechanism, the satellite node periodically transfers the local critic network gradient It is uploaded to the ground station, globally averaged, and then transmitted back to the constellation network.

[0103] In some feasible embodiments, step S6 further includes:

[0104] Update the model parameters according to the adaptive adjustment learning rate:

[0105] Adaptive learning rate adjustment: According to the network dynamic fluctuation index, the frequency of topological changes of the satellite network f topo (Unit: times / second), the initial learning rate α0 = 0.001, the attenuation coefficient γ = 0.2 used to suppress policy oscillations under high-frequency topology changes, and the learning rate α is adaptively adjusted.

[0106]

[0107] After the parameters are updated, the new model is immediately applied to perform the next routing decision and the updated global model parameters are output.

[0108] This preferred embodiment proposes a distributed model update method for low-orbit constellations, reducing computational and communication overhead and accelerating model convergence. It also innovatively incorporates the frequency of topology changes into learning rate adjustments, enhancing the model's stability in highly dynamic environments.

[0109] Based on the above scheme, the rectification data flow direction of the method of the present invention is referenced Figure 3 .

[0110] In some feasible embodiments, step S7 specifically includes:

[0111] Using STK+NS-3 simulation, we evaluate indicators such as routing delay, bandwidth utilization, and packet loss rate, verify the real-time and effectiveness of decision execution in the simulation, and output performance verification results.

[0112] Reference Figure 4 , an Internet constellation dynamic routing optimization system, comprising:

[0113] A model building module, configured to execute step S1;

[0114] An embedded module, configured to execute step S2;

[0115] A reinforcement learning design module, configured to execute step S3;

[0116] The collaborative optimization module is used to execute step S4.

[0117] A recording module, configured to execute step S5;

[0118] A parameter updating module, configured to execute step S6;

[0119] The verification module is used to execute step S7.

[0120] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0121] An Internet constellation dynamic routing optimization device:

[0122] at least one processor;

[0123] at least one memory for storing at least one program;

[0124] When the at least one program is executed by the at least one processor, the at least one processor implements the Internet constellation dynamic routing optimization method as described above.

[0125] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0126] A storage medium stores processor-executable instructions, which, when executed by a processor, are used to implement the above-mentioned method for dynamic routing optimization of an Internet constellation.

[0127] The contents of the above method embodiments are all applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0128] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for dynamic routing optimization of Internet constellations, characterized in that: The following steps are involved: Set orbital plane parameters and satellite parameters, and build a constellation model; Based on the ephemeris data and the constellation model, the visibility and link quality between each pair of satellites are calculated, and a time-varying graph is generated; Based on the time-varying graph, define the state space, action space and action reward function, and arrange the Actor-Critic network; Based on the constellation model, the state space, the action space, the Actor-Critic network and the action reward function, a multi-agent collaborative decision-making mechanism is introduced for simulation, and the collaboratively optimized action selection and execution results are output.

2. The method for dynamic routing optimization of an Internet constellation according to claim 1, characterized in that: Also includes: Record paths and performance indicators, and update the experience pool.

3. The method for dynamic routing optimization of an Internet constellation according to claim 2, characterized in that: Also includes: The local critic network is globally averaged and then transmitted back.

4. The method for dynamic routing optimization of an Internet constellation according to claim 3, characterized in that: Also includes: Adaptively adjust the learning rate according to the frequency of satellite network topology changes; The model parameters of the Actor-Critic network are updated according to the adaptively adjusted learning rate.

5. The method for dynamic routing optimization of an Internet constellation according to claim 4, characterized in that: The step of calculating the visibility and link quality between each pair of satellites based on the ephemeris data and the constellation model, and generating a time-varying graph, specifically includes: Calculate relative delay based on satellite distance; Calculate the remaining bandwidth based on the total link capacity and the currently occupied bandwidth; Estimate the bit error rate through the signal-to-noise ratio and Gaussian tail function; Construct a state vector for the satellite; A time-varying graph is generated with satellites as nodes, inter-satellite links as edges, and current link state parameters as weights.

6. The method for dynamic routing optimization of an Internet constellation according to claim 5, characterized in that: The step of defining the state space, action space, actor-critic network and action reward function specifically includes: According to the link state vector and the task requirement vector, the link state and the task requirement are integrated to construct a state tensor and obtain the state space; Map the next hop selection into a discrete action space based on the set of visible neighbor nodes of the current satellite; Construct an Actor network and a Critic network. The Actor network updates the policy using a policy gradient method based on the Critic's feedback. Generate dynamic rewards based on the current network state, execution action and task type to obtain a real-time reward function.

7. The method for dynamic routing optimization of an Internet constellation according to claim 6, characterized in that: Also includes: According to the real-time time-varying topology, if a neighbor node is unavailable, the corresponding action is removed from the discrete action space.

8. The method for dynamic routing optimization of an Internet constellation according to claim 6, characterized in that: The real-time reward function is expressed as follows: Among them, w d ,w b ,w s are the weights of delay, bandwidth and stability respectively, I task is the task priority penalty term, η is the penalty coefficient, L delay Represents the relative delay, B avl It indicates the remaining bandwidth, and BER indicates the bit error rate.

9. The method for dynamic routing optimization of an Internet constellation according to claim 1, characterized in that: The step of introducing a multi-agent collaborative decision-making mechanism to simulate based on the constellation model, the state space, the action space, the Actor-Critic network, and the action reward function, and outputting the collaboratively optimized action selection and execution results specifically includes: The satellite is used as an independent intelligent agent to run the Actor-Critic network, which is responsible for generating routing decisions based on local information and neighbor status; Predict link congestion probability based on local observations and neighbor state summaries; Performing conflict judgment based on the link congestion probability; If a conflict is detected, a new decision is triggered and a new action is executed, and finally the collaboratively optimized action selection and execution results are output.

10. An Internet constellation dynamic routing optimization system, characterized in that: include: Model building module, setting orbital plane parameters and satellite parameters, and building constellation model; An embedding module calculates visibility and link quality between each pair of satellites based on ephemeris data and the constellation model, and generates a time-varying graph; A reinforcement learning design module defines the state space, action space, and action reward function based on the time-varying graph, and arranges the Actor-Critic network; The collaborative optimization module introduces a multi-agent collaborative decision-making mechanism for simulation based on the constellation model, the state space, the action space, the actor-critic network and the action reward function, and outputs the action selection and execution results after collaborative optimization.

Citation Information

Cited By

  • Variable-scale satellite group orbit planning method based on graph near-end optimization algorithm

    CN121032282A

  • On-demand routing agent system for low earth orbit satellite network and generation method

    CN121567192A

  • An On-Demand Routing Agent System and Generation Method for Low-Earth Orbit Satellite Networks

    CN121567192B