Low earth orbit satellite constellation dynamic hopping beam variation optimization method based on deep reinforcement learning

Through the deep reinforcement learning method of distributed multi-agents, the problem of dynamic traffic and channel conditions changes in the LEO satellite network is solved, efficient resource optimization and throughput improvement are achieved, latency and communication overhead are reduced, and it is suitable for large-scale LEO satellite constellations.

CN120342471APending Publication Date: 2025-07-18NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510647691.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing low-orbit satellite network, the traditional static beam distribution method cannot effectively deal with dynamic traffic demand and uneven user distribution, resulting in low resource usage efficiency. The existing dynamic beam jump distribution method ignores channel condition fluctuations, high computational complexity, and large communication overhead.

Method used

The distributed multi-agent deep reinforcement learning method is adopted, and the local interaction Markov game model and multi-agent deep Q network are combined with the local collaboration reward mechanism to realize distributed decision-making and resource optimization among satellites, and a multi-objective optimization model is built to balance throughput and delay.

Benefits of technology

It significantly improves the throughput of LEO satellite network, reduces communication latency, and reduces communication overhead, has good scalability, and is suitable for large-scale LEO satellite constellation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342471A_ABST
    Figure CN120342471A_ABST
Patent Text Reader

Abstract

The invention discloses a low earth orbit satellite constellation dynamic hopping beam variation optimization method based on deep reinforcement learning, solves the problem of dynamic hopping beam optimization in LEO satellite constellations, and realizes intelligent decision making of satellites only depending on local information through a distributed multi-agent deep reinforcement learning method. The system throughput is improved; and the communication delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite communication, and particularly relates to a dynamic hopping beam optimization method for low-earth orbit satellite constellations based on deep reinforcement learning. Background Art

[0002] With the rapid development of low-earth orbit (LEO) satellite networks, traditional static beam allocation methods have been unable to effectively address the challenges brought by dynamic traffic demands and uneven user distributions. Due to its fast orbital motion, the network topology of LEO satellites is highly dynamic. At the same time, the dense deployment of satellites leads to a complex interference environment. Coupled with the changing characteristics of user traffic, these factors together make resource optimization extremely complex.

[0003] In traditional multi-beam satellite communication systems, resources are restricted to specific beams, resulting in resource fragmentation and reduced flexibility. This static allocation method cannot adapt to the dynamic characteristics of LEO satellite networks, leading to low resource utilization efficiency. To address these limitations, researchers have begun to explore beam hopping (BH) technology, which allows dynamic scheduling of coverage areas and beam dwell times, and allocates beam resources on demand according to the real-time traffic distribution.

[0004] Existing beam hopping allocation methods mainly focus on static scenarios and pre-determine the beam illumination pattern within a certain period. Although these static methods can reduce system complexity by making advance plans, they lack the flexibility to adapt to the rapid changes in traffic demands and channel conditions in LEO satellite systems, resulting in increased traffic delays and degraded service quality.

[0005] To overcome these limitations, researchers have begun to develop dynamic beam hopping allocation schemes that adjust the beam illumination pattern and resource allocation according to real-time demands and environmental changes. In this context, deep reinforcement learning (DRL) technology has become a promising method. After effective training, the DRL agent can achieve efficient decision-making in a dynamic environment. Nevertheless, there are still several gaps in the current research on DRL for beam hopping allocation: most studies ignore the impact of channel condition fluctuations and mainly focus on traffic demands; most studies focus on centralized training frameworks, and the training process requires global information, which may lead to challenges such as higher computational complexity, communication overhead, and difficulty in scalability.

[0006] To address these issues, the present invention proposes a distributed multi-agent deep reinforcement learning method, specifically a method for optimizing dynamic hopping beams in a Low Earth Orbit (LEO) satellite constellation based on distributed multi-agent deep reinforcement learning (Multi-Agent Deep Reinforcement Learning, MARL), to solve the complex problem of hopping beam scheduling in LEO satellite constellations, and is particularly applicable to scenarios with dynamic traffic demands and unpredictable channel conditions. Summary of the Invention

[0007] The present invention provides a method for optimizing dynamic hopping beams in a LEO satellite constellation based on deep reinforcement learning, which solves the complex problem of hopping beam scheduling in LEO satellite constellations and is applicable to scenarios with dynamic traffic demands and unpredictable channel conditions.

[0008] To achieve the above objectives, the technical solution adopted by the present invention to solve the above technical problems is:

[0009] A method for optimizing dynamic hopping beams in a LEO satellite constellation based on deep reinforcement learning, which realizes the optimization method based on an integrated ground and space network system. The integrated ground and space network system consists of a ground gateway, a LEO satellite constellation with a Walker structure, and multiple ground cells; the ground gateway transmits data to the cells through LEO satellites;

[0010] Let the set of LEO satellites be denoted as M = {1, 2,..., m,..., M}, and the set of cells be denoted as N = {1, 2,..., n,..., N}; the set of beams of the m-th LEO satellite is denoted as C m = {1, 2,..., c,..., C}, and each beam precisely covers a single cell;

[0011] Let the set of cells served by the m-th LEO satellite be denoted as C m where and the coverage areas of different LEO satellites do not overlap; thus, the entire LEO constellation can comprehensively cover the set of cells N, specifically:

[0012] and for m ≠ m'

[0013] Since the total number of beams of LEO satellites is much less than the number of cells, C·M << N. LEO satellites implement the hopping beam technology through a time slicing mechanism to effectively provide access services to all covered cells;

[0014] Let N mDenote the buffer queue number of the m-th LEO satellite, and each data queue is associated with a predetermined cell; when the ground gateway transmits data to the satellite, if a transmission link is established between the satellite and the target cell, the data is successfully transmitted; if the transmission condition is not satisfied, the data enters the corresponding queue and waits for transmission in the next time slot; if the data waiting time exceeds the delay threshold T th , the data packet is discarded;

[0015] The downlink spectrum of the LEO satellite beam adopts a multi-color multiplexing technology. The total system duration is divided into consecutive time intervals Δt, denoted as T = {1, 2,..., t,..., T}; assume that the LEO satellite position and the cells illuminated by its beam remain static within each time slot; new data packets from the ground gateway are received by the satellite at the end of each time slot.

[0016] Set λ n (t) represents the number of data packets arriving at cell n within time slot t. Assume the data packet size is a constant M0, and λ n (t) follows a Poisson distribution;

[0017] Construct a multi-objective optimization model to effectively balance the two performance metrics of throughput and delay in the LEO satellite network; the specific optimization model is as follows:

[0018]

[0019] Among them, ω1, ω2 ∈ [0, 1] represent the weighted coefficients for weighing the trade-off between parameterized throughput maximization and delay minimization under the normalization constraint ω1 + ω2 = 1;

[0020] Constraint C1 establishes a quality of service threshold by stipulating that the signal-to-interference-plus-noise ratio of each cell n in each time slot t must exceed the minimum acceptable threshold γ min to ensure reliable communication connectivity under different interference conditions; Constraint C2 imposes a non-negativity condition on the remaining data packets in the buffer queue to maintain the physical feasibility of the queue system; Constraint C3 sets an upper bound Q on the cumulative queue length of each cell max , prevent buffer overflow, and enforce strict quality of service requirements for delay-sensitive services in the entire satellite network.

[0021] Furthermore, based on the multi-objective optimization model, construct a locally interactive Markov game model, defined as a tuple G = (M, S, A, N, T, R, γ), where M represents the set of satellites, S represents the joint state space, A represents the joint action space, N represents the neighbor structure, T represents the state transition function, R represents the set of reward functions, and γ represents the discount factor.

[0022] Furthermore, based on the Markov game model with local interaction, a multi-agent deep Q-network MDQN-LCR algorithm with local cooperation rewards is constructed to solve the hopping beam scheduling problem in the LEO satellite constellation through a distributed decision-making and cooperation optimization mechanism. The Markov game model with local interaction includes a distributed Q-learning structure and a cooperation reward mechanism. In the distributed Q-learning structure, each satellite agent obtains its own and neighboring satellite information through an abstract state space to achieve state sharing among neighboring satellites. Each agent maintains an independent evaluation network to calculate the action value Q(s, a) and selects the optimal action based on the policy π(s|θ). The decision-making process realizes stable learning through a dual-network architecture, including a target network with periodic soft updates and an evaluation network optimized by policy gradients. Multiple satellite agents generate actions in parallel based on environmental observations to form a joint action decision, and the joint action decision is directly mapped to the illumination pattern of K beam positions, affecting the overall throughput and latency performance of the system. The cooperation reward mechanism enables satellites to share local utility information and promotes global coordination of distributed decision-making. Each agent stores interaction data through an independent experience memory and updates network parameters through mini-batch sampling.

[0023] Furthermore, the MDQN-LCR specifically includes a distributed Q-learning framework and a cooperation reward mechanism.

[0024] The distributed Q-learning framework enables each satellite to evaluate action values and make decisions based on local and neighboring observations.

[0025] The cooperation reward mechanism promotes cooperative behavior among agents through information sharing.

[0026] At each time slot t, each agent m first observes its local state S from the environment t,m , including information from neighboring satellites (N m ), to form an abstract state space. The agent then uses its evaluation network to calculate the Q value Q m (s, a) and selects an action a through an ε-greedy policy t,m . After all agents execute the joint action, the environment transitions to the next state, and at the same time, feedback is provided through the cooperation reward mechanism.

[0027] A dynamic hopping beam optimization method for low Earth orbit (LEO) satellite constellations based on distributed multi-agent deep reinforcement learning (MARL) of the present invention aims to solve the resource allocation problem caused by dynamic traffic demands and uneven user distribution in LEO satellite networks. For the hopping beam scheduling problem in LEO satellite constellations, the present invention first establishes a multi-objective optimization model that takes into account both maximizing system throughput and minimizing communication latency. Then, a Markov game model with local interaction is constructed. On this basis, a multi-agent deep Q-network combined with local cooperative rewards (MDQN-LCR) algorithm is designed, enabling satellites to make intelligent decisions based only on local observations without global information exchange, thereby effectively improving system throughput and reducing communication latency, while significantly reducing communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a schematic diagram of the MDQN-LCR algorithm framework of the present invention;

[0030] Figure 2 It is a schematic diagram of the local Q-network structure of the present invention;

[0031] Figure 3 It is a schematic diagram of the reward curve during the training process of the present invention;

[0032] Figure 4 It is a schematic diagram of the loss curve during the training process of the present invention;

[0033] Figure 5 It is a schematic diagram of the performance under different traffic demands of the present invention;

[0034] Figure 6 It is a schematic diagram of the performance under different time slot durations of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0036] The present invention provides a dynamic hopping beam optimization method for a low-earth orbit satellite constellation based on deep reinforcement learning. The optimization method is implemented based on an integrated ground and space network system, and the integrated ground and space network system consists of a ground gateway, a low-earth orbit (LEO) satellite constellation with a Walker structure, and multiple ground cells. The ground gateway transmits data to the cells through the LEO satellites.

[0037] Let the set of LEO satellites be denoted as M = {1, 2,..., m,..., M}, and the set of cells be denoted as N = {1, 2,..., n,..., N}. The set of beams of the m-th LEO satellite is denoted as C m = {1, 2,..., c,..., C}, and each beam precisely covers a single cell.

[0038] Let the set of cells served by the m-th LEO satellite be denoted as C m , where and the coverage areas of different LEO satellites do not overlap. Therefore, the entire LEO constellation can fully cover the set of cells N, specifically:

[0039] and for m ≠ m'

[0040] Since the total number of beams of LEO satellites is much less than the number of cells, C·M << N. The LEO satellites implement the hopping beam technology through a time-division multiplexing mechanism to effectively provide access services to all covered cells.

[0041] Let N m represent the number of buffer queues of the m-th LEO satellite, and each data queue is associated with a predetermined cell. When the ground gateway transmits data to the satellite, if a transmission link is established between the satellite and the target cell, the data is successfully transmitted. If the transmission conditions are not met, the data enters the corresponding queue and waits for transmission in the next time slot. If the data waiting time exceeds the delay threshold T th , the data packet is discarded.

[0042] The downlink spectrum of the LEO satellite beam adopts a multi-color multiplexing technology. The total system duration is divided into consecutive time intervals Δt, denoted as T = {1, 2,..., t,..., T}. It is assumed that the LEO satellite position and the cells illuminated by its beam remain static within each time slot. New data packets from the ground gateway are received by the satellite at the end of each time slot.

[0043] Let λ n (t) represent the number of data packets arriving at cell n within time slot t. The data packet size is assumed to be a constant M0, and λ n (t) follows a Poisson distribution.

[0044] A multi-objective optimization model is constructed to effectively balance the two performance metrics of throughput and delay in the LEO satellite network. The specific optimization model is as follows:

[0045]

[0046] Among them, ω1, ω2 ∈ [0, 1] represent the weighted coefficients for weighing the trade-off between parameterized throughput maximization and delay minimization under the normalization constraint ω1 + ω2 = 1. Constraint C1 establishes a quality-of-service threshold by stipulating that the signal-to-interference-plus-noise ratio of each cell n in each time slot t must exceed the minimum acceptable threshold γ min to ensure reliable communication connectivity under different interference conditions. Constraint C2 imposes a non-negativity condition on the remaining data packets in the buffer queue to maintain the physical feasibility of the queue system. Constraint C3 sets an upper bound Q on the cumulative queue length of each cell to prevent buffer overflow and enforce strict quality-of-service requirements for delay-sensitive services in the entire satellite network. max

[0047] Furthermore, based on the multi-objective optimization model, a locally interactive Markov game model is constructed, defined as a tuple G = (M, S, A, N, T, R, γ), where M represents the set of satellites, S represents the joint state space, A represents the joint action space, N represents the neighbor structure, T represents the state transition function, R represents the set of reward functions, and γ represents the discount factor.

[0048] ​Furthermore, based on the Markov game model of local interaction, a multi-agent deep Q-network MDQN-LCR algorithm with local cooperation rewards is constructed to solve the hopping beam scheduling problem in the LEO satellite constellation through a distributed decision-making and cooperation optimization mechanism. The Markov game model of local interaction includes a distributed Q-learning structure and a cooperation reward mechanism. In the distributed Q-learning structure, each satellite agent obtains its own and neighboring satellite information through an abstract state space to achieve state sharing among neighboring satellites. Each agent maintains an independent evaluation network to calculate the action value Q(s, a) and selects the optimal action based on the policy π(s|θ). The decision-making process realizes stable learning through a dual-network architecture, including a target network updated periodically in a soft manner and an evaluation network optimized through policy gradients. Multiple satellite agents generate actions in parallel based on environmental observations to form a joint action decision, and the joint action decision is directly mapped to the illumination patterns of K beam positions, affecting the overall throughput and latency performance of the system. The cooperation reward mechanism enables satellites to share local utility information and promotes global coordination of distributed decision-making. Each agent stores interaction data through an independent experience memory and updates network parameters through mini-batch sampling.

[0049] Furthermore, the MDQN-LCR specifically includes a distributed Q-learning framework and a cooperation reward mechanism;

[0050] The distributed Q-learning framework enables each satellite to evaluate action values and make decisions based on local and neighboring observations;

[0051] The cooperation reward mechanism promotes cooperative behavior among agents through information sharing;

[0052] At each time slot t, each agent m first observes its local state S from the environment t,m , including information from neighboring satellites (N m ), to form an abstract state space. The agent then uses its evaluation network to calculate the Q value Q m (s, a) and selects the action a through an ∈-greedy policy t,m ; after all agents execute the joint action, the environment transitions to the next state, and at the same time, feedback is provided through the cooperation reward mechanism.

[0053] The integrated ground and space network system consists of a ground gateway, a LEO satellite constellation in a Walker formation, and a series of ground cells. Since users within a cell cannot directly access the ground Internet, the ground gateway transmits data to the cell through LEO satellites. The set of LEO satellites is denoted as M = {1, 2,..., m,..., M}, and the set of cells is denoted as N = {1, 2,..., n,..., N}. The beam set of the m-th LEO satellite is denoted as C m={1, 2, ..., c, ..., C}, and each beam precisely covers a single cell.

[0054] Suppose the set of cells served by the m-th LEO satellite is denoted as C m , where and the coverage areas of different LEO satellites do not overlap. Therefore, the entire LEO constellation can comprehensively cover the cell set N. Specifically:

[0055] and For m ≠ m′

[0056] Considering that the total number of beams of LEO satellites is much less than the number of cells, i.e., C·M << N, LEO satellites implement the hopping beam technology through the time slicing mechanism to effectively provide access services to all covered cells. The specific process is as follows: Let N m denote the number of buffer queues of the m-th LEO satellite, and each data queue is associated with a predetermined cell. When the ground gateway transmits data to the satellite, if a transmission link is established between the satellite and the target cell, the data is successfully transmitted; if the transmission condition is not met, the data enters the corresponding queue to wait for transmission in the next time slot. If the data waiting time exceeds the delay threshold T th , the data packet is discarded.

[0057] The downlink spectrum of LEO satellite beams adopts the multi-color multiplexing technology, ignoring the interference between different beams of the same satellite. The total duration of the system is divided into consecutive time intervals Δt, denoted as T = {1, 2, ..., t, ..., T}. Assume that the position of the LEO satellite and the cells illuminated by its beams remain static in each time slot. In addition, new data packets from the ground gateway are received by the satellite at the end of each time slot. Due to the uneven distribution of users, the traffic demands of each cell fluctuate significantly. Let λ n (t) represent the number of data packets arriving at cell n in time slot t, and the data packet size is assumed to be a constant M o , and λ n (t) follows a Poisson distribution. The symbols used in this invention are summarized in Table 1.

[0058] Table 1 Symbol Definitions

[0059]

[0060] Construct a multi-objective optimization model: This invention first constructs a multi-objective optimization model to effectively balance the two key performance indicators of throughput and delay in the LEO satellite network.

[0061]

[0062] In this formula, ω1, ω2 ∈ [0, 1] represent the weighted coefficients that trade off between parameterized throughput maximization and latency minimization under the normalization constraint ω1 + ω2 = 1. Constraint C1 establishes the quality-of-service threshold by stipulating that the signal-to-interference-plus-noise ratio of each cell n at each time slot t must exceed the minimum acceptable threshold γ min to ensure reliable communication connectivity under different interference conditions. Constraint C2 imposes a non-negativity condition on the remaining data packets in the buffer queue to maintain the physical feasibility of the queue system. Constraint C3 sets an upper bound Q on the cumulative queue length of each cell max , preventing buffer overflow and enforcing strict quality-of-service requirements for latency-sensitive traffic in the entire satellite network.

[0063] Based on the proposed multi-objective optimization model, the present invention proposes a locally interacting Markov game model (LocallyInteracting Markov Game Model), formally defined as the tuple G = (M, S, A, N, T, R, γ), where M represents the set of satellites, S represents the joint state space, A represents the joint action space, N represents the neighbor structure, T represents the state transition function, R represents the set of reward functions, and γ represents the discount factor.

[0064] This model has the following characteristics:

[0065] (1) Each LEO satellite acts as an agent and makes decisions based on local state information (including channel gain, buffer queue state, and interference level from neighboring satellites), without the need for global information exchange.

[0066] (2) The interaction relationship between satellites is defined by the neighbor structure N. When the distance between two satellites is less than the preset threshold d th , they are regarded as neighbors.

[0067] (3) The utility function of each satellite consists of three parts: local reward, cooperation reward, and interference penalty, formally expressed as:

[0068]

[0069] Through strict mathematical proof, this model is proven to be an exact potential game, with a potential function:

[0070]

[0071] where Θ t represents the system throughput, and Λ tDenote the system delay, and ω1 and ω2 are weight coefficients. This paper proves that when any satellite m unilaterally changes its strategy, the change in its utility function is equal to the change in the potential function:

[0072] U m (π′m, π-m) - U m (π m , π -m ) = Φ(π′m, π-m) - Φ(π m , π -m )

[0073] This theoretical proof guarantees that there exists at least one pure-strategy Nash equilibrium, corresponding to the local maximum of the potential function. This property provides a solid theoretical basis for the convergence of distributed learning algorithms, especially suitable for the case where centralized coordination is difficult to achieve due to communication constraints and computational limitations in satellite networks.

[0074] Based on the Markov game model with local interaction, this patent proposes a Multi-Agent Deep Q-Network with Local Cooperative Rewards (MDQN-LCR) algorithm to solve the hopping beam scheduling problem in LEO satellite constellations.

[0075] Figure 1 Figure shows the MDQN-LCR algorithm framework proposed by the present invention. This framework solves the hopping beam scheduling problem in LEO satellite constellations through a distributed decision-making and collaborative optimization mechanism. The core design of this framework is based on the Markov game model with local interaction and consists of two major components: a distributed Q-learning structure and a collaborative reward mechanism. In the distributed Q-learning structure, each satellite agent can be seen from the left process. By abstracting the state space, it obtains information about itself and neighboring satellites, realizing state sharing among neighboring satellites. Each agent maintains an independent evaluation network to calculate the action value Q(s, a), and selects the optimal action based on the policy π(s|θ). The decision-making process achieves stable learning through a dual-network architecture, including a target network with regular soft updates and an evaluation network optimized by policy gradients. This design effectively alleviates the instability problem in multi-agent environments and ensures the convergence performance of the algorithm. The environmental interaction process is clearly presented on the right. The environmental state includes information such as the data queue state and interference level of each cell. Multiple satellite agents generate actions in parallel based on environmental observations to form a joint action decision. These decisions are directly mapped to the illumination patterns of K beam positions, affecting the overall throughput and delay performance of the system.

[0076] The Cooperative Reward Mechanism is a key innovation of this framework and is located at the bottom right (Cooperative Reward Mechanism). This mechanism enables satellites to share local utility information and promotes global coordination of distributed decision-making. Each agent stores interaction data through independent Experience Memory and updates network parameters through Mini-batch sampling. The training process (marked by the red dashed line in the figure) forms a closed-loop optimization system that continuously iteratively improves the policy parameter θ, enabling the satellite to adapt to dynamically changing traffic demands and channel conditions. This design significantly reduces communication overhead while maintaining approximately optimal system performance, making it particularly suitable for distributed resource optimization problems in large-scale LEO satellite constellation scenarios.

[0077] Two key components of the MDQN-LCR algorithm are described as follows:

[0078] (1) Distributed Q-learning framework: Enables each satellite to evaluate action values and make decisions based on local and neighbor observations.

[0079] (2) Cooperative Reward Mechanism: Promotes cooperative behavior among agents through information sharing.

[0080] In each time slot t, each agent m first observes its local state S from the environment t,m , including information from neighboring satellites (N m ), to form an abstract state space. The agent then uses its evaluation network to calculate the Q-value Q m (s, a) and selects an action a through an ε-greedy policy t,m . After all agents execute the joint action, the environment transitions to the next state while providing feedback through the Cooperative Reward Mechanism.

[0081] For each LEO satellite agent m, a local Q-network Q m (s t , m, a m |θ m ) is designed to approximate the optimal action value function. This network has the following structure as shown Figure 2 and specifically includes:

[0082] (1) Input layer: Processes state information from the current satellite and its neighboring satellites, including channel state, queue length, and interference level.

[0083] (2) Hidden layer: Through multiple fully connected hidden layers (h (1) = ReLU(W (1) st,m + b (1) ) and h (l) = ReLU(W(l) h (l-1) +b (l) )(5) Extract relevant features from the combined state representation and learn the important associations between the agent's state and the states of other agents. Three hidden layers are adopted, each containing 40 nodes.

[0084] (3) Output layer: Generate an extended Q-value function considering the multi-agent environment:

[0085] Q m (s t,m , a t,m ) = W (out) h (L) +b (out) .

[0086] The training of the local Q-network follows the DQN framework, and its loss function is defined as:

[0087]

[0088] Among them, D m represents the experience replay buffer of agent m, Q'm represents the target network, and θ' m represents its parameters. During the training process, the stochastic gradient descent algorithm is used to optimize the loss function. The parameter update formula is:

[0089]

[0090] In addition, a soft update strategy is adopted to synchronize the parameters of the target network:

[0091]

[0092] Among them, τ << 1 is used as the soft update coefficient.

[0093] To encourage the cooperative behavior between satellites, the present invention proposes a cooperative reward mechanism based on local information exchange. At the end of each time slot, satellite m shares its current local utility value m with the satellites in the neighbor satellite set N and receives the utility information from these neighbors. This information exchange enables satellite m to comprehensively evaluate the impact of its actions on the overall system performance, thereby promoting more intelligent decision-making.

[0094] When updating the local Q-network, the cooperative reward R coop m(t) is incorporated into the immediate reward:

[0095]

[0096] Among them, represents the local reward, represents the cooperative reward, Represents the interference penalty term. Hyperparameters α, β, η are used to balance the weights of these reward components, allowing for adaptive adjustment of cooperation and local optimization. These parameters are calibrated according to the potential game formula to ensure convergence to the Nash equilibrium.

[0097] The pseudo-code of the MDQN-LCR algorithm is as follows:

[0098]

[0099]

[0100] Formulas 1, 2, and 3 are as follows:

[0101] Formula 1: N m = k ∈ M | d(m, k) < d th

[0102] Formula 2:

[0103] Formula 3:

[0104] Through this algorithm design, the MDQN-LCR algorithm can significantly reduce the communication overhead, while achieving excellent performance in terms of throughput, latency, convergence speed, and stability, and is particularly suitable for the hopping beam scheduling problem of LEO satellite constellations.

[0105] The specific implementation steps summarized according to the pseudo-code are as follows:

[0106] (1) System initialization: Set the orbital parameters, number of beams, number of serving cells, etc. of the LEO satellite constellation, and initialize the Q-network parameters of each satellite.

[0107] (2) State observation and action selection: At each time slot, each satellite observes its local state (including channel gain, buffer queue state, neighbor satellite interference level, etc.), calculates the action value using the Q-network, and selects the hopping beam pattern according to the ε-greedy policy.

[0108] (3) Execute action and feedback: All satellites execute the selected hopping beam pattern, and the environment feedbacks the reward according to the current state and action, and transitions to the next state.

[0109] (4) Q-value update and training: Each satellite updates its Q-value using the cooperation reward mechanism and stores the experience tuple in the local experience replay buffer. By sampling a small batch of data, the parameters of the Q-network are updated using the policy gradient method.

[0110] (5) Repeat execution: Repeat steps (2) to (4) until the algorithm converges or reaches the preset number of training rounds.

[0111] Verify the effectiveness of the MDQN-LCR algorithm through comprehensive simulation experiments and compare its performance with existing methods.

[0112] Construct a constellation system consisting of 16 LEO satellites with an orbital altitude of 600 km. Each satellite is equipped with 4 beams to serve 64 ground cells in total. The communication link operates in the Ka band (20 GHz) with a bandwidth of 250 MHz. Simulate the communication demand based on the existing channel model and the packet arrival rate that follows a Poisson distribution. Regarding the algorithm hyperparameters, set the discount factor γ = 0.95, the target network soft update coefficient τ = 0.001, and train for 1000 episodes. For comprehensive evaluation, select four comparison algorithms: the random algorithm, the greedy algorithm, the centralized QMIX, and the MDQN-NLCR variant without cooperative rewards.

[0113] Experiments show that MDQN-LCR exhibits obvious advantages in terms of learning efficiency and stability, such as Figure 3 and Figure 4 . It achieves a faster convergence speed and higher reward values at the initial stage of training; compared with the centralized QMIX, the reward variance is reduced by 42%, showing more stable performance; compared with the MDQN-NLCR without a cooperation mechanism, the final reward value is significantly improved.

[0114] In a dynamic traffic environment, as Figure 5 shown, in terms of throughput: in the case of high load (80 Mbps), MDQN-LCR reaches a peak throughput of 9×104 Mbps, which is about 2.3% higher than that of the centralized QMIX. In terms of latency: in the traffic range of 50 - 80 Mbps, it maintains a low latency of 18.66 - 21.5 ms, significantly superior to other methods.

[0115] In the test where the slot duration ranges from 64 ms to 256 ms, as Figure 6 shown, the throughput of MDQN-LCR approaches 6.3×104 Mbps in the 256-ms scenario, which is about 3.2% higher than that of the centralized QMIX. It achieves an optimal latency performance of about 4.5 ms in the shortest slot (64 ms), which is 13.5% lower than that of the greedy method. It demonstrates excellent resource utilization scalability and maintains high efficiency as the available resources increase.

[0116] The test with the number of satellites increasing from 16 to 96 shows that the throughput of MDQN-LCR exceeds that of the centralized QMIX by 2.6% in large-scale deployment (96 satellites). In large-scale scenarios, the latency stabilizes at about 10.2 ms, close to the theoretical minimum. Compared with other methods, it shows better scalability, especially in high-density network configurations.

[0117] In summary, the experimental results confirm that the MDQN-LCR algorithm has excellent performance in various scenarios and is particularly suitable for the hopping beam scheduling application of large-scale LEO satellite constellations. Specifically, it has the following advantages:

[0118] (1) Improve system performance: Through the distributed decision-making and cooperative reward mechanism, the present invention effectively improves the throughput of the LEO satellite network and reduces the communication delay.

[0119] (2) Reduce communication overhead: Compared with the centralized optimization method, the present invention only relies on local information for decision-making, significantly reducing the communication overhead.

[0120] (3) Improve scalability: The distributed algorithm structure makes the present invention more suitable for large-scale LEO satellite constellation scenarios and has good scalability.

[0121] The present invention proposes a hopping beam optimization method for LEO satellite constellations based on distributed multi-agent deep reinforcement learning. In response to the challenges of dynamic traffic demands and uneven user distributions, it achieves a significant improvement in system performance. The main contributions of the present invention include: establishing a multi-objective optimization model that balances throughput and delay; proposing a local interaction Markov game model and proving it to be an exact potential game; developing a multi-agent deep Q-network (MDQN-LCR) algorithm with local cooperative rewards.

[0122] Compared with the prior art, the present invention has significant advantages in the following aspects: effectively reducing communication overhead and computational complexity; improving system throughput and reducing transmission delay; having good scalability and being particularly suitable for large-scale LEO satellite constellation applications. The present invention provides a new idea for solving the resource allocation problem in satellite communication systems and is of great significance for promoting the development of satellite communication technologies and enhancing global connectivity.

[0123] The above specific embodiments do not constitute a limitation to the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A dynamic hopping beam variable optimization method for low-earth orbit satellite constellations based on deep reinforcement learning, characterized in that The optimization method is implemented based on an integrated terrestrial and space network system, which consists of a terrestrial gateway, a low-Earth orbit (LEO) satellite constellation in a Walker formation, and multiple terrestrial cells; the terrestrial gateway transmits data to the cells via LEO satellites; Let the set of LEO satellites be denoted as \(M = \{1, 2, \ldots, m, \ldots, M\}\), and the set of cells be denoted as \(N=\{1, 2, \ldots, n, \ldots, N\}\); the beam set of the \(m\)-th LEO satellite is denoted as \(C\) m \(=\{1, 2, \ldots, c, \ldots, C\}\), and each beam precisely covers a single cell; Let the set of cells served by the m-th LEO satellite be denoted as C m , where and the coverage areas of different LEO satellites do not overlap; therefore, the entire LEO constellation fully covers the set of cells N, specifically: and for m ≠ m' Since the total number of beams of LEO satellites is less than the number of cells, i.e., C·M << N, the LEO satellites implement the hopping beam technology through a time slicing mechanism to effectively provide access services to all covered cells; Set N m Indicates the buffer queue number of the m-th LEO satellite, and each data queue is associated with a predetermined cell; when the ground gateway transmits data to the satellite, if a transmission link is established between the satellite and the target cell, the data is successfully transmitted. If the transmission conditions are not met, the data enters the corresponding queue and waits for transmission in the next time slot; if the data waiting time exceeds the delay threshold T th , the data packet is discarded; The downlink spectrum of the LEO satellite beams adopts a multi-color multiplexing technology, and the total system duration is divided into consecutive time intervals Δt, denoted as T = {1, 2,..., t,..., T}; it is assumed that the LEO satellite positions and the cells illuminated by their beams remain static within each time slot; New data packets from the terrestrial gateway are received by the satellite at the end of each time slot; Set λ n (t) represents the number of data packet arrivals in cell n during time slot t. The data packet size is assumed to be a constant M0, and λ n (t) follows a Poisson distribution; A multi-objective optimization model is constructed to effectively balance the two performance metrics of throughput and latency in the LEO satellite network; the specific optimization model is as follows: Among them, ω1, ω2 ∈ [0, 1] represent the weighted coefficients that trade off between parameterized throughput maximization and latency minimization under the normalization constraint ω1 + ω2 = 1; Θ(t) is the total throughput of all links in time slot t, and Λ(t) is the average delay of all links in time slot t. Constraint C1 establishes a quality of service threshold by stipulating that the signal-to-interference-plus-noise ratio of each cell n in each time slot t must exceed the minimum acceptable threshold γ min to ensure reliable communication connectivity under different interference conditions; Constraint C2 imposes a non-negativity condition on the remaining data packets in the buffer queue to maintain the physical feasibility of the queue system; Constraint C3 sets an upper bound Q on the cumulative queue length of each cell max , prevent buffer overflow, and enforce strict quality of service requirements for delay-sensitive services in the entire satellite network.

2. The method according to claim 1, wherein Based on the multi-objective optimization model, a locally interactive Markov game model is constructed, defined as a tuple G = (M, S, A, N, T, R, γ), where M represents the set of satellites, S represents the joint state space, A represents the joint action space, N represents the neighbor structure, T represents the state transition function, R represents the set of reward functions, and γ represents the discount factor.

3. The method according to claim 2, wherein Based on the locally interactive Markov game model, a multi-agent deep Q-network MDQN-LCR algorithm with local collaborative rewards is constructed to solve the hopping beam scheduling problem in the LEO satellite constellation through a distributed decision-making and collaborative optimization mechanism. The locally interactive Markov game model includes a distributed Q-learning structure and a collaborative reward mechanism; in the distributed Q-learning structure, each satellite agent obtains its own and neighbor satellite information through an abstract state space to achieve state sharing among neighboring satellites; each agent maintains an independent evaluation network to calculate the action value Q(s, a) and selects the optimal action based on the policy π(s|θ); the decision-making process realizes stable learning through a double-network architecture, including a target network with regular soft updates and an evaluation network optimized by policy gradients; multiple satellite agents generate actions in parallel based on environmental observations to form a joint action decision, and the joint action decision is directly mapped to the illumination pattern of K beam positions, affecting the overall throughput and latency performance of the system; the collaborative reward mechanism enables satellites to share local utility information and promotes the global coordination of distributed decision-making; each agent stores interaction data through an independent experience memory and updates network parameters through mini-batch sampling.

4. The method according to claim 3, wherein The MDQN-LCR specifically includes a distributed Q-learning framework and a collaborative reward mechanism; The distributed Q-learning framework enables each satellite to evaluate the action value and make decisions based on local and neighbor observations; The collaborative reward mechanism promotes collaborative behavior among agents through information sharing; At each time slot t, each agent m first observes its local state S from the environment t,m , including information from neighboring satellites (N m ), to form an abstract state space; the agent then uses its evaluation network to compute the Q-value Q m (s, a), and selects an action a through an ε-greedy policy t,m ; after all agents execute the joint action, the environment transitions to the next state, while providing feedback through a collaborative reward mechanism.

Citation Information

Cited By

  • Multi-satellite hopping beam scheduling method for random access

    CN121462066A

  • A multi-satellite hop beam scheduling method for random access

    CN121462066B

  • Communication and electromagnetic map sampling balanced low-orbit satellite hopping beam resource allocation method

    CN122339537A