Dynamic beam hopping method for low-orbit satellite communication system and satellite communication system

Cell division and load balancing are performed through improved P-center algorithm and greedy heuristic algorithm, combined with MADDPG reinforcement learning algorithm to optimize beam hopping strategy, the dynamic balance problem of interference suppression and resource allocation in low-orbit satellite communication systems is solved, and the cell average delay and packet loss rate are reduced and the stability of time delay variance is achieved.

CN120499741APending Publication Date: 2025-08-15SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510531830.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

There are problems in the low-orbit satellite communication system with dynamic balancing problems of interference suppression and resource allocation, complexity of multi-dimensional resource joint optimization, and challenges in dynamic service requirements and load balancing coordination control, resulting in high average cell delay and packet loss rate and large delay variance.

Method used

The improved P-center algorithm is used to divide cells, load balancing is achieved in combination with greedy heuristic algorithms, and multi-satellite joint optimization is used to build a cell-satellite affiliation relationship through multi-objective optimization problems, and beam hopping strategy is optimized.

Benefits of technology

While reducing the average cell delay and packet loss rate, the cell average delay variance is maintained, and it has stability and adaptability to adapt to different traffic load conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499741A_ABST
    Figure CN120499741A_ABST
Patent Text Reader

Abstract

The invention relates to a low-orbit satellite communication system dynamic wave beam hopping method and a satellite communication system, and the method comprises the steps: taking the minimization of the number of cells as a target and the minimum center spacing as a constraint, constructing a cell division model, and solving the cell division model to obtain a cell division result; taking minimization of the maximum load gap between the satellites as a target, taking the cell division result as a constraint, constructing a load balancing model, solving the load balancing model, and determining a cell-satellite ownership relationship; and with user service quality index weighted optimization as a target and the cell-satellite affiliation relationship as a constraint, satellites are used as intelligent agents, and multi-satellite joint optimization is realized by adopting a reinforcement learning algorithm. According to the method, the cell average time delay and the packet loss rate can be reduced, and meanwhile, the low cell average time delay variance can still be kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite communication technology, and in particular to a dynamic beam hopping method for a low-orbit satellite communication system and a satellite communication system. Background Art

[0002] In recent years, low-Earth orbit (LEO) satellite communication networks have played a key role in the evolution of sixth-generation (6G) communication systems, providing crucial support for achieving seamless global communication coverage. To enhance the flexibility of communication resource allocation, beam hopping (BH) technology, a new dynamic resource scheduling solution, has been implemented in existing LEO constellations such as Starlink and OneWeb.

[0003] In existing research, some work has proposed a dynamic beam allocation framework based on deep reinforcement learning (DRL), which maximizes system throughput by adjusting beam pointing in real time while ensuring low latency and latency fairness for real-time services. However, these methods are mainly limited to single-satellite application scenarios. To address the interference suppression challenge in large-scale low-orbit satellite networking scenarios, existing research has introduced power optimization strategies to balance the two competing goals of long-term network throughput and average latency reduction by dynamically adjusting transmit power. Some studies have further considered the inter-satellite load balancing problem and proposed a low-complexity iterative algorithm based on fixed cell division to achieve efficient utilization of network resources by minimizing traffic load differences.

[0004] In summary, the core challenges facing beam-hopping technology research include: (1) The dynamic balance between interference suppression and resource allocation in multi-satellite coordination scenarios: interference cancellation and resource allocation need to be optimized simultaneously in a satellite network with dynamically changing topology. (2) The complexity of multi-dimensional resource joint optimization: The coupling relationship between multi-dimensional resources such as beam-time-frequency-power leads to an explosion in the solution space dimension. (3) Coordinated control of dynamic service demand and load balancing: Time-varying service distribution and inter-satellite load differences place stringent requirements on real-time resource scheduling. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a low-orbit satellite communication system dynamic beam hopping method and a satellite communication system, which can reduce the average cell delay and packet loss rate while still maintaining a low average cell delay variance.

[0006] The technical solution adopted by the present invention to solve the technical problem is: to provide a dynamic beam hopping method for a low-orbit satellite communication system, comprising the following steps:

[0007] With minimizing the number of cells as the goal and the minimum center distance as the constraint, a cell division model is constructed, and the cell division model is solved to obtain a cell division result;

[0008] With the goal of minimizing the maximum load difference between satellites, using the cell division result as a constraint, a load balancing model is constructed, and the load balancing model is solved to determine the cell-satellite affiliation relationship;

[0009] With the weighted optimization of user service quality indicators as the goal and the cell-satellite affiliation as a constraint, the satellites are regarded as intelligent agents and a reinforcement learning algorithm is used to realize multi-satellite joint optimization.

[0010] When solving the cell division model and obtaining the cell division result, an improved P-center algorithm is used to solve the cell division model.

[0011] The improved P-center algorithm is used to solve the cell division model, specifically including:

[0012] Randomly initialize the cell center;

[0013] Calculate the cell belonging of each user;

[0014] The minimum circumscribed circle algorithm is used to update the center position of the cell and cluster users;

[0015] When the cell distance is less than the minimum center distance, merging redundant centers;

[0016] For uncovered users, a new center is added according to the maximum distance criterion, and the calculation of the cell affiliation of each user is repeated and iterated.

[0017] When solving the load balancing model and determining the cell-satellite affiliation relationship, a greedy heuristic algorithm is used to solve the load balancing model.

[0018] The greedy heuristic algorithm is used to solve the load balancing model, specifically including:

[0019] Estimate the total traffic demand of the community in the future based on the historical traffic data of users within the community;

[0020] Sort the cells by their total traffic demands in the future from largest to smallest, and assign the cell with the largest total traffic demand in the future to the satellite with the smallest load based on a greedy strategy;

[0021] Iteratively fine-tune the cell allocation between satellites to narrow the maximum load gap between satellites.

[0022] The user service quality indicators include average delay, average delay variance and packet loss rate.

[0023] The reinforcement learning algorithm is based on the MADDPG framework. The state space of the MADDPG framework includes the average delay, total cumulative flow, and maximum delay of each cell. The state space observed by each intelligent agent is only the average delay, total cumulative flow, and maximum delay of the cell to which the intelligent agent currently belongs. The action space of the MADDPG framework is the beam-hopping decision of the intelligent agent at the current moment. The reward function of the MADDPG framework is the weighted sum of the average delay, average delay variance, and packet loss rate of the cell to which the intelligent agent currently belongs.

[0024] The training process of the reinforcement learning algorithm is as follows:

[0025] An independent action-evaluation network structure is constructed for each agent, in which the action network is used to generate beam-hopping strategies based on the action space, and the evaluation network is used to evaluate the value of actions based on the state space. During the centralized training phase, all agents share the interaction data stored in the experience replay pool and reuse experience through batch sampling. During the policy optimization process, the evaluation network updates parameters by minimizing the temporal difference error, while the action network improves the expected return along the policy gradient.

[0026] The technical solution adopted by the present invention to solve the technical problem is to provide a low-orbit satellite communication system, including:

[0027] Multi-beam phased array antenna module to support dynamic beamforming;

[0028] Distributed queue management module, used to record the status of each cell service queue;

[0029] The processor module is used to execute the above-mentioned low-orbit satellite communication system dynamic beam hopping method to achieve real-time beam scheduling.

[0030] The technical solution adopted by the present invention to solve its technical problem is: providing a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned low-orbit satellite communication system dynamic beam hopping method are implemented.

[0031] Beneficial effects

[0032] Due to the adoption of the above-mentioned technical solution, the present invention has the following advantages and positive effects compared to existing technologies: It constructs a multi-objective optimization problem based on user QoS indicators and decomposes the problem into three sub-problems to be solved sequentially. It reduces the number of cells through a cell partitioning algorithm, achieves inter-satellite load balancing through a greedy heuristic algorithm, and finally implements beam-hopping decisions using the MADDPG algorithm. This method can reduce the average cell delay and packet loss rate while maintaining a low average cell delay variance. It also demonstrates superior performance under varying traffic load conditions, demonstrating stability and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 2 is a schematic diagram of a downlink of a LEO satellite communication system in accordance with a first embodiment of the present invention;

[0034] Figure 2 This is a flow chart of a dynamic beam hopping method for a low-orbit satellite communication system according to a first embodiment of the present invention;

[0035] Figure 3 2 is a schematic diagram of a beam hopping optimization solution based on MADDPG in the first embodiment of the present invention. DETAILED DESCRIPTION

[0036] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0037] The first embodiment of the present invention relates to a dynamic beam hopping method for a low-orbit satellite communication system, which can be applied to Figure 1 The downlink architecture of a LEO satellite communication system is shown in Figure 1. Various users (mobile phones, vehicles, etc.) are distributed within the satellite coverage area. The downlink consists of four components: a ground network control center (CNC), a gateway, satellites, and ground users. The CNC divides users into independent cells based on their geographic location. It also determines the cell-satellite affiliation based on each cell's historical traffic information and satellite-cell visibility. The gateway receives user traffic data from the cloud and, based on both the user-cell affiliation and the cell-satellite affiliation, packages the data and sends it to the corresponding satellite. The satellite then transmits the data to the ground cells in a time-division manner through beam hopping. Finally, the ground users receive the traffic data.

[0038] The dynamic beam hopping method for the low-orbit satellite communication system of this embodiment divides the service quality optimization problem of the LEO satellite communication system into three sub-problems: cell division, load balancing, and beam hopping strategy, which correspond to the three stages of user-cell affiliation determination, cell-satellite affiliation determination, and beam hopping in downlink traffic transmission, respectively. Figure 2 As shown, specifically including:

[0039] Step 1: With minimizing the number of cells as the goal and the minimum center distance as the constraint, a cell division model is constructed, and the cell division model is solved to obtain a cell division result.

[0040] This step addresses the cell division subproblem. This implementation aims to minimize the number of cells while adding a minimum distance constraint. This constructs a cell division model and proposes an improved P-center algorithm to solve the cell division model. The details are as follows:

[0041] The process randomly initializes cell centers; calculates each user's cell affiliation; updates cell center locations using the minimum circumscribed circle algorithm and clusters users; merges redundant centers when cell spacing falls below the minimum center-to-center spacing; adds centers for uncovered users based on the maximum distance criterion, and repeats the process of calculating each user's cell affiliation. This cell division method minimizes the number of cells while ensuring full user coverage.

[0042] Step 2: With the goal of minimizing the maximum load difference between satellites and the cell division result as a constraint, a load balancing model is constructed, and the load balancing model is solved to determine the cell-satellite affiliation relationship.

[0043] This step focuses on the load balancing subproblem. With the goal of minimizing the maximum load difference between satellites, a load balancing model is constructed and solved based on a greedy heuristic algorithm. The details are as follows:

[0044] The method uses historical traffic data from users within a cell to estimate the total future traffic demand of the cell. The cells are ranked from highest to lowest in terms of total future traffic demand, and based on a greedy strategy, the cell with the highest total future traffic demand is assigned to the satellite with the lowest load. The inter-satellite cell allocation is then iteratively fine-tuned to narrow the maximum load gap between satellites. This method allows cells to be assigned to the satellite with the lowest load in descending order of traffic demand, optimizing traffic distribution between satellites and achieving inter-satellite load balancing.

[0045] Step 3: With the goal of achieving the optimal weighted user quality of service (QoS) indicator and the cell-satellite affiliation as a constraint, a reinforcement learning algorithm is used to realize multi-satellite joint optimization by treating satellites as intelligent agents.

[0046] This step addresses the beam-hopping strategy subproblem. Aiming to optimize the weighted QoS metrics, a multi-satellite coordinated beam-hopping decision algorithm based on MADDPG is proposed. The QoS metrics in this step include average delay, average delay variance, and packet loss rate.

[0047] The average delay can be calculated as follows:

[0048]

[0049] in, Indicates cell c j The average delay at time t is Indicates that at time t, cell c j In satellites i The traffic that has waited for m time slots, where M is the total number of time slots.

[0050] The packet loss rate can be calculated as follows:

[0051]

[0052] in, Indicates cell c j The packet loss rate at time t, Represents cell c at time t+Δt j In satellites i The traffic that has been waiting for M time slots, is cell c at time t j communication rate.

[0053] The average delay variance can be calculated as follows:

[0054]

[0055] in, Indicates time t, satellite s i The average delay variance between served cells, Satellites i The set of cells served, It represents the average of the average delays of all cells at time t.

[0056] like Figure 3 As shown, the state space, action space, reward function, and training process of the MADDPG framework in this implementation are as follows:

[0057] State Space: To achieve optimal weighting of QoS metrics while avoiding degradation of algorithm convergence due to an excessively large state space, this implementation introduces three components to the state space: average latency across all cells, total cumulative traffic, and maximum latency. For each agent (i.e., satellite), the observed state space is limited to the average latency, total cumulative traffic, and maximum latency of the cell to which it currently belongs.

[0058] Action space: A satellite can generate up to K beams simultaneously. The action space is the satellite's beam-hopping decision at the current moment, that is, the cell number illuminated by the satellite's beam selection. It can use Gumbel-Softmax to implement differentiable sampling of discrete beam selection, representing the satellite's beam hopping in each time slot.

[0059] Reward function: The reward consists of the average delay, data packet loss rate, and average delay variance of the cell at the current moment, which correspond to three QoS indicators respectively. The reward function is the weighted sum of the average delay, average delay variance, and packet loss rate of the cell to which the agent currently belongs.

[0060] Training Process: First, a separate actor-critic network is constructed for each satellite. The actor network generates beam-hopping strategies based on the action space, while the critic network evaluates the value of actions based on global state information in the state space. During the centralized training phase, all satellites share interaction data (state, action, reward, and next state) stored in the experience replay pool, enabling experience reuse through batch sampling. During the policy optimization process, the evaluation network updates its parameters by minimizing the temporal difference error (TDE). This means that the evaluation network of each agent concatenates the observations of all agents, including itself, into an observation vector and the actions of all agents into an action vector. These two vectors serve as the input to the online evaluation network, outputting a one-dimensional Q-value. The temporal difference error (TDE) between the current Q-value and the target Q-value is then calculated using samples from the experience replay buffer. A mean squared error loss function is then constructed, and the parameters of the evaluation network are updated using gradient descent to minimize this loss. The action network improves the expected return along the policy gradient. During training, the action network of each agent selects actions based on its own observations, while the evaluation network evaluates the strategy based on the observations and actions of all agents. To improve the expected return, the policy gradient is calculated: the gradient of the Q-value output by the evaluation network with respect to its own action, multiplied by the gradient of the action output by the action network with respect to its parameters. This results in the policy gradient direction, which is then used to update the parameters of the action network, adjusting the agent's action strategy in a direction that increases the expected return. To enhance training stability, this implementation also introduces a delayed target network update mechanism and adds strategic noise to facilitate exploration. Ultimately, through a distributed execution architecture, each satellite relies solely on its own local observations to achieve decentralized beam-hopping decisions.

[0061] It's easy to see that this method constructs a multi-objective optimization problem based on user QoS metrics and decomposes it into three sub-problems, each solved sequentially. It uses a cell partitioning algorithm to reduce the number of cells, a greedy heuristic algorithm to achieve inter-satellite load balancing, and finally, a MADDPG algorithm to implement beam-hopping decisions. This method can reduce the average cell latency and packet loss rate while maintaining a low average cell latency variance. It also demonstrates superior performance under varying traffic load conditions, demonstrating stability and adaptability.

[0062] A second embodiment of the present invention relates to a low-orbit satellite communication system, comprising:

[0063] Multi-beam phased array antenna module to support dynamic beamforming;

[0064] Distributed queue management module, used to record the status of each cell service queue;

[0065] The processor module is used to execute the dynamic beam hopping method of the low-orbit satellite communication system of the first embodiment to achieve real-time beam scheduling.

[0066] In this embodiment, the processor module includes:

[0067] The central network controller part is used to perform cell division and cell-satellite affiliation calculation;

[0068] The MADDPG decision engine is used to implement real-time beam scheduling.

[0069] A third embodiment of the present invention relates to a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the steps of the dynamic beam hopping method for a low-orbit satellite communication system of the first embodiment.

[0070] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) that contain computer-usable program code.

[0071] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0072] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction method, which is implemented in the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0074] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A dynamic beam hopping method for a low-orbit satellite communication system, characterized in that: The following steps are involved: With minimizing the number of cells as the goal and the minimum center distance as the constraint, a cell division model is constructed, and the cell division model is solved to obtain a cell division result; With the goal of minimizing the maximum load difference between satellites, the cell division result is used as a constraint to build a load balancing model. and solving the load balancing model to determine the cell-satellite ownership relationship; With the weighted optimization of user service quality indicators as the goal and the cell-satellite affiliation as a constraint, the satellites are regarded as intelligent agents and a reinforcement learning algorithm is used to realize multi-satellite joint optimization.

2. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 1, wherein: When solving the cell division model and obtaining the cell division result, an improved P-center algorithm is used to solve the cell division model.

3. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 2, wherein: The improved P-center algorithm is used to solve the cell division model, specifically including: Randomly initialize the cell center; Calculate the cell belonging of each user; The minimum circumscribed circle algorithm is used to update the center position of the cell and cluster users; When the cell distance is less than the minimum center distance, merging redundant centers; For uncovered users, a new center is added according to the maximum distance criterion, and the calculation of the cell affiliation of each user is repeated and iterated.

4. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 1, wherein: When solving the load balancing model and determining the cell-satellite affiliation relationship, a greedy heuristic algorithm is used to solve the load balancing model.

5. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 4, characterized in that: The greedy heuristic algorithm is used to solve the load balancing model, specifically including: Estimate the total traffic demand of the community in the future based on the historical traffic data of users within the community; Sort the cells by their total traffic demands in the future from largest to smallest, and assign the cell with the largest total traffic demand in the future to the satellite with the smallest load based on a greedy strategy; Iteratively fine-tune the cell allocation between satellites to narrow the maximum load gap between satellites.

6. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 1, wherein: The user service quality indicators include average delay, average delay variance and packet loss rate.

7. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 6, wherein: The reinforcement learning algorithm is based on the MADDPG framework. The state space of the MADDPG framework includes the average delay, total cumulative flow, and maximum delay of each cell. The state space observed by each intelligent agent is only the average delay, total cumulative flow, and maximum delay of the cell to which the intelligent agent currently belongs. The action space of the MADDPG framework is the beam-hopping decision of the intelligent agent at the current moment. The reward function of the MADDPG framework is the weighted sum of the average delay, average delay variance, and packet loss rate of the cell to which the intelligent agent currently belongs.

8. The dynamic beam hopping method for a low-orbit satellite communication system according to claim 7, wherein: The training process of the reinforcement learning algorithm is as follows: An independent action-evaluation network structure is constructed for each agent, in which the action network is used to generate beam-hopping strategies based on the action space, and the evaluation network is used to evaluate the value of actions based on the state space. During the centralized training phase, all agents share the interaction data stored in the experience replay pool and reuse experience through batch sampling. During the policy optimization process, the evaluation network updates parameters by minimizing the temporal difference error, while the action network improves the expected return along the policy gradient.

9. A low-orbit satellite communication system, characterized in that: include: Multi-beam phased array antenna module to support dynamic beamforming; Distributed queue management module, used to record the status of each cell service queue; A processor module, configured to execute the dynamic beam hopping method for a low-orbit satellite communication system as described in any one of claims 1-8 to achieve real-time beam scheduling.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the dynamic beam hopping method for a low-orbit satellite communication system as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Multi-satellite hopping beam scheduling method for random access

    CN121462066A