Multi-beam cooperative scheduling method and system for low-orbit satellite communication

By adopting adaptive clustering and deep reinforcement learning methods in multi-beam low-orbit satellite communication, the efficiency of beam scheduling and resource allocation in high dynamic environments is solved, and more efficient resource utilization and system performance improvement is achieved.

CN120091415APending Publication Date: 2025-06-03XIDIAN UNIV

Patent Information

Application Number
CN202510203922.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In multi-beam low-orbit satellite communication scenarios, it is difficult for the prior art to effectively schedule beam and resource allocation in a high dynamic environment, resulting in reduced resource utilization efficiency and unstable system performance.

Method used

By splitting large-scale global optimization problems into smaller sub-problems, combining wide beam and beam hopping technology, the beam hopping bit is adopted to design the beam hopping bits, and using deep reinforcement learning for beam hopping scheduling, the beam hopping time frame structure is optimized to improve resource allocation efficiency.

Benefits of technology

It improves the flexibility of beam scheduling and resource allocation, enhances the adaptability of scheduling strategies to time-varying needs, continuously improves decision-making quality, and improves resource utilization efficiency and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091415A_ABST
    Figure CN120091415A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-beam cooperative scheduling method and a multi-beam cooperative scheduling system for low-orbit satellite communication, which mainly solve the problems of poor beam scheduling flexibility and difficulty in adapting to time-varying requirements in a high-dynamic environment in the prior art, and the implementation scheme of the multi-beam cooperative scheduling method comprises the following steps of: collecting online user information by a low-orbit satellite through a wide beam; a hopping beam position is generated by using adaptive clustering under the constraint of channel resources; a hopping beam scheduling algorithm based on deep reinforcement learning outputs an optimal scheduling strategy according to the real-time environment variable; and optimizing a beam hopping time frame structure in each scheduling period, covering accessed users in a beam position through spot beams, and sending service data and control information to complete efficient service transmission. The method can adapt to changes of user distribution and service requirements in low-orbit satellite communication, the utilization rate of spectrum resources is improved, communication delay is reduced, the overall performance and service quality of the system in a high-dynamic environment can be enhanced, and the method can be used for dynamic beam position division, hopping beam scheduling and hopping beam resource allocation in a multi-beam low-orbit satellite communication scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and further relates to a method and system for multi-beam cooperative scheduling, which can be used for dynamic beam position division, hopping beam scheduling, and hopping beam resource allocation in multi-beam low-earth orbit satellite communication scenarios. Background Art

[0002] With the rapid growth of global broadband access demand, traditional large-beam satellites are difficult to meet the evolving service requirements. The user distribution and traffic volume exhibit highly time-varying and spatially heterogeneous characteristics. Traditional satellite coverage methods, such as single large-beam coverage, cannot achieve flexible acceleration of hotspots and are also difficult to dynamically adjust according to temporary peak demands. For this reason, multi-beam satellite systems have emerged. Multi-beam satellite systems divide the entire coverage area into many smaller service units by using high-gain, narrow beams, and achieve high-frequency spectrum reuse and dynamic resource allocation between these units. Such systems can obtain higher system capacity under limited spectrum and can maximize resource utilization and quality of service QoS through flexible scheduling and power allocation strategies between beams. However, with the increase in system complexity, how to efficiently schedule multi-beam resources and perform coverage planning has become a key challenge.

[0003] The patent document with the application number CN202310363998.0 discloses a "multi-beam satellite communication resource allocation system and method based on deep reinforcement learning", which uses deep reinforcement learning to allocate multi-beam satellite communication resources, including satellite communication devices, ground communication modules, satellite available resource discriminators, and user service performance discriminator devices. The satellite communication devices are set in a preset satellite communication area for receiving data to be communicated; the ground communication module is used for receiving the communication data transmitted by the satellite devices. The resource allocation is optimized through a deep reinforcement learning algorithm; during the resource allocation process, it continuously self-learns and iteratively updates to optimize the resource allocation result; according to the current online user information, channel historical allocation situation, etc., without exceeding the satellite transmission power limit, the current available channel resources are counted, and the allocation feedback is obtained based on the Shannon formula. Although this method can improve the resource allocation performance and stability, in the face of large-scale communication systems and complex dynamic environments, the action space of resource allocation will increase sharply, and the high-dimensional action space will lead to a slow convergence speed of the algorithm, making it difficult to learn the optimal strategy within a limited time.

[0004] Peng Mingyang, Zhang Chen et al. from Nanjing University of Posts and Telecommunications proposed a hopping beam resource scheduling strategy applicable to the low-earth orbit (LEO) constellation scenario in the paper "Hopping Beam Resource Scheduling Strategy for LEO Constellations". This strategy formulates two hopping beam resource scheduling strategies based on iterative algorithms and convex optimization for two cases where the user service is either a priori unknown or a priori known, considering the co-channel interference between beams. Compared with traditional strategies, this strategy can significantly improve the system throughput, has good delay performance, and can adapt to the uneven distribution of user services. Although this strategy considers the two cases where the user service is either a priori unknown or a priori known, in practical applications, due to the dynamic changes in user requirements, the adaptability of the strategy may be insufficient. Moreover, the rapid movement of LEO satellites will cause their coverage areas and user requirements to change continuously, increasing the complexity of beam switching and inter-satellite switching, resulting in the inability to adjust resource allocation in a timely manner when dealing with a rapidly changing dynamic environment, and causing a decline in resource utilization efficiency.

[0005] Li Hongguang, Shi Jinglin et al. from the Institute of Computing Technology, Chinese Academy of Sciences gave a multi-beam scheduling strategy for LEO satellites based on delay and genetic algorithms in the paper "Multi-beam Scheduling Strategy for LEO Satellites Based on DWGA". They designed a multi-beam scheduling architecture for LEO satellites, analyzed the impact of co-frequency reuse distance on the carrier-to-interference-plus-noise ratio (CINR) through the constructed beam interference model, and determined the interference avoidance scheme for full frequency reuse. Then, a service model was constructed using the real population density distribution and user scheduling rate. The service capacity to be sent to different beam cells was weighted with the service tolerance delay to obtain the capacity scheduling factor to avoid the timeout of service data. Finally, with the goal of maximizing satellite capacity, the genetic algorithm was used to solve the hopping beam pattern under different scheduling orders. However, due to the high dynamicity of LEO satellites, their coverage areas and user requirements change continuously, making this method may need to be frequently re-optimized in a dynamic environment, increasing the complexity and computational burden of the algorithm. Moreover, in the face of bursty traffic demands, it may not be able to quickly adjust the resource allocation strategy, resulting in a decline in system performance. Summary of the Invention

[0006] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies, and propose a multi-beam cooperative scheduling method and system for LEO satellite communication, so as to enhance the flexibility of beam scheduling and resource allocation in a high-dynamic environment, improve the adaptability of the scheduling strategy to time-varying demands, enable it to continuously improve the decision-making quality in a time-varying environment, and further improve the resource utilization efficiency.

[0007] The key technology to achieve the object of the present invention is: by splitting a large-scale global optimization problem into several smaller sub-problems, in the multi-beam low-Earth orbit satellite communication scenario, relying on the existing wide-beam and hopping-beam technology bases and the transmission and control separation mechanism, through adaptive clustering, to achieve dynamic hopping-beam position design; by using deep reinforcement learning to achieve the scheduling of hopping beams, and by optimizing the frame structure at fixed positions, using spot beams to cover online users during service time slots.

[0008] The implementation steps include:

[0009] 1. A multi-beam cooperative scheduling method for low-Earth orbit satellite communication, characterized by including:

[0010] The low-Earth orbit satellite collects online user information through a wide beam, and uses this information to adaptively cluster and generate hopping-beam positions under channel resource constraints;

[0011] The hopping-beam baseband monitors the change of wide-beam information in real time, reads the position information as an environmental variable, and uses this environmental variable to perform periodic scheduling of the hopping beam based on deep reinforcement learning;

[0012] In each scheduling period, optimize the hopping-beam frame structure, use the online user information provided by the wide beam to allocate service time slot resources for users, and use spot beams to cover users within the position, completing the beam cooperative scheduling of low-Earth orbit satellite communication.

[0013] Further, the low-Earth orbit satellite collects online user information through a wide beam, including:

[0014] In the initial stage, the low-Earth orbit satellite adopts a wide-beam coverage strategy for full-area coverage. Within the coverage area, user terminals attempt to access the system under the specified random access mechanism; during the access process, the user terminal reports its own location information, channel resource requirements, and current service cache situation to the low-Earth orbit satellite through an access request; the low-Earth orbit satellite responds to the access request of the user terminal, allocates control channel resources for the user according to the current in-satellite resource situation, and records and statistics the information reported by the user, completing the collection of online user information.

[0015] Further, the adaptive clustering to generate hopping-beam positions using online user information under channel resource constraints includes: mapping the wide-beam coverage range of the low-Earth orbit satellite to a coordinate plane, mapping the reported geographical location information to the coordinate plane when the user accesses, and using the horizontal and vertical coordinates as user identifiers; when clustering, on the basis of giving priority to position constraints, generate the single-clustering effect, and then verify whether the total resource requirements of all users within each position meet the total channel resource limit:

[0016] Further, the periodic scheduling of the hopping beam based on deep reinforcement learning includes:

[0017] (a) The online policy network and the target network are constructed using the same neural network structure. The state vector from the environment is input into the online policy network, and the Q-values corresponding to each possible action are output. Among them, the state vector includes buffer traffic volume, transmission rate, and wave position; the actions include wave position selection and coverage time selection; the Q-value is used to represent the expected future cumulative reward that can be obtained by taking different actions in the current state, and the reward is the single-step reward that can be obtained when performing the action.

[0018] (b) The experience tuple composed of the state, action, reward, and next state is stored in a pre-set experience replay buffer. The experience replay mechanism is used to randomly sample a small batch of data from the buffer regularly for subsequent network parameter updates to break the correlation between the data.

[0019] (c) Using the Q-values output by the current online policy network, the temperature annealing Softmax strategy is used to select actions, where the temperature parameter value is dynamically adjusted according to the number of training steps.

[0020] (d) In the target network, the target Q'-value is obtained by estimating the optimal reward that may be obtained from the reward of the current online policy network and the next state. The mean square error is used as the loss function to calculate the difference between the output Q-value of the current online policy network and the target Q'-value, and the parameters of the online policy network are updated through the backpropagation algorithm.

[0021] (e) Repeat the above steps (b) to (d) until the reward of the online policy network converges, and an approximate optimal state-action value function is obtained, which represents the optimal expected value of the future cumulative discounted reward that can be obtained by taking a certain action in a given state and continuing to act according to a specific policy.

[0022] (f) The hopping beam is scheduled according to the optimal state-action value function, that is, at each scheduling moment, the wave position and coverage duration selected for coverage can both obtain the optimal reward.

[0023] 2. A multi-beam cooperative scheduling system for low-earth orbit satellite communication, characterized in that it includes:

[0024] A random access module for random access of users within the wide beam coverage area and collection of user information;

[0025] An adaptive clustering wave position design module for generating hopping beam wave positions through adaptive clustering;

[0026] A hopping beam scheduling module based on deep reinforcement learning for agent training and loading the network parameters after training to perform hopping beam scheduling;

[0027] A hopping beam time frame optimization module for designing the hopping beam time frame structure and optimizing resource allocation.

[0028] The spot beam service transmission module is used for the spot beam to complete multi-user service transmission within the wave position.

[0029] Furthermore, the adaptive clustering wave position design module includes:

[0030] The forward mapping sub-module is used to map the actual geographical location information of the user into plane coordinates;

[0031] The mobility detection sub-module is used to detect the position change and online status of the user;

[0032] The clustering planning sub-module is used for wave position planning of the current user set;

[0033] The reverse mapping sub-module is used to reverse map the plane coordinate information of the wave position design result into actual geographical location information; the wave position visualization sub-module is used for visual display of the wave position coverage scheme.

[0034] Furthermore, the hopping beam scheduling module based on deep reinforcement learning includes:

[0035] The environment simulation sub-module is used to simulate the task state of the hopping beam and its dynamic evolution;

[0036] The experience replay sub-module is used to store the state transition data collected by the agent in the environment to alleviate the correlation between samples;

[0037] The policy network sub-module is used for neural network approximation to help the agent estimate the value of each action from the environmental state;

[0038] The target network sub-module is used to generate stable target values to assist in optimizing the training process;

[0039] The model storage and loading sub-module is used to save the trained model parameters and can load the model for inference. Description of the Drawings

[0040] Figure 1 is the implementation flow chart of the multi-beam cooperative scheduling method for low-orbit satellite communication of the present invention;

[0041] Figure 2 is the sub-flow chart of the adaptive clustering wave position design under the wide beam in the method of the present invention;

[0042] Figure 3 is the sub-flow chart of the hopping beam scheduling algorithm based on deep reinforcement learning in the method of the present invention;

[0043] Figure 4 is the schematic diagram of the hopping beam time frame structure in the method of the present invention;

[0044] Figure 5It is the structural block diagram of the beam collaborative scheduling system for low-earth orbit satellite communication of the present invention;

[0045] Figure 6 It is the structural diagram of the adaptive clustering wave position design module in the system of the present invention;

[0046] Figure 7 It is the structural diagram of the hopping beam scheduling module in the system of the present invention; Detailed implementation manners

[0047] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] It should be noted that the step numbers in the specification and claims of the present invention are only for clearly describing the implementation solutions of the present invention for easy understanding, and their sequence numbers are not limited.

[0049] The implementation scenarios of the present invention cover the collaborative work of the space segment and the ground segment. The space segment mainly consists of low-earth orbit satellites, which have the functions of wide beams and hopping beams. The wide beams are used to achieve large-scale signal coverage to ensure that user terminals within a wide area can receive stable communication signals; while the hopping beams provide more concentrated signal coverage for specific regions or user groups to improve communication quality and efficiency. The ground segment includes user terminals, which can be mobile devices such as mobile phones and tablet computers, and they establish communication connections with low-earth orbit satellites through ground stations.

[0050] Refer to Figure 1 , the implementation steps of this example are as follows:

[0051] Step 1, collect online user information. (1.1) The low-earth orbit satellite adopts a wide beam coverage strategy for full-area coverage, and the satellite broadcasts its own ephemeris information to all users under the wide beam in real time, where the ephemeris information provides basic parameters of the satellite to the users;

[0052] (1.2) User terminals within the coverage area attempt to access the system under the specified random access mechanism, and at the same time send user information to the satellite, including geographical location, resource requirements, service cache information, and user priority information, where:

[0053] The geographical location information refers to the longitude and latitude obtained through the positioning system;

[0054] The resource requirement refers to the amount of resources requested by the user according to its own priority;

[0055] The service cache information refers to the uplink service cache volume of the user;

[0056] The user priority is generated by the user himself and serves as the basic attribute of the user;

[0057] (1.3) The satellite responds to the user's access request, detects the user's legitimacy, distributes the accessed users to different hopping beam basebands according to the location information of the user, and through the access response, sends information such as the beam frequency, power, modulation method, etc. assigned to the user, and counts the user information of the legally accessed users for subsequent wave position planning;

[0058] Step 2, generate the hopping beam wave positions according to the collected information.

[0059] Refer to Figure 2 , the specific implementation of this step includes the following:

[0060] (2.1) Read the statistically collected user geographical location information for forward coordinate mapping, that is, map the wide beam coverage range of the low-earth orbit satellite to the two-dimensional region Ω = [0, L] × [0, L] of the coordinate plane; map the user longitude and latitude information to the horizontal and vertical coordinates X i =(x i , y i ) ∈ Ω, where i ∈ N and N represents the number of accessed users;

[0061] (2.2) Read the user resource request records, and record the control channel resource request volume r i of the i-th user in the current wave position. The set of resource requests of all accessed users is R = {r 1 , r 1 , …, r N};

[0062] (2.3) Set the channel resource upper limit, requiring the upper limit of the channel resources of a single beam to be Q s , then the total resource request within each wave position does not exceed the upper limit, that is:

[0063] (2.4) Use the K-means clustering algorithm to cluster and divide the user coordinate points in the plane:

[0064] (2.4.1) Initialize the minimum number of clusters K = K min ;

[0065] (2.4.2) Initially select K user coordinate points as the cluster centers O = {o 1 , o 2 , …, o K}, where o i represents

[0066] The central point coordinates of the i-th category;

[0067] (2.4.3) Calculate the Euclidean distance from each user to each cluster center, and assign each user point X i to the cluster center with the smallest Euclidean distance, that is

[0068]

[0069] (2.4.4) After the user assignment is completed, recalculate the cluster center c of each classification in the clustering result j :

[0070]

[0071] (2.4.5) Determine whether the updated cluster center has changed:

[0072] If there is no change, obtain the clustering result of this time and execute step (2.4.6);

[0073] Otherwise, return to step (2.4.3);

[0074] (2.4.6) Check all classifications in the clustering result to see if they meet the resource constraint conditions

[0075] If all clustering results meet the resource constraints, the clustering is completed and step (2.5) is executed;

[0076] Otherwise, increase the number of clusters, that is, let K = K + 1, and return to step (2.4.3);

[0077] (2.5) Output all clustering results, use the coordinate values of each cluster center as the central point coordinates of each wave position, and define the wave position coverage radius λ j as the maximum value of the distances from all user points in the cluster to the cluster center:

[0078]

[0079] Among them, j is the classification identifier, and X j represents the set of user coordinates in the j-th category, and o j represents the cluster center of the j-th category.

[0080] (2.6) Output the wave position design result M including the wave position radius, central coordinates, and wave position user set:

[0081]

[0082] Among them, K is the number of wave positions finally output, i ∈ [1, K], represents the abscissa of the central point of wave position i and the ordinate λ i represents the radius of wave position i, X i ' represents the set of user coordinates included in wave position i;

[0083] (2.7) Reverse-map the wave position design result into actual geographical location information for visual output, and while displaying the wave position design result, store the wave position information for use by the hopping beam scheduling module.

[0084] Step 3, perform periodic scheduling on the hopping beam based on deep reinforcement learning.

[0085] The hopping beam scheduling can be implemented through fixed wave position polling, greedy algorithms, heuristic algorithms, and intelligent machine learning methods. The present invention uses, but is not limited to, the DQN algorithm in intelligent machine learning methods to implement the hopping beam scheduling.

[0086] Refer to Figure 3 , and the specific implementation of this step includes the following:

[0087] (3.1) Obtain environmental parameters, which include the information of each wave position, the buffer traffic volume of each wave position, and the communication rate of each wave position;

[0088] (3.2) Construct a Markov decision process model:

[0089] (3.2.1) Define the state space S t :

[0090] S t ={D t ,C t ,L t}={{d 1 ,…,d i ,…,d N},{c 1 ,…,c i ,…,c N},{l 1 ,…,l i ,…,l N}}

[0091] Where:

[0092] D t ={d 1 ,…,d i ,…,d N} represents the remaining traffic volume of each wave position;

[0093] C t ={c 1 ,…,c i ,…,cN} represents the current transmission rate of each wave position;

[0094] L t = {l 1 , …, l i , …, l N} represents the current service delay time of each wave position;

[0095] d i represents the remaining traffic volume of wave position i;

[0096] c i represents the current transmission rate of wave position i;

[0097] l i represents the service delay time of wave position i;

[0098] N represents the number of wave positions;

[0099] (3.2.2) Define the action space:

[0100] Among them, represents the wave position selected at the current moment; represents the coverage time selected at the current moment;

[0101] (3.2.3) When the current state s t is given, and the agent takes an action a t , the probability that the system transfers to the next state s t+1 is defined as the state transition probability: P(s t+1 |(s t , a t ));

[0102] (3.2.4) Define the reward function r t as the immediate reward obtained after each execution of an action. It is the negative value of the weighted delay of the entire network of the system, expressed as:

[0103]

[0104] Among them, l i is the delay record of wave position i, and d i is the remaining traffic volume of wave position i;

[0105] (3.3) After the Markov decision model is constructed, the ε-greedy strategy is used as the balance method for exploration and exploitation in the DQN algorithm. The parameter ε is used to control the probability of the agent choosing "exploration" or "exploitation" at each step. That is, with probability ε, the agent randomly selects an action for exploration; with probability 1 - ε, the agent selects the action with the largest Q value for exploitation;

[0106] (3.4) The agent selects an action a based on the current state s t and undergoes a state transition to obtain the next state s t and the current action reward r t+1 ; t

[0107] (3.5) The quadruple (s t , a t , r t , s t+1 ) is stored as an experience tuple in the experience replay buffer. The agent uses the experience replay mechanism to randomly sample a small batch of experience tuples from the experience replay buffer for training the online policy network;

[0108] (3.6) For each sampled experience tuple (s t , a t , r t , s t+1 ), the target value y t

[0109] y t = r t + γ max a' Q(s t+1 , a'; θ - )

[0110] where r t represents the immediate reward obtained after executing the current action; γ represents the discount factor, which is used to balance the importance of immediate rewards and future rewards; s t+1 represents the next state the agent enters after executing the action; a' represents the action that may be selected in the s t+1 state; max a' Q(s t+1 , a', θ - ) represents the maximum expected return obtained for all possible actions in the next state, and θ - are the parameters of the target network;

[0111] (3.7) Calculate the predicted value Q(s t , a t , θ) using the parameters θ of the online policy network, and then use the mean squared error MSE to calculate the loss function L(θ)

[0112]

[0113] where s t represents the current state, a t represents the action taken in the current state, r t+1 ​$r_{t + 1}$ represents the immediate reward obtained at step $t + 1$, and $\gamma$ represents the discount factor;

[0114] (3.8) Calculate the gradient according to the loss function, and use the gradient descent algorithm to update the parameters of the online network. Every certain number of steps, copy the parameters of the current network to the target network and update the parameters of the target network;

[0115] (3.9) Iteratively execute steps (3.3) to (3.8) until the optimal state - action value function is reached, and save the parameters of the online policy network and the target network at this time;

[0116] (3.10) In the hopping beam scheduling, after loading the saved network parameters into the model, directly use the loaded model to perform inference and analysis on the current scheduling requirements, and select the action with the highest value for hopping beam scheduling.

[0117] Step 4: Optimize the hopping beam time frame structure within each scheduling period, and use the spot beam to cover the users within the wave position.

[0118] Refer to Figure 4 , and the specific implementation of this step is as follows:

[0119] (4.1) Divide the entire hopping beam scheduling period into several sub - frames, and each sub - frame represents the single - time coverage time of the hopping beam.

[0120] (4.2) Further divide each sub - frame into a guard interval, a synchronization sub - frame, a control sub - frame, an optimization interval, and a service sub - frame;

[0121] (4.3) The guard interval provides blank time between adjacent sub - frames, ensures that the frame order does not conflict, and provides time for the system to adjust, so as to effectively avoid frame conflicts caused by signal interference or processing delay, thereby ensuring the accuracy and reliability of data transmission;

[0122] (4.4) The synchronization sub - frame establishes time synchronization between the satellite and the users, ensuring that there is a unified clock constraint among the communication nodes of the communication system, so as to avoid data transmission errors and communication conflicts caused by inconsistent time;

[0123] (4.5) After time synchronization is completed, divide the control sub - frame into several control time slots according to a fixed time slot length, and allocate different control time slots to users to transmit control information, ensuring that each user can receive control information within a specific time slot;

[0124] (4.6) The satellite obtains the current service transmission status through the optimization interval, provides environmental feedback for the scheduling algorithm, monitors and evaluates the status of service transmission in real - time, and the scheduling algorithm makes dynamic adjustments according to the current transmission status to adapt to the changing communication requirements;

[0125] (4.7) Divide the service sub-frame into several service time slots according to a fixed time slot length. The service time slot is the most basic resource unit for service transmission. Based on time synchronization, users perform service sending and receiving in different time slots respectively, effectively avoiding communication conflicts and improving the utilization rate and transmission efficiency of service resources;

[0126] (4.8) According to the optimized hopping beam time frame structure completed in (4.1) to (4.7), the satellite allocates service time slot resources for users within the beam position according to the needs and priorities of the users, and sends control information to the users through hopping beam in the control sub-frame, including beam position information, total amount of service time slot resources, and service resource allocation information;

[0127] (4.9) After receiving the control information, the user configures the service time slot resources. The satellite reduces the hopping beam angle to form a spot beam, and covers the corresponding user terminal in the allocated service time slot to complete service transmission.

[0128] Refer to Figure 5 , the multi-beam cooperative scheduling system for low-earth orbit satellite communication provided in this example includes a random access module 1, an adaptive clustering beam position design module 2, a hopping beam scheduling module 3 based on deep reinforcement learning, a hopping beam time frame optimization module 4, and a spot beam service transmission module 5. Among them:

[0129] The adaptive clustering beam position design module 2 includes: a forward mapping sub-module 21, a mobility detection sub-module 22, a clustering planning sub-module 23, a reverse mapping sub-module 24, and a beam position visualization sub-module 25, as Figure 6 shown.

[0130] The hopping beam scheduling module 3 based on deep reinforcement learning includes: an environment simulation sub-module 31, a policy network sub-module 32, an experience replay sub-module 33, a target network sub-module 34, and a model storage and loading sub-module 35, as Figure 7 shown.

[0131] The working principle of the entire system is as follows:

[0132] The random access module 1 performs random access of users and collection of user information within the wide beam coverage area, and stores the user information in the satellite; in the adaptive clustering beam position design module 2, the forward mapping sub-module 21 reads the accessed user information, maps it into plane coordinates and outputs it to the mobility detection sub-module 22 to detect the change of the user's plane coordinates, and outputs an update instruction to the clustering planning sub-module 23; after receiving the update instruction, the clustering planning sub-module 23 divides the beam positions through adaptive clustering, and outputs the clustering result to the reverse mapping sub-module 24; the reverse mapping sub-module 24 maps the clustering result to the spatial geographical location information, and outputs the actual beam position information to the beam position visualization sub-module 25; the beam position visualization sub-module 25 visually outputs the mapped beam position information and saves it. The hopping beam scheduling module 3 based on deep reinforcement learning conducts agent training and loads the network parameters after training completion for hopping beam scheduling. Its specific implementation is as follows: the environment simulation sub-module 31 reads the stored beam position information, transmission rate, and service cache information, and establishes an environment model based on this information; the policy network sub-module 32 selects actions according to the current state in the environment model, obtains the corresponding rewards and the transferred state to form an experience tuple, and stores the experience tuple in the experience buffer in the experience replay sub-module 33; the experience replay sub-module 33 samples data in small batches from the experience buffer to calculate the target value, and updates the network parameters of the policy network sub-module 32 through backpropagation; the target network sub-module 34 periodically copies the network parameters of the policy network sub-module 32 and maintains a stable output during the training process to reduce fluctuations. After training completion, the model storage and loading sub-module 35 is responsible for saving the trained model parameters and loading the model for scheduling inference during hopping beam scheduling; the hopping beam time frame optimization module 4 maintains the hopping beam time frame structure to meet the requirements of hopping beam scheduling and allocates resources for users; the spot beam service transmission module 5 transmits service data to users according to the resource allocation scheme in the hopping beam time frame optimization module 4 to complete beam collaborative scheduling.

Claims

1. A multi-beam coordinated scheduling method for low-orbit satellite communications, characterized in that: include: The low-orbit satellite collects online user information through a wide beam, and uses this information to adaptively cluster and generate beam-hopping positions under the constraints of channel resources; The beam-hopping baseband monitors the changes in wide-beam information in real time, reads the beam position information as an environmental variable, and uses the environmental variable to periodically schedule beam-hopping based on deep reinforcement learning. The beam-hopping time frame structure is optimized in each scheduling cycle. The online user information provided by the wide beam is used to allocate service time slot resources to users. Point beams are used to cover users within the beam position to complete beam coordinated scheduling of low-orbit satellite communications.

2. The method according to claim 1, characterized in that The low-orbit satellite collects online user information through a wide beam, including: (2a) In the initial stage, the low-orbit satellite adopts a wide-beam coverage strategy to cover the entire area. Within the coverage area, user terminals attempt to access the system under a specified random access mechanism; (2b) During the access process, the user terminal reports its own location information, channel resource requirements, and current service cache status to the low-orbit satellite through an access request; (2c) The low-orbit satellite responds to the access request of the user terminal, allocates control channel resources to the user according to the current on-board resource situation, and records and compiles statistics on the information reported by the user to complete the collection of online user information.

3. The method according to claim 1, characterized in that The method of utilizing online user information to adaptively cluster and generate beam hopping positions under channel resource constraints includes: (3a) Mapping the low-orbit satellite wide beam coverage area into a coordinate plane, and mapping the reported geographic location information into the coordinate plane when the user accesses, using the horizontal and vertical coordinates as the user identification; (3b) When clustering, the location constraints are given priority, a single clustering effect is generated, and then the sum of the resource requirements of all users in each wave position is checked to see whether it meets the total channel resource limit: When any wave position does not meet the resource limit, the cluster radius is reduced, the number of clusters is increased, and a new cluster result is generated to continue to determine whether the total channel resource limit is met; When all beam positions meet the resource constraints, the optimal beam hopping beam position scheme is obtained, and the beam hopping beam position generation process is completed. The optimal scheme can use the least number of beam positions to cover the user terminals in the area, the users of each beam position do not overlap in space, and the total resource demand is less than the total channel resource.

4. The method according to claim 1, characterized in that: The periodic scheduling of beam hopping based on deep reinforcement learning includes: (4a) Specifically, the DQN algorithm in deep reinforcement learning is used to schedule beam hopping. The same neural network structure is used to form an online policy network and a target network. The state vector from the environment is input into the online policy network, and the Q value corresponding to each possible action is output. The state vector includes the cache traffic volume, transmission rate and wave position; the action includes wave position selection and coverage time selection; the Q value is used to represent the expectation of the future cumulative reward that can be obtained by taking different actions in the current state, and the reward is the single-step reward that can be obtained when the action is executed; (4b) The experience tuple consisting of state, action, reward and next state is stored in a pre-set experience replay buffer. The experience replay mechanism is used to randomly sample small batches of data from the buffer periodically for subsequent network parameter updates to break the correlation between data. (4c) Using the Q value output by the current online policy network, the temperature annealing Softmax strategy is used to select actions, where the temperature parameter value is dynamically adjusted according to the number of training steps; (4d) In the target network, the target Q' value is estimated by the reward of the current online policy network and the optimal reward that may be obtained in the next state. The mean square error is used as the loss function to calculate the difference between the output Q value of the current online policy network and the target Q' value. The parameters of the online policy network are updated through the back propagation algorithm. (4e) Repeat the above steps (4b) to (4d) until the reward of the online policy network converges and obtains an approximately optimal state-action value function, which represents the optimal expected value of the future cumulative discounted reward that can be obtained by continuing to act according to a specific strategy after taking an action in a given state. (4f) The beam hopping is scheduled according to the optimal state-action value function, that is, the coverage position and coverage duration are selected at each scheduling moment to obtain the optimal reward.

5. The method according to claim 4, characterized in that The Q value corresponding to each possible action output by the online policy network in step (4a) is expressed as follows: Among them, s represents the current state, a represents the action taken in the current state, and r t+1 represents the immediate reward obtained in step t+1, and γ represents the discount factor.

6. The method according to claim 4, characterized in that The state-action value function in step (4c) is expressed as follows: Among them, s represents the current state, a represents the action taken in the current state, r represents the reward obtained by executing the current action a, γ represents the discount factor, s' represents the next state transferred after executing action a, max(Q * (s',a')) represents the maximum future cumulative reward that can be obtained in state s'.

7. The method according to claim 1, characterized in that The optimizing the beam hopping time frame structure in each scheduling period includes: (7a) The beam hopping scheduling period is divided into several subframes, and a single subframe is used as the coverage period of a single scheduling of any virtual beam hopping position. (7b) Each subframe is further divided into a guard interval, a synchronization subframe, a control subframe, an optimization interval, and a service subframe; (7c) Using the result of the division in (7b), the optimization of the beam hopping time frame structure is realized: (7c1) Using the guard interval to provide blank time between adjacent subframes to ensure that the order of frames will not conflict and provide time for the system to adjust; (7c2), establishing time synchronization between the satellite and the user in the synchronization subframe; (7c3) After time synchronization is completed, the control subframe is divided into several control time slots according to the fixed time slot length, and different time slots are allocated to users to transmit control information. The control information sends the beam hopping scheduling strategy to the user, and obtains the current service transmission status through the optimization interval to provide environmental feedback for the scheduling algorithm; (7c4) The service subframe is divided into several service time slots according to the fixed time slot length, and the service time slot resources are allocated to the users in the beam position using the resource allocation algorithm. After the user resource configuration is completed, the beam hopping executes the scheduling strategy and feeds back the scheduling results in the optimization interval to complete the optimization of the beam hopping time frame structure.

8. A multi-beam coordinated scheduling system for low-orbit satellite communications, characterized in that: include: Random access module, used for random access and user information collection of users in the wide beam coverage area; An adaptive clustering wave position design module is used to generate beam hopping wave positions through adaptive clustering; The beam-hopping scheduling module based on deep reinforcement learning is used for agent training and loads the trained network parameters for beam-hopping scheduling; The beam hopping time frame optimization module is used to design the beam hopping time frame structure and optimize resource allocation. The spot beam service transmission module is used to complete multi-user service transmission within the spot beam.

9. The system according to claim 8, characterized in that The adaptive clustering wave position design module comprises: The forward mapping submodule is used to map the user's actual geographic location information into plane coordinates; Mobility detection submodule, used to detect the location change and online status of the user; Cluster planning submodule, used for wave planning of the current user set; The reverse mapping submodule is used to reversely map the plane coordinate information of the wave position design results into the actual geographical location information; The wave position visualization submodule is used to visualize the wave position coverage plan.

10. The system according to claim 8, characterized in that The beam-hopping scheduling module based on deep reinforcement learning includes: The environment simulation submodule is used to simulate the beam-hopping mission status and its dynamic evolution; The experience replay submodule is used to store the state transition data collected by the agent in the environment to alleviate the correlation between samples; The policy network submodule is used for neural network approximation to help the agent estimate the value of each action from the state of the environment; The target network submodule is used to generate stable target values ​​and assist in optimizing the training process; The model storage and loading submodule is used to save the trained model parameters and load the model for inference.

Citation Information

Patent Citations

  • Multi-beam satellite communication resource allocation system and method based on deep reinforcement learning

    CN116505998A

Cited By

  • Intelligent beam forming system and method for millimeter wave communication

    CN120582653A

  • Multi-point beam flexible coverage planning method and system

    CN120825716A

  • Method and system for multi-point beam flexible coverage planning

    CN120825716B

  • Satellite beam coverage adjustment method and system based on reinforcement learning

    CN120896625A

  • Multi-satellite hopping beam scheduling method for random access

    CN121462066A