Low-orbit satellite beam resource allocation method based on deep reinforcement learning

By applying deep reinforcement learning and improved clustering algorithms in low-orbit satellite communication systems, dynamically adjusting beam resource allocation, solving the problem of unreasonable resource allocation in traditional technologies, and achieving the effect of maximizing user fairness and throughput.

CN120150800APending Publication Date: 2025-06-13HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510347166.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional low-orbit satellite beam hopping technology is difficult to dynamically match the communication needs and beam transmission capabilities of ground cells, resulting in unreasonable resource allocation and unable to meet the needs of user fairness and maximizing throughput.

Method used

The deep reinforcement learning method is adopted and combined with the improved clustering algorithm, the communication needs of ground users are clustered, the beam resource allocation strategy is dynamically adjusted, narrow beam resources are allocated to hot spot areas, and wide beam resources are allocated to non-hot spot areas.

Benefits of technology

It realizes that while ensuring user fairness, it improves the throughput of the satellite system, dynamically adapts to user needs, optimizes resource allocation, and improves the flexibility and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120150800A_ABST
    Figure CN120150800A_ABST
Patent Text Reader

Abstract

The invention discloses a low-orbit satellite beam resource allocation method based on deep reinforcement learning, and the method comprises the steps: firstly building a satellite system model according to a low-orbit satellite system hopping beam scene, obtaining the communication traffic demand and geographic position of a ground user through a satellite traffic monitoring facility, and carrying out the preprocessing; secondly, carrying out user clustering on information of communication traffic demands of ground users, and dividing a ground cell into a hot spot area and a non-hot spot area; and then based on a satellite system model, establishing an optimization problem of multi-beam resource allocation and control, adopting a deep reinforcement learning method, dynamically adjusting a beam resource allocation strategy, allocating narrow beam resources to a hot spot area, and allocating wide beam resources to a non-hot spot area to realize system throughput maximization and user fairness. According to the method, the resource allocation problem in the low-orbit satellite hopping beam scene is effectively solved, and the throughput of a satellite system is improved while the user fairness is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite communication, and particularly relates to a method for dynamically allocating and covering low-earth orbit satellite beams based on deep reinforcement learning. Background Art

[0002] The low-earth orbit satellite network can provide broadband access services for users globally and is the main development direction of the current satellite communication network. The hopping beam scheduling technology is the key technology to improve the throughput performance of multi-beam satellite systems with limited resources.

[0003] Traditional fixed grid partitioning or predefined priority allocation of hopping beams cannot effectively cope with complex and uncertain environmental changes, and there are problems of inability to dynamically match the communication requirements of ground cells and the beam transmission capabilities, lacking flexibility and adaptability. The hopping beam technology can achieve the allocation of resources such as bandwidth, power, and beam shape to the correct beam unit in the most effective way at the correct time.

[0004] How to reasonably utilize satellite beam resources to achieve dynamic allocation and coverage of beams, so as to ensure fairness among user terminals while maximizing throughput during the entire communication process, is a problem that needs to be solved. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method for dynamically allocating low-earth orbit satellite hopping beam resources based on deep reinforcement learning. First, considering the uneven distribution and high dynamicity of the communication requirements of ground users, the communication requirements of ground users are clustered based on an improved clustering algorithm, which can fully adapt to the high dynamicity of ground users and solve the problem of unreasonable beam resource allocation due to large differences in communication requirements. Furthermore, a method based on deep reinforcement learning is proposed for dynamic allocation and coverage control of satellite beams, which improves the throughput of the satellite system while ensuring fairness among users.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions:

[0007] A method for covering and controlling low-earth orbit satellite beams based on deep reinforcement learning includes the following steps:

[0008] S1. Establish a satellite system model according to the hopping beam scenario of the low-earth orbit satellite system

[0009] The research scenario of the present invention is for the hopping beam scenario of low-earth orbit (LEO) satellites, and a multi-user, multi-beam low-earth orbit satellite communication system model is established. It is assumed that the system is served by a single low-earth orbit satellite, and the satellite is equipped with multiple adjustable beams.

[0010] S2. Based on the satellite system model in S1, obtain the communication volume requirements and geographical locations of ground users through satellite traffic monitoring facilities, and preprocess the data.

[0011] S3. Input the information of the traffic demand of terrestrial users into the model for user clustering, divide the terrestrial cells into hotspots and non-hotspots, and provide a basis for subsequent beam resource allocation.

[0012] S4. Based on the satellite system model constructed in S1, establish an optimization problem for multi-beam resource allocation and control. The goal is to maximize the system throughput and ensure delay fairness among users under the premise of meeting the system constraints.

[0013] S5. Based on the optimization problem established in S4, adopt the method of deep reinforcement learning to dynamically adjust the beam resource allocation strategy, allocate narrow beam resources to hotspots and wide beam resources to non-hotspots to maximize the system throughput and user fairness.

[0014] As a preferred solution, in step S1, for constructing the low-earth orbit satellite hopping beam scenario, assume that a single low-earth orbit satellite carries N beam transmitters and can allocate at most N beams. The entire terrestrial coverage area is divided into K cells according to user information. Assume that there is a satellite ground station at the center of each cell to receive and send communication requirements. The beam set is defined as N = {1, 2, … N}, the terrestrial cell set is defined as K = {1, 2, … K}, and the number of beams is much smaller than the number of cells (N << K).

[0015] As a preferred solution, in step S3, construct a terrestrial user clustering model based on an improved clustering algorithm, specifically including:

[0016] S31. Use the DBSCAN algorithm to cluster the user geographical locations and automatically identify areas with high density.

[0017] S32. On the basis of the DBSCAN clustering results obtained in step S31, further increase the traffic demand of the hotspots for clustering.

[0018] S33. Combine the DBSCAN clustering results and the user traffic demand, design a multi-dimensional distance metric formula, comprehensively consider the distance metrics of geographical location and communication demand, and improve the accuracy of clustering. The formula is designed as:

[0019]

[0020] In the formula, α and β are weight coefficients, which are parameters controlling geographical location and traffic demand respectively, λ i represents the traffic demand of the sample point xi, represents the average traffic demand of the clustering center cj, x is represents the s-th dimensional value of the sample point xi, c jsRepresents the s-th dimensional value of the clustering center cj.

[0021] S34. After completing the clustering of the sample points, calculate the centroid c of each cluster according to the formula j .

[0022] S35. Introduce the roulette wheel method to select the next clustering center, and preferentially select points that are farther away from the current clustering center to avoid over-concentration of the clustering centers. Calculate the distance dis(x 1 , x 2 , …, x m ) of the sample space X = {x m - c i ). min The probability P(x m ) that the point is selected as the next clustering center point.

[0023] S36. Repeat the above steps until the target number k of clustering is reached, complete the initialization of the clustering center points, and divide the ground users into k cells.

[0024] As a preferred solution, in step S4, based on the low-earth orbit satellite hopping beam scenario model constructed in S1, establish an optimization problem for multi-beam resource allocation and control:

[0025] S41. At time t, the number of data packets to be served recorded in the satellite buffer of cell c i is:

[0026]

[0027] where l is the queuing delay of the data packet, and l th+1 is the maximum tolerable queuing delay of the data packet. The number of data packets requested by cell c i at time t is and follows a Poisson distribution;

[0028] S42. The correlation between the number of unserved data packets at time t and the number of unserved data packets at adjacent times is:

[0029]

[0030] where, represents the number of unserved data packets at time t - 1, represents the number of data packets transmitted at time t.

[0031] S43. The relationship between the angle θ and the transmit antenna gain is expressed as:

[0032]

[0033] where a and b are weight coefficients, and L s is a constant, θ represents the orthogonal angle in degrees, and θ b represents the 3dB beamwidth. G m represents the maximum gain of the transmitting antenna.

[0034] S44. Calculate the channel gain h i from beam k to cell c k,i and the channel capacity

[0035] S45. The number of data packets sent by beam k to cell c at time t n is:

[0036]

[0037] S46. The long-term delay fairness among all terrestrial cells can be described as the delay gap between cells, which is in the form of:

[0038]

[0039] where is the average queue delay of each data packet in cell n at time slot t;

[0040] S47. While dynamically selecting the beam coverage strategy, it is necessary to maximize the system throughput to meet more communication requirements while ensuring user fairness. The expression of the multi-objective optimization problem is:

[0041]

[0042] C2: p k ≤ p max

[0043]

[0044] where P1 is to maximize the system throughput and ensure the delay fairness among cells. The constraint C1 means that the sum of the powers allocated to all beams should not exceed the total system power. The constraint C2 indicates that the transmit power of a single beam should not exceed the maximum power that a single beam can carry. C3 requires that each beam size is less than the threshold, represents the coverage radius of beam k at t, and γ is a weight parameter.

[0045] As a preferred solution, in step S5, a method based on deep reinforcement learning is used to allocate multi-beam satellite resources and perform a coverage control model, specifically including:

[0046] S51. Solve the optimal handover problem based on the deep reinforcement learning method, and construct a Markov (MDP) decision process to gradually solve the optimization problem;

[0047] S52. Obtain the state information recorded in the satellite buffer of the multi-beam satellite system to obtain the state space for deep reinforcement learning:

[0048]

[0049] Among them, represents the number of data packets to be served in cell c i ;

[0050] S53. The action space includes the covered cell number and the radius size of the beam. Wide beams are used to cover hot cells, and narrow beams are used to cover non-hot cells. The action space is defined as:

[0051]

[0052] Among them, n t represents the central cell ID of beam selection k, represents the coverage radius of beam k;

[0053] S54. In the beam allocation and coverage control based on the reinforcement learning algorithm, the goal is to maximize the system throughput and ensure the fairness of cell users. The reward function is defined as the expression for the multi-objective optimization problem we want to achieve:

[0054]

[0055] As an optimal solution, in step S5, the deep reinforcement learning algorithm includes:

[0056] S510. Initialize the Q network, the target network, and the step size parameter with random parameters, and set up an experience replay buffer. Initialize the attention mechanism. An attention mechanism module is introduced into the Q network (policy network) to dynamically calculate the weights of the hot regions.

[0057] S520. Receive the initial state space S t , calculate the attention weights. The weights of the hot regions are higher. Based on the current environmental observation data, determine the actions of the target low-earth orbit satellite and execute them.

[0058] S530. The reward obtained by the target low-earth orbit satellite after executing the execution action and the target environmental observation value of the next state collected are stored in the experience replay buffer together with the current environmental observation data in S520, the execution action of the target low-earth orbit satellite in S430, the reward of the deterministic policy network, and the target environmental observation data.

[0059] S540. Each satellite randomly samples a batch of experiences from the experience replay buffer for training and updates the target network parameter θ - , and updates it once every G steps using the parameter θ of the policy network Q to maximize the reward Q and minimize the loss value L(θ) to train the Q network.

[0060] S550. Repeat steps S20 - S540 until the algorithm converges to obtain the optimal action value function.

[0061] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0062] Aiming at the problems of low efficiency of the current low - earth - orbit satellite hopping beam technology algorithm, inability to well meet dynamic users, and uneven distribution of user traffic volume, the present invention proposes a hopping beam resource allocation method based on deep reinforcement learning, effectively solving the resource allocation problem in the low - earth - orbit satellite hopping beam scenario. Through DBSCAN hot - spot area detection and improved k - means clustering algorithm, it can not only dynamically adapt to real - time changing user needs, but also preferentially cover hot - spot areas. By using a double - network deep reinforcement learning algorithm with a discrete - continuous hybrid action space and introducing an attention mechanism to focus on hot - spot areas, the system can process hot - spot areas more intelligently and efficiently, achieving maximum system throughput and user fairness, providing an efficient and flexible solution for the resource optimization of satellite communication systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a flowchart of a method for low - earth - orbit satellite beam allocation and coverage control based on deep reinforcement learning according to the present invention;

[0064] Figure 2 is a schematic diagram of the satellite beam resource coverage scenario of an embodiment of the present invention;

[0065] Figure 3 is a framework diagram of the design of a deep reinforcement learning decision network structure of an embodiment of the present invention;

[0066] Figure 4 is a comparison chart of the average rewards of the adaptive reinforcement learning algorithm proposed by the present invention and other algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The following describes the implementation manners of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0068] Embodiment 1:

[0069] As Figure 2 shown, this example is a method for low-earth orbit satellite beam allocation and coverage control based on deep reinforcement learning. Based on an improved clustering algorithm, ground cells are divided into hot spots and non-hot spots. The double-network deep reinforcement learning (DDQN) algorithm is used to allocate the beam resources of low-earth orbit satellites, so as to achieve beam allocation and coverage control while maximizing throughput and ensuring user fairness.

[0070] As Figure 1 shown, it specifically includes the following steps:

[0071] S1. According to the hopping beam scenario of the low-earth orbit satellite system, establish a satellite system model;

[0072] Further, the establishment of the satellite system model in step S1 includes:

[0073] Assume that a satellite carries N beam transmitters, and the maximum number of beams that can be allocated is N. In addition, each satellite has a buffer to record the communication requirements of each cell at different time slots. The entire ground coverage area is divided into K cells according to user information. Assume that there is a satellite ground station at the center of each cell to receive and send communication requirements. The beam set is defined as N = {1, 2,... N}, the ground cell set is defined as K = {1, 2,... K}, and the number of beams is much smaller than the number of cells (N << K);

[0074] S2. The satellite traffic monitoring facility acquires information data.

[0075] Further, the information acquired in step S2 includes: the communication volume requirements and geographical locations of each ground user. Preprocess the acquired data, including data cleaning (removing noise and outliers by cleaning the data) and formatting the data to meet the requirements of subsequent model input.

[0076] S3. Based on an improved k-means clustering algorithm, perform ground user clustering. Input the information of the ground users acquired in S2 into the model for user clustering, and divide the ground cells into hot spots and non-hot spots.

[0077] Further, in step S3, an improved k-means clustering algorithm is used to divide ground cells into hotspots and non-hotspots. The specific steps are as follows.

[0078] S3.1 Use the DBSCAN algorithm to cluster user geographical locations, automatically identify areas with higher density as potential hotspot areas, and mark users in low-density areas as noise points. Subsequently, reassign through multi-dimensional distance metrics to avoid missing potential hotspot areas and affecting the subsequent dynamic allocation of beams.

[0079] S3.2 Based on the DBSCAN clustering results obtained in step S3.1, further enhance the traffic demand in hotspot areas. DBSCAN clustering is only based on the geographical location density of users and does not consider the communication requirements of users. In some areas, the user density may be low, but the communication requirements are high. Based on the DBSCAN clustering results, introduce the communication demand data λ i , and perform weighted adjustment on the clustering results to make the user demands in hotspot areas more prominent.

[0080] S3.3 Combine the DBSCAN clustering results and user traffic demands, design a multi-dimensional distance metric formula, comprehensively consider the distance metrics of geographical location and communication requirements, and improve the accuracy of clustering. The formula is designed as:

[0081]

[0082] In the formula, α and β are weight coefficients, which control the parameters of geographical location and traffic demand respectively, and λ i represents the traffic demand of sample point xi, represents the average traffic demand of cluster center cj, and x is represents the s-th dimension value of sample point xi, and c js represents the s-th dimension value of cluster center cj.

[0083] In specific calculations, first calculate the dis from sample point xi to all cluster centers, and then take the minimum value, that is, satisfy the condition of the formula:

[0084] dis(x i , c j ) = min{dis(x i , c 1 ), dis(x i , c 2 ), …, dis(x i , c j )} (2)

[0085] After completing the classification of sample points in S3.4, the centroid is weighted according to the user's communication requirements to make the centroid more biased towards the area with high communication requirements. The formula for calculating the weighted centroid is:

[0086]

[0087] where x m represents the currently selected sample point, c i represents the existing clustering center, and dis(x m -c i ) 2 represents the distance from the sample point x m to the existing clustering center c i . X represents the sample space, which contains all sample points;

[0088] In S3.5, according to the distance between the sample points and the current clustering center, the roulette method is used to select the next clustering center, giving priority to selecting points with a longer distance to avoid the clustering centers being too concentrated. Then calculate the shortest dis(x 1 ,x 2 ,..x m ) between each point in the sample space X = {x m -c i} and the existing clustering center points. Considering that the data differences in the communication requirements of ground users are relatively large, points farther from the current clustering center are more likely to become the next clustering center than those closer. Therefore, the probability P(x min ) that the sample point is selected as the next clustering center is defined by the formula for calculating the centroid of each cluster: m :

[0089]

[0090] where x m represents the currently selected sample point, c i represents the existing clustering center, and dis(x m -c i ) 2 represents the distance from the sample point x m to the existing clustering center c i . X represents the sample space, which contains all sample points;

[0091] Finally, the clustering process needs to be iterated multiple times to optimize the clustering results. Repeat steps S33 - S35, update the clustering center and cluster division, and calculate the stability index of the clustering results until the change in the stability index is less than the threshold ∈ and the target number k of clusters is reached, then stop the iteration, complete the initialization of the clustering center, and divide the ground users into k cells.

[0092] S4. Establish the optimization problem of satellite multi-beam resource allocation and control;

[0093] Further, in step S4, to establish the optimization problem of satellite multi-beam resource allocation and control, the specific steps include:

[0094] S4.1 At time t, the number of data packets to be served recorded in the satellite buffer of cell c i is

[0095]

[0096] where l is the queuing delay of the data packet, and l th+1 is the maximum tolerable queuing delay of the data packet. The number of data packets requested by cell c i at time t is and follows a Poisson distribution.

[0097] S4.2 The correlation between the number of unserved data packets at time t and the number of unserved data packets at adjacent times is:

[0098]

[0099] where represents the number of unserved data packets at time t - 1, represents the number of data packets transmitted at time t.

[0100] S4.3 The relationship between the angle θ and the transmit antenna gain is expressed as:

[0101]

[0102] where a = 2.88, b = 6.32, L s = -25dB, θ represents the orthogonal angle in degrees, and θ b represents the 3dB beamwidth. G m represents the maximum gain of the transmit antenna, and its form is:

[0103]

[0104] S4.4 The channel gain from beam k to cell c i can be expressed as:

[0105] h k,i = G tx G r L f (9)

[0106] where G r is the receive antenna gain, and L fis the free space loss.

[0107] S4.5 From beam k to cell c i The channel capacity can be expressed as:

[0108]

[0109] where p k and p i are the transmission powers of beam k and beam i, N 0 represents Gaussian white noise, W k represents the bandwidth of beam k, is used to determine whether cell c n is a cell covered by beam k. If so, otherwise

[0110] S4.6 The number of data packets sent by beam k to cell c at time t n is:

[0111]

[0112] S4.7 The long-term delay fairness among all ground cells can be described as the delay gap between cells, in the form of:

[0113]

[0114] where is the average queue delay of each data packet in cell n at time slot t.

[0115] S4.8 While dynamically selecting the beam coverage strategy, it is necessary to maximize the system throughput to meet more communication requirements while ensuring fairness for users. The expression of the multi-objective optimization problem is:

[0116]

[0117] C2: p k ≤ p max

[0118]

[0119] where P1 is to maximize the system throughput and ensure delay fairness among cells. Constraint C1 means that the sum of the powers allocated to all beams should not exceed the total system power. Constraint C2 indicates that the transmission power of a single beam should not exceed the maximum power that a single beam can carry. C3 requires that the size of each beam is less than the threshold, represents the coverage radius of beam k at t.

[0120] In the first embodiment, a method for low-earth orbit satellite beam allocation and coverage control based on deep reinforcement learning is based on the DDQN model. The diagram of this model is shown in Figure 2 As shown. In this network, the target network receives the state information and actions composed of the data in the satellite buffer as inputs during the training phase, calculates the action value function of the satellite based on this, and trains according to the estimated Q value and the actual Q value. The satellite updates its policy based on the feedback of the target network.

[0121] It includes the following steps:

[0122] S1. Solve the optimal handover problem based on the deep reinforcement learning method, and construct a Markov (MDP) decision process to gradually solve the optimization problem;

[0123] S2. Obtain the state information recorded in the satellite buffer of the multi-beam satellite system to obtain the state space for deep reinforcement learning:

[0124]

[0125] Among them, represents the number of data packets to be served in cell c i .

[0126] S3. The action space includes the covered cell number and the radius size of the beam. Wide beams are used to cover hot cells, and narrow beams are used to cover non-hot cells. The action space is defined as:

[0127]

[0128] Among them, n t represents the central cell ID of beam selection k, represents the coverage radius of beam k.

[0129] S4. In the beam allocation and coverage control based on the reinforcement learning algorithm, the goal is to maximize the system throughput and ensure the fairness of cell users. The reward function comprehensively considers the system throughput and user fairness and is defined as:

[0130]

[0131] Among them, μ and γ are weight coefficients that respectively control the importance of throughput and fairness, represents the number of data packet transmissions in cell c n at time t, and F represents the delay fairness between cells, which is defined as:

[0132]

[0133] Among them, is the average queue delay of each data packet in cell n at time slot t.

[0134] S5. As Figure 3 shown, initialize the Q-network and the target network, the step-size parameters (such as the learning rate, discount factor, etc.), and set up an experience replay buffer for storing historical experience data. Initialize the attention mechanism, and introduce an attention mechanism module in the Q-network (policy network) to dynamically calculate the weights of hot regions.

[0135] S6. Receive the initial state space S t , use the attention mechanism module to calculate the attention weights of each region, which reflect the importance of each region. The weights of hot regions are higher. Based on the current state S t , select an action a k (i.e., select the cell and the beam coverage radius) through the policy network. Among them, the discrete-continuous hybrid action space samples the discrete cell selection through the Categorical distribution and maps the continuous beam coverage radius through the Sigmoid function, realizing a flexible and fine resource allocation strategy. After executing the action, observe the feedback of the environment (including the reward r and the next state S t+1 ).

[0136] S7. The reward obtained by the target low-earth orbit satellite after executing the execution action and the target environmental observation value of the next state collected are stored in the experience replay buffer together with the current environmental observation data in S6, the execution action of the target low-earth orbit satellite, the reward of the deterministic policy network, and the target environmental observation data;

[0137] S8. Each satellite randomly samples a batch of experiences from the experience replay buffer for training, and updates the parameters θ of the Q-network using the sampled data to minimize the loss function:

[0138] L(θ) = E(Q(S t , a k ; θ) - y t ) 2 (18)

[0139] where θ represents the parameters of the policy network, E represents the expected value, Q(S t , a k ; θ) represents the predicted value of the policy network in state S t and action a k , S t represents the state at time t, a k represents the kth action in the action space, y t = r + γmax a′ Q(S t+1 , a′; θ -) is the target Q value,

[0140] Update the target network parameter θ - , and update the target network parameter θ every G steps using the parameter θ of the policy network Q - , and train the Q network by maximizing the reward Q and minimizing the loss value L(θ);

[0141] S9. Repeat steps S6 - S8 until the algorithm converges to obtain the optimal action value function.

[0142] S10 To evaluate the effectiveness of the model and algorithm, experimental simulations were carried out for the hopping beam satellite in the Ka band. The scenario design is as follows: The satellite coverage area is divided into 20 equal - sized beam position cells. The satellite uses 5 dynamically schedulable beams to perform hopping coverage on these 20 cells. In the simulation scenario, 1000 - round experiments were carried out with communication demands generated by Poisson distribution. Different traffic request volumes were set for each beam position cell covered by the satellite and converted into different arrival rates. The total communication demand for each cell at each moment is 2400 Mbps. The constructed Q network uses a 5 - layer convolutional neural network. First, the input state matrix extracts spatial features through two convolutional layers. Then, an attention score is calculated through a convolutional layer, and the Softmax function is used to generate attention weights to weight the feature map to enhance the attention to hot regions. Then, the weighted feature map is flattened into a one - dimensional vector, and the Q value of each action is generated through two fully - connected layers. Figure 4 shows the comparison of the average rewards of the adaptive deep reinforcement algorithm proposed in the present invention and three other fixed - radius algorithms based on DRL where the radius is fixed at 5, 10, and 20, as well as the random algorithm, greedy algorithm, and genetic algorithm.

[0143] It can be obtained by comparison that in the algorithm average reward comparison graph, the fixed - radius DRL has a slow convergence speed and low final performance because its fixed beam allocation strategy cannot adapt to dynamic demand changes. The genetic algorithm optimizes the beam allocation strategy through global search, but due to its high computational complexity, it has a slow convergence speed and limited performance improvement. Although the greedy algorithm has high real - time performance, due to its local optimization characteristics, it is easy to fall into sub - optimal solutions, so the performance improvement is slow and the final reward value is low. As a benchmark comparison algorithm, the random algorithm has the worst performance and the slowest convergence speed due to the lack of optimization ability. The adaptive deep reinforcement learning (Adaptive DRL) proposed in the present invention performs the best. Its curve rises rapidly in the initial stage of training, reaches a high reward value in a short time, and finally stabilizes. This fast - convergence characteristic benefits from the dynamic adjustment ability of adaptive deep reinforcement learning, which can optimize the beam allocation strategy according to real - time demands.

[0144] The above description only elaborates in detail on the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, based on the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.

[0145] The above description only elaborates in detail on the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, based on the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.

Claims

1. A low-orbit satellite beam resource allocation method based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Establish a satellite system model based on the beam hopping scenario of the low-orbit satellite system; S2. Based on the satellite system model, the communication volume demand and geographical location of ground users are obtained through satellite traffic monitoring facilities, and the data is pre-processed; S3, inputting the information of the communication volume demand of ground users into the model to perform user clustering, and dividing the ground cells into hotspot areas and non-hotspot areas; S4. Based on the satellite system model, establish the optimization problem of multi-beam resource allocation and control to maximize the system throughput and ensure delay fairness among users; S5. Based on the optimization problem, deep reinforcement learning is used to dynamically adjust the beam resource allocation strategy, allocate narrow beam resources to hotspot areas, and allocate wide beam resources to non-hotspot areas to achieve maximum throughput and user fairness.

2. The method for allocating low-orbit satellite beam resources based on deep reinforcement learning according to claim 1, characterized in that: The step S1 is specifically implemented as follows: for the low-orbit satellite LEO beam hopping scenario, a multi-user, multi-beam low-orbit satellite communication system model is established; assuming that the system is served by a single low-orbit satellite, and the satellite is equipped with multiple adjustable beams; Assume that a single low-orbit satellite carries N beam transmitters and can allocate up to N beams. The entire ground coverage area is divided into K cells based on user information. Assume that there is a satellite ground station at the center of each cell to receive and send communication needs; the beam set is defined as N = {1, 2, ...N}, and the ground cell set is defined as K = {1, 2, ...K}. The number of beams is much smaller than the number of cells.

3. The method for allocating low-orbit satellite beam resources based on deep reinforcement learning according to claim 2, characterized in that: The specific implementation process of user clustering is as follows: S31. Clustering user geographic locations using the DBSCAN algorithm; S32, based on the DBSCAN clustering result obtained in step S31, adding the communication volume demand of the hotspot area for clustering; S33. Combining the DBSCAN clustering results and user communication volume requirements, a multidimensional distance measurement formula is designed to comprehensively consider the distance measurement of geographical location and communication requirements. The formula is designed as follows: In the formula, α and β are weight coefficients, which are parameters controlling the geographical location and traffic demand, respectively, and λ i represents the communication volume demand of sample point xi, represents the mean value of the communication demand of cluster center cj, x is represents the s-th dimension value of the sample point xi, c js Represents the sth dimension value of the cluster center cj; S34. After completing the clustering of the sample points, calculate the centroid c of each cluster j : S35, introduce the roulette wheel method to select the next cluster center and calculate the sample space X = {x1, x2, ..., x m }Distance dis(x m -c i ) min The probability of a point being selected as the next cluster center is P(x m ); S36. Repeat the above steps until the target number k of clusters is reached, the initialization of the cluster center points is completed, and the ground users are divided into k cells.

4. The method for allocating low-orbit satellite beam resources based on deep reinforcement learning according to claim 3 is characterized in that: The specific process of establishing the optimization problem in step S4 is as follows: S41. At time t, in cell c i The number of packets to be served recorded in the satellite buffer is: Where, l is the queuing delay of the data packet, l th+1 is the maximum tolerable queuing delay of the data packet; cell c i The number of packets requested at time t is Satisfies Poisson distribution; S42. The correlation between the number of unserved data packets at time t and the number of unserved data packets at adjacent times is: in, represents the number of packets not served at time t-1, represents the number of data packets transmitted at time t; The relationship between S43, angle θ and transmit antenna gain is expressed as: Where a and b are weight coefficients, L s is a constant, θ represents the orthogonal angle in degrees, θ b Indicates 3dB beam width; G m Indicates the maximum gain of the transmitting antenna G m ; S44. Calculate the distance from beam k to cell c i The channel gain h k,i and channel capacity S45, beam k to time c n The number of data packets sent by the cell is S46. The long-term delay fairness between all terrestrial cells is described as the delay gap F between cells: in is the average queue delay of each packet in cell n at time slot t; S47, while dynamically selecting the beam coverage strategy, it is necessary to maximize the system throughput to meet more communication needs while ensuring user fairness; the expression of the multi-objective optimization problem is: C2:p k ≤p max Among them, P1 is to maximize the system throughput and ensure the delay fairness between cells. Constraint C1 means that the sum of the power allocated to all beams should not exceed the total system power. Constraint C2 means that the transmission power of a single beam must not exceed the maximum power that a single beam can carry. C3 requires that the size of each beam is smaller than the threshold. represents the coverage radius of beam k at point t, and γ is the weight parameter.

5. The method for allocating low-orbit satellite beam resources based on deep reinforcement learning according to claim 4, characterized in that: The specific implementation process of step S5 is as follows: S51. Solve the optimal switching problem based on deep reinforcement learning method, and construct the Markov MDP decision process to gradually solve the optimization problem; S52, acquiring state information of the multi-beam satellite system recorded in the satellite buffer, and obtaining a state space for deep reinforcement learning; S53, the action space includes the number of the covered cell and the radius of the beam. Hotspot cells are covered by wide beams, and non-hotspot cells are covered by narrow beams. The action space is defined as: Among them, n t represents the center cell ID of beam selection k, represents the coverage radius of beam k; S54. In the beam allocation and coverage control based on reinforcement learning algorithm, the goal is to maximize the system throughput and ensure fairness among cell users. The reward function is defined as:

6. The method for allocating low-orbit satellite beam resources based on deep reinforcement learning according to claim 5, characterized in that: In the deep reinforcement learning, the Q network and the target network are initialized with random parameters, the step size parameters are set, the experience replay buffer is set, the attention mechanism is initialized, the attention mechanism module is introduced into the convolutional network of the Q network, and the weights of the hot spot areas are dynamically calculated.

Citation Information

Cited By

  • Satellite beam resource adjusting method and device, electronic equipment and storage medium

    CN121012566A

  • Methods, devices, electronic equipment, and storage media for adjusting satellite beam resources

    CN121012566B

  • Method for adjusting beam coverage area of low-altitude satellite group

    CN121036841A

  • Method for adjusting beam coverage area of low-altitude satellite group

    CN121036841B

  • Satellite resource allocation method and device, electronic equipment and storage medium

    CN121217211A