Low-orbit satellite resource allocation method and device based on wave beam and frequency domain collaborative optimization
By dynamically adjusting spectrum utilization and beam coverage through a resource allocation method based on collaborative optimization of beam and frequency domains, the problems of dynamic user demand and frequent switching in LEO satellite communication systems are solved, efficient space-time and frequency resource management is achieved, and system performance and user experience are improved.
Patent Information
- Application Number
- CN202510853207.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-05
AI Technical Summary
Existing multi-beam LEO satellite communication systems are insufficient in dynamically sensing the needs of ground users. Traditional resource management methods are unable to respond to real-time changes, resulting in limited control of co-channel interference between beams and limited system performance improvement. Frequent beam switching and high-speed satellite movement cause complex interference patterns, affecting communication quality and user experience.
A resource allocation method based on collaborative optimization of beam and frequency domain is adopted. Through dynamic soft frequency reuse, user clustering, deep reinforcement learning and autonomous switching mechanism, three-dimensional resource management of time, space and frequency is realized. Spectrum utilization and beam coverage are dynamically adjusted, channel model and switching matrix are optimized, multi-objective optimization problem model is constructed and solved efficiently using proximal strategy optimization algorithm.
It significantly improves the system's service efficiency and resource utilization, reduces co-channel interference, improves user density matching accuracy, ensures communication quality and service continuity, and breaks through the spectrum utilization limitations of traditional strategies.
Smart Images

Figure CN120601945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of low-orbit satellite resource allocation, and in particular to a method and device for low-orbit satellite resource allocation based on beam and frequency domain collaborative optimization. Background Art
[0002] As low-Earth-orbit (LEO) satellite communication systems mature, multi-beam architectures are being widely adopted in LEO constellations as a key means of improving system capacity and spectrum efficiency. Multi-beam LEO satellites can provide differentiated services to different geographic regions by reusing the same frequency resources through large-scale narrow beam coverage within spectrum-constrained conditions. However, given the highly time-varying and highly uneven spatial distribution of terrestrial user traffic demand, traditional resource management methods struggle to achieve efficient scheduling, limiting further improvements in overall system performance.
[0003] Existing multi-beam satellite beam-hopping mechanisms have significant shortcomings in dynamically sensing the needs of ground users. Specifically, beam hopping timing and activation sequences are typically based on preset rules and cannot flexibly respond to real-time changes in indicators such as ground user density and traffic intensity. Furthermore, traditional beam-hopping mechanisms often employ fixed or low-frequency reuse strategies, with limited means of controlling co-channel interference between beams. This limits the system's ability to expand concurrent service capabilities due to spectrum conflicts and makes it unable to meet the needs of scenarios with large numbers of users online simultaneously.
[0004] On the other hand, as LEO constellations expand and network dynamics increase, frequent beam switching and high-speed satellite motion trigger complex interference patterns. Existing frequency reuse schemes lack accurate modeling and response mechanisms for real-time interference conditions, making it difficult to suppress high-intensity co-channel interference at the beam edges. This leads to large fluctuations in communication quality and poor connection stability, impacting the user experience. Furthermore, current resource allocation systems, which are mostly centralized, are inefficient in handling high-dimensional, multi-parameter scheduling tasks, easily forming system bottlenecks, and struggle to adapt to highly dynamic orbital and user environments.
[0005] To address the above issues, there is an urgent need to build an adaptive resource scheduling solution that integrates beam hopping and frequency reuse mechanisms to achieve efficient collaborative management of spatial, time, and frequency domain resources in highly dynamic and high-density scenarios. Summary of the Invention
[0006] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method and device for allocating low-orbit satellite resources based on beam and frequency domain coordinated optimization. The technical solution is as follows:
[0007] On the one hand, a low-orbit satellite resource allocation method based on beam and frequency domain coordinated optimization is provided, the method comprising:
[0008] S1. Obtain satellite system status information and ground user service demand information. Based on the ground user distribution and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference.
[0009] S2: All ground users are spatially clustered according to their geographic location and traffic demand to obtain multiple user clusters. Based on these user clusters, coverage cells are dynamically divided to achieve balanced service.
[0010] S3: Based on dynamic cell division, select user clusters with high traffic volume according to priority, generate beam activation sets, and update the beam hopping matrix to achieve dynamic beam scheduling;
[0011] S4: Build a channel model between the user and the beam, calculate the channel gain, path loss, and signal-to-interference-and-noise ratio, and establish an inter-satellite handover matrix to handle link interruptions.
[0012] S5. Based on the established channel model, construct the system utility goal. Taking the system utility goal as the main consideration, comprehensively consider interference, switching overhead and delay penalty, build a multi-objective optimization problem model, and add various constraints.
[0013] S6. A multi-dimensional resource collaborative optimization framework based on deep reinforcement learning is used to reconstruct the multi-objective optimization problem model, including designing the state space, action space and reward function of the target optimization reinforcement learning network, and using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network. Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0014] Optionally, the S1 specifically includes:
[0015] The ground control center periodically collects the status information of LEO satellites, including: position trajectory, the number of active beams N of the satellite currently b , beam activation matrix A, the upper limit of the transmit power of each beam P max , bandwidth W and power resource occupancy state matrix P;
[0016] At the same time, collect demand information of ground user terminals, including: real-time location collection Flow density D u , communication delay tolerance L u ;
[0017] Based on the distribution of ground users and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a central sub-band and an edge sub-band. The central sub-band is dedicated to serving high-demand users in the center of the beam, and the edge sub-band is reused by non-adjacent beams. The ratio of the two is adaptively adjusted according to the real-time user distribution. For beam b, the edge sub-band ratio is Dynamically adjust the spatial density of edge users to balance the contradiction between spectrum reuse efficiency and co-channel interference suppression. Edge users are defined as users located outside half the half-power beamwidth of the beam. The edge subband ratio is defined as:
[0018]
[0019] The power shaping algorithm based on interference perception calculates the equivalent interference strength of adjacent beams in real time, and uses the inverse proportional attenuation function to dynamically adjust the transmission power of the edge sub-channel, so that the system can maintain high spectrum utilization while reducing the co-channel interference level at the edge of the beam to below the threshold. Through the joint optimization of the power domain and the frequency domain, it breaks through the limitation of the traditional fixed multiplexing strategy on the concurrent service capability of the system. The power shaping algorithm based on interference perception calculates the edge sub-channel The power is:
[0020]
[0021] Among them, P b,t Indicates the initial transmit power of the beam subchannel, the interference term I j,k,t Indicates the interference of this beam to other activated beams, I b,k,t represents the co-channel interference from other active beams, N0 is the noise power, W k Corresponds to the bandwidth allocated to subchannel k.
[0022] Optionally, the S2 specifically includes:
[0023] Assume that the total number of users is N, the clustering result is K clusters, and each user cluster is recorded as According to the number of active beams in the current satellite time slot and ground active user collection A variable grid strategy is used to initially screen users: for any two ground users u u with u v , if the spatial angular distance β u,v Satisfy β u,v ≤Δθ -3dB , then these two users are classified into the same user cluster, where Δθ -3dB represents the half-power beamwidth, which is determined by the variable grid strategy and used as the threshold control parameter for angular domain clustering;
[0024] Next, the P-center minimum cover optimization framework based on angular distance is used to achieve iterative refinement of user partitioning across clusters, which is formally expressed as the following optimization objective:
[0025]
[0026] in, represents the geographic center of beam b, represents the geographic center of user u, represents user u and beam center The angular distance between the user Assigned to the number of currently active beams In the clusters, the maximum coverage radius of all beams is minimized. In the initial stage, randomly select The user is used as the initial cluster center, and the remaining users are assigned to the nearest cluster center in turn to form the initial cluster grouping. For any cluster The coverage radius is defined as the maximum distance from the user in the cluster to the center, which is expressed as follows:
[0027]
[0028] Subsequently, the extreme cluster with the largest radius is identified The corresponding coverage radius is recorded as Add it to the subset As the optimization target of the current iteration, for each Try to migrate it to the adjacent cluster in turn And judge whether the global maximum radius can be reduced. The specific judgment conditions are:
[0029]
[0030] If the condition is met, user u is removed from the original cluster Adjust to Cluster
[0031] Iteratively identify extreme clusters with the largest radius Until the maximum cluster coverage radius cannot be further reduced until.
[0032] Optionally, the S3 specifically includes:
[0033] First, calculate the weight of the traffic to be processed, the cell c of a certain user cluster j The priority weight calculation method at time slot t is:
[0034]
[0035] Among them, D uis the flow density, τ is the coefficient, It means giving priority to underserved users;
[0036] Next, select the top N servers with the highest volume of business to be processed. b The cell of the user cluster generates an active beam set To ensure priority service in high-demand areas, a time slot beam activation framework is adopted. The system timeline is divided into N time slots with fixed durations, each of which lasts T s milliseconds, represents the beam activation matrix, which is defined as:
[0037]
[0038] The beam activation sequence is based on beam b m The cumulative flow demand D at time slot t b (t) Dynamic adjustment, cumulative flow demand D b The expression of (t) is:
[0039]
[0040] in, is the set of users served by beam b in time slot t.
[0041] Optionally, the S4 specifically includes:
[0042] First, the channel is modeled, user u n Receive beam b at time slot t m The gain is calculated as:
[0043]
[0044] Among them, α m,n (t) = θ m,n (t) / Δθ -3dB is the normalized off-axis ratio, which represents the angular offset of the user relative to the half-power width of the beam, θ m,n (t) is beam b m With user u n The angular deviation of the satellite is calculated as follows:
[0045]
[0046] in, is the geocentric coordinate of satellite s at time t, is the user position coordinate, η is the antenna aperture efficiency factor, J ω (·) is the first kind ω-order Bessel function;
[0047] Secondly, calculate the path loss and channel gain, user un With beam b m The time-varying distance is defined as:
[0048]
[0049] Among them, h s is the satellite orbit height, R e is the average radius of the Earth’s equator, and the free space path loss is calculated as:
[0050]
[0051] Where λ represents the wavelength, so the channel gain calculation formula is defined as:
[0052] h m,n (t) = G m,n (t)L m,n (t)κ(t),
[0053] Where κ(t) is the additional loss factor caused by atmospheric absorption and shadowing;
[0054] Next, the signal to interference noise ratio is calculated. When adjacent beams share spectrum resources, user u n The signal to interference and noise ratio is modeled as:
[0055]
[0056] Among them, P m (t) is beam b m The transmission power, ∑ j≠m P j (t)h j,n (t) represents the total interference caused by frequency reuse;
[0057] The on-board processing and forwarding equipment is used to trigger the switching autonomously, and the binary switching matrix H = [ho u,t ] u×T for:
[0058]
[0059] Among them, θ b,max is the maximum coverage angle of the satellite beam.
[0060] Optionally, the S5 specifically includes:
[0061] First, according to Shannon's formula, the instantaneous rate of user u in time slot t is defined as:
[0062] r u,k,t =W k log2(1+γ u,k,t )
[0063] Where W k Corresponding to the bandwidth allocated to subchannel k, the total data rate of user u during the beam hopping period of T time slots is:
[0064]
[0065] Among them, w u,k,t ∈{0,1} is the subchannel allocation matrix for user u;
[0066] Furthermore, to ensure balanced throughput among users with different channel conditions and requirements, the system utility is defined as:
[0067]
[0068] Taking the system utility goal as the main objective, the multi-objective optimization problem is constructed by integrating interference, switching overhead and delay penalty, which can be expressed as:
[0069]
[0070] Where A is the beam activation matrix, W = [w u,k,t ] is the subchannel allocation matrix, P = [P b,k,t ] is the power allocation matrix, H=[ho u,t ] is the switching decision matrix, the interference term I b,k,t =∑ j≠m P j (t)h j,n (t), represents the co-channel interference from other activated beams, and the total switching overhead is Delay L u =Q u / R u , is the queue length Q u Ratio to data rate, Q u By the flow density D u (t) is accumulated, and the weights λ, μ and ν are used to balance these terms;
[0071] Furthermore, to ensure that the number of beams activated simultaneously by each satellite s in any time slot t does not exceed the maximum number of beams that can be activated simultaneously under the interference limit N b,max , subject to the following constraints:
[0072]
[0073] Subchannel allocation follows the orthogonal allocation principle to avoid intra-beam interference. Subchannel k in beam b is only available when the beam is activated (a b,t =1) can be allocated to the user, expressed as:
[0074]
[0075] To mitigate co-channel interference in frequency reuse scenarios, edge subbands are subject to spatial isolation constraints to prevent adjacent beams from using the same edge subchannel simultaneously in the same time slot. If beam b is assigned an edge subchannel k∈K edge , whose spatially adjacent beams The same subchannel should not be activated in the same time slot t, which is expressed as:
[0076]
[0077] in, represents the set of beams adjacent to beam b. This constraint ensures that edge subbands are reused only in non-adjacent beams, balancing spectrum efficiency and interference suppression.
[0078] The total transmit power of all beams and subchannels must be lower than the satellite's maximum power budget P max , to prevent the power amplifier from saturating and ensure energy-efficient operation, expressed as:
[0079]
[0080] To meet the user-specific service level agreement requirements, each user u must achieve a minimum data rate R min,u , to ensure basic service quality, expressed as:
[0081] C5:R u ≥R min,u ,
[0082] For delay-sensitive applications, the end-to-end delay L u Must not exceed the threshold L max , expressed as:
[0083] C6:L u ≤L max ,
[0084] In order to comply with the International Telecommunication Union Radio Regulations, the equivalent power flux density needs to be dynamically controlled to ensure that the interference from the LEO system to the geostationary orbit GEO network remains within the international standard range, which is expressed as:
[0085]
[0086] Among them, P i (t) is the transmission power of the i-th LEO satellite, G i (θ i (t)) is the antenna gain with respect to the off-axis angle θ i Function, G GEO (φi (t)) is the GEO ground station antenna gain with respect to the off-axis angle φ i function, d i (t) is the Euclidean distance between the LEO satellite and the GEO ground station.
[0087] Optionally, the S6 specifically includes:
[0088] Real-time collection of network state parameters to construct a composite state space, where the state vector state t , including: beam real-time business requirements D b (t), user channel gain h b,u (t), user queue length Q u (t), beam interference level I b (t), user switching indication matrix ho u (t);
[0089] Generate a hybrid action space where the action vector act t Including: sub-channel allocation decision matrix w b,k,t ∈{0,1}, indicating whether subchannel k is allocated to beam b in time slot t, and the power adjustment decision matrix ΔP b,t ∈[-P step ,P step ], represents the power change step;
[0090] The reward function feeds back the system utility and guides the policy gradient update, which is defined as:
[0091]
[0092] The proximal policy optimization algorithm is used to update the executor network and the evaluator network of the target optimization reinforcement learning network. In the executor network, the subchannel allocation adopts the Bernoulli strategy and outputs the binary probability. The power adjustment adopts the Gaussian strategy and outputs the mean and variance. In the evaluator network, the state value function V(state t ), trained with temporal difference targets:
[0093]
[0094] Where γ is the discount factor;
[0095] The actor network updates the policy by clipping the alternative objective function:
[0096]
[0097] in, is the strategy probability ratio, is the advantage function,∈control strategy update amplitude;
[0098] Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0099] On the other hand, a low-orbit satellite resource allocation device based on beam and frequency domain coordinated optimization is provided, the device comprising:
[0100] The acquisition and division module is used to obtain satellite system status information and ground user service demand information. Based on the distribution of ground users and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference.
[0101] The spatial clustering module is used to spatially cluster all ground users according to their geographical location and traffic demand to obtain multiple user clusters. The coverage cells are dynamically divided based on the user clusters to achieve balanced service.
[0102] The generation module is used to select user clusters with high traffic volume according to priority based on dynamic cell division, generate beam activation sets and update the beam hopping matrix to achieve dynamic beam scheduling;
[0103] A module is used to build the channel model between users and beams, calculate channel gain, path loss and signal-to-interference-and-noise ratio, and establish the inter-satellite handover matrix to cope with link interruptions;
[0104] The construction module is used to build the system utility target based on the established channel model. Taking the system utility target as the main factor, it integrates interference, switching overhead and delay penalty to build a multi-objective optimization problem model and add various constraints.
[0105] A reconstruction module is used to reconstruct the multi-objective optimization problem model based on a multi-dimensional resource collaborative optimization framework of deep reinforcement learning, including designing the state space, action space and reward function of the target optimization reinforcement learning network, and using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network. Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0106] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned low-orbit satellite resource allocation method based on coordinated optimization of beam and frequency domain.
[0107] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the above-mentioned low-orbit satellite resource allocation method based on coordinated optimization of beam and frequency domains.
[0108] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0109] This paper addresses the shortcomings of existing multi-beam LEO satellite communication systems in dynamic resource allocation, spectrum utilization, and switching stability. By integrating user clustering modeling, dynamic beam scheduling, deep reinforcement learning optimization, and frequency reuse mechanisms, a resource allocation method based on three-dimensional joint optimization of time, space, and frequency is proposed. This method solves problems such as beam coverage mismatch, increased co-channel interference, and poor service continuity in highly dynamic scenarios, significantly improving the system's service efficiency and resource utilization. Specifically:
[0110] 1) The present invention realizes efficient collaborative management of time-space-frequency multi-dimensional resources in the high-dynamic communication scenarios of multi-beam LEO satellites by integrating dynamic beam scheduling and spectrum reuse mechanisms. Compared with the traditional static resource allocation method, the present invention constructs a time-space adaptive beam activation framework based on user clustering and dynamic cell division. Through the variable grid strategy and the angular distance-driven P-center optimization algorithm, the modeling accuracy of user spatial distribution is significantly improved. The global minimization of the beam coverage radius is taken as the optimization goal. The identification and elimination of extreme clusters are realized in the iterative reallocation process, so that the beam resources can accurately match the uneven characteristics of the ground user density and reduce the invalid coverage loss in the space division multiplexing scenario.
[0111] 2) In the dimension of frequency domain resource management, the present invention divides the frequency band into central sub-band and edge sub-band through a dynamic soft frequency reuse mechanism, and combines the sub-band ratio adjustment adaptive to the edge user density to effectively balance the contradiction between spectrum reuse efficiency and co-channel interference suppression. The power shaping algorithm based on interference perception calculates the equivalent interference intensity of adjacent beams in real time, and adopts an inverse proportional attenuation function to dynamically adjust the transmission power of the edge sub-channel, so that the system can maintain high spectrum utilization while reducing the co-channel interference level at the edge of the beam to below the threshold. Through the joint optimization of the power domain and the frequency domain, it breaks through the limitations of the traditional fixed multiplexing strategy on the system's concurrent service capabilities.
[0112] 3) This paper uses a deep reinforcement learning framework to achieve collaborative optimization of multi-dimensional parameters. By constructing a composite state space containing beam activation status, channel gain, inter-satellite handover indication, and user queue length, and designing a hybrid action space containing sub-channel allocation, power adjustment, and handover triggering, the policy network can autonomously perceive system dynamics and generate decisions. The proximal policy optimization algorithm, through an executor-evaluator dual network architecture, achieves efficient solution to high-dimensional non-convex optimization problems while ensuring the stability of policy updates. The introduction of a balancing mechanism between the logarithmic utility function and the interference penalty term in the reward function enables the system to improve the total throughput while ensuring the fairness of resource allocation among users.
[0113] 4) The on-board processing and forwarding equipment of the present invention autonomously triggers the switching mechanism, effectively alleviating the frequent switching problem caused by the high-speed movement of the satellite. Based on real-time off-axis angle monitoring and the minimum circle coverage criterion, a binary switching matrix is generated and the optimal switching timing is predicted. Through autonomous on-board decision-making, the dependence on the ground control center is reduced, and the service continuity of the system in the scenario of dynamic orbit changes is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0114] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0115] Figure 1 This is a flow chart of a low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization provided by an embodiment of the present invention;
[0116] Figure 2 Schematic diagram of a multi-beam LEO satellite system communication model provided by an embodiment of the present invention;
[0117] Figure 3 This is a block diagram of a low-orbit satellite resource allocation device based on beam and frequency domain collaborative optimization provided by an embodiment of the present invention;
[0118] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0119] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0120] An embodiment of the present invention provides a low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization. The method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of the method is shown, and the processing flow may include the following steps:
[0121] S1. Obtain satellite system status information and ground user service demand information. Based on the ground user distribution and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference.
[0122] Optionally, the S1 specifically includes:
[0123] The ground control center periodically collects the status information of LEO satellites, including: position trajectory, the number of active beams N of the satellite currently b , beam activation matrix A, the upper limit of the transmit power of each beam P max , bandwidth W and power resource occupancy state matrix P;
[0124] At the same time, collect demand information of ground user terminals, including: real-time location collection Flow density D u , communication delay tolerance L u ;
[0125] Based on the distribution of ground users and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a central sub-band and an edge sub-band. The central sub-band is dedicated to serving high-demand users in the center of the beam, and the edge sub-band is reused by non-adjacent beams. The ratio of the two is adaptively adjusted according to the real-time user distribution. For beam b, the edge sub-band ratio is Dynamically adjust the spatial density of edge users to balance the contradiction between spectrum reuse efficiency and co-channel interference suppression. Edge users are defined as users located outside half the half-power beamwidth of the beam. The edge subband ratio is defined as:
[0126]
[0127] The power shaping algorithm based on interference perception calculates the equivalent interference strength of adjacent beams in real time, and uses the inverse proportional attenuation function to dynamically adjust the transmission power of the edge sub-channel, so that the system can maintain high spectrum utilization while reducing the co-channel interference level at the edge of the beam to below the threshold. Through the joint optimization of the power domain and the frequency domain, it breaks through the limitation of the traditional fixed multiplexing strategy on the concurrent service capability of the system. The power shaping algorithm based on interference perception calculates the edge sub-channel The power is:
[0128]
[0129] Among them, Pb,t Indicates the initial transmit power of the beam subchannel, the interference term I j,k,t Indicates the interference of this beam to other activated beams, I b,k,t represents the co-channel interference from other active beams, N0 is the noise power, W k Corresponds to the bandwidth allocated to subchannel k.
[0130] Traditional frequency reuse strategies suppress co-channel interference by adjusting the frequency reuse factor, but are unable to dynamically adjust the transmit power, resulting in low spectrum utilization. To address this issue, the dynamic soft frequency reuse scheme in an embodiment of the present invention defines the extent to which a frequency is used by adjusting the transmit power. For edge and center sub-bands, the transmit power is adaptively adjusted through interference sensing, thereby suppressing co-channel interference and improving spectrum utilization.
[0131] S2: All ground users are spatially clustered according to their geographic location and traffic demand to obtain multiple user clusters. Based on these user clusters, coverage cells are dynamically divided to achieve balanced service.
[0132] Optionally, the S2 specifically includes:
[0133] Assume that the total number of users is N, the clustering result is K clusters, and each user cluster is recorded as According to the number of active beams in the current satellite time slot and ground active user collection A variable grid strategy is used to initially screen users: for any two ground users u u with u v , if the spatial angular distance β u,v Satisfy β u,v ≤Δθ -3dB , then these two users are classified into the same user cluster, where Δθ -3dB represents the half-power beamwidth, which is determined by the variable grid strategy and used as the threshold control parameter for angular domain clustering;
[0134] Next, the P-center minimum cover optimization framework based on angular distance is used to achieve iterative refinement of user partitioning across clusters, which is formally expressed as the following optimization objective:
[0135]
[0136] in, represents the geographic center of beam b, represents the geographic center of user u, represents user u and beam center The angular distance between the user Assigned to the number of currently active beams In the clusters, the maximum coverage radius of all beams is minimized. In the initial stage, randomly select The user is used as the initial cluster center, and the remaining users are assigned to the nearest cluster center in turn to form the initial cluster grouping. For any cluster The coverage radius is defined as the maximum distance from the user in the cluster to the center, which is expressed as follows:
[0137]
[0138] Subsequently, the extreme cluster with the largest radius is identified The corresponding coverage radius is recorded as Add it to the subset As the optimization target of the current iteration, for each Try to migrate it to the adjacent cluster in turn And judge whether the global maximum radius can be reduced. The specific judgment conditions are:
[0139]
[0140] If the condition is met, user u is removed from the original cluster Adjust to Cluster
[0141] Iteratively identify extreme clusters with the largest radius Until the maximum cluster coverage radius cannot be further reduced until.
[0142] Traditional static user grouping methods cannot adapt to the strong time-varying and non-balanced characteristics of terrestrial user distribution, resulting in inefficient resource allocation. The embodiments of the present invention construct an angular distance-driven P-center minimum coverage optimization framework, adopt a variable grid strategy for initial user screening, and implement cross-cluster iterative redistribution based on the angular domain clustering threshold. With minimizing the maximum cluster coverage radius as the optimization goal, through extreme cluster identification, user migration and cluster center dynamic update mechanism, users are accurately divided into geographic clusters that match the beam half-power width, thereby reducing invalid coverage loss and improving the spatial matching accuracy of beam resources and user density.
[0143] S3: Based on dynamic cell division, select user clusters with high traffic volume according to priority, generate beam activation sets, and update the beam hopping matrix to achieve dynamic beam scheduling;
[0144] Optionally, the S3 specifically includes:
[0145] First, calculate the weight of the traffic to be processed, the cell c of a certain user cluster j The priority weight calculation method at time slot t is:
[0146]
[0147] Among them, D u is the flow density, τ is the coefficient, It means giving priority to underserved users;
[0148] Next, select the top N servers with the highest volume of business to be processed. b The cell of the user cluster generates an active beam set To ensure priority service in high-demand areas, a time slot beam activation framework is adopted. The system timeline is divided into N time slots with fixed durations, each of which lasts T s milliseconds, let A = [a m,t ] B×N represents the beam activation matrix, which is defined as:
[0149]
[0150] The beam activation sequence is based on beam b m The cumulative flow demand D at time slot t b (t) Dynamic adjustment, cumulative flow demand D b The expression of (t) is:
[0151]
[0152] in, is the set of users served by beam b in time slot t.
[0153] An embodiment of the present invention proposes a dynamic cell division and priority-driven beam activation strategy. The traditional beam hopping mechanism activates beams based on preset rules and is difficult to respond to real-time changes in business demand. The embodiment of the present invention calculates the cell priority by weighting the amount of business to be processed, selects high-demand areas to generate an activation beam set, and adopts a time slot beam activation framework to dynamically adjust the beam service sequence. In low-load scenarios, combined with orthogonal sub-channel allocation constraints and power budget limitations, flexible switching of resource allocation is achieved, allowing the system to reduce ineffective energy consumption while ensuring coverage fairness.
[0154] S4: Build a channel model between the user and the beam, calculate the channel gain, path loss, and signal-to-interference-and-noise ratio, and establish an inter-satellite handover matrix to handle link interruptions.
[0155] Optionally, the S4 specifically includes:
[0156] First, the channel is modeled, user u n Receive beam b at time slot t m The gain is calculated as:
[0157]
[0158] Among them, α m,n (t) = θ m,n (t) / Δθ -3dB is the normalized off-axis ratio, which represents the angular offset of the user relative to the half-power width of the beam, θ m,n (t) is beam b m With user u n The angular deviation of the satellite is calculated as follows:
[0159]
[0160] in, is the geocentric coordinate of satellite s at time t, is the user position coordinate, η is the antenna aperture efficiency factor, J ω (·) is the first kind ω-order Bessel function;
[0161] Secondly, calculate the path loss and channel gain, user u n With beam b m The time-varying distance is defined as:
[0162]
[0163] Among them, h s is the satellite orbit height, R e is the average radius of the Earth’s equator, and the free space path loss is calculated as:
[0164]
[0165] Where λ represents the wavelength, so the channel gain calculation formula is defined as:
[0166] h m,n (t) = G m,n (t)L m,n (t)κ(t),
[0167] Where κ(t) is the additional loss factor caused by atmospheric absorption and shadowing;
[0168] Next, the signal to interference noise ratio is calculated. When adjacent beams share spectrum resources, user u n The signal to interference and noise ratio is modeled as:
[0169]
[0170] Among them, P m (t) is beam b m The transmission power, ∑ j≠m P j (t)h j,n (t) represents the total interference caused by frequency reuse;
[0171] The on-board processing and forwarding equipment is used to trigger the switching autonomously, and the binary switching matrix H = [ho u,t ] u×T for:
[0172]
[0173] Among them, θ b,max is the maximum coverage angle of the satellite beam.
[0174] The embodiment of the present invention designs an autonomously triggered switching mechanism for on-board processing and forwarding equipment. To address the frequent switching problem caused by the high-speed movement of satellites, based on real-time off-axis angle monitoring and the minimum circle coverage criterion, a binary switching matrix is generated and the optimal switching timing is predicted. Combined with the sub-channel pre-allocation strategy of the target beam and the power gradient control of the service beam and the target beam, the number of switching times is effectively reduced. This mechanism reduces dependence on the ground control center through autonomous on-board decision-making, and ensures service continuity in scenarios where the satellite orbit changes dynamically.
[0175] S5. Based on the established channel model, construct the system utility goal. Taking the system utility goal as the main consideration, comprehensively consider interference, switching overhead and delay penalty, build a multi-objective optimization problem model, and add various constraints.
[0176] Optionally, the S5 specifically includes:
[0177] First, according to Shannon's formula, the instantaneous rate of user u in time slot t is defined as:
[0178] r u,k,t =W k log2(1+γ u,k,t )
[0179] Where W k Corresponding to the bandwidth allocated to subchannel k, the total data rate of user u during the beam hopping period of T time slots is:
[0180]
[0181] Among them, w u,k,t ∈{0,1} is the subchannel allocation matrix for user u;
[0182] Furthermore, to ensure balanced throughput among users with different channel conditions and requirements, the system utility is defined as:
[0183]
[0184] Taking the system utility goal as the main objective, the multi-objective optimization problem is constructed by integrating interference, switching overhead and delay penalty, which can be expressed as:
[0185]
[0186] Where A is the beam activation matrix, W = [w u,k,t ] is the subchannel allocation matrix, P = [P b,k,t ] is the power allocation matrix, H=[ho u,t ] is the switching decision matrix, the interference term I b,k,t =∑ j≠m P j (t)h j,n (t), represents the co-channel interference from other activated beams, and the total switching overhead is Delay L u =Q u / R u , is the queue length Q u Ratio to data rate, Q u By the flow density D u (t) is accumulated, and the weights λ, μ and ν are used to balance these terms;
[0187] Furthermore, to ensure that the number of beams activated simultaneously by each satellite s in any time slot t does not exceed the maximum number of beams that can be activated simultaneously under the interference limit N b,max , subject to the following constraints:
[0188]
[0189] Subchannel allocation follows the orthogonal allocation principle to avoid intra-beam interference. Subchannel k in beam b is only available when the beam is activated (a b,t =1) can be allocated to the user, expressed as:
[0190]
[0191] To mitigate co-channel interference in frequency reuse scenarios, edge subbands are subject to spatial isolation constraints to prevent adjacent beams from using the same edge subchannel simultaneously in the same time slot. If beam b is assigned an edge subchannel k∈K edge , whose spatially adjacent beams The same subchannel should not be activated in the same time slot t, which is expressed as:
[0192]
[0193] in, represents the set of beams adjacent to beam b. This constraint ensures that edge subbands are reused only in non-adjacent beams, balancing spectrum efficiency and interference suppression.
[0194] The total transmit power of all beams and subchannels must be lower than the satellite's maximum power budget P max , to prevent the power amplifier from saturating and ensure energy-efficient operation, expressed as:
[0195]
[0196] To meet the user-specific service level agreement requirements, each user u must achieve a minimum data rate R min,u , to ensure basic service quality, expressed as:
[0197] C5:R u ≥R min,u ,
[0198] For delay-sensitive applications, the end-to-end delay L u Must not exceed the threshold L max , expressed as:
[0199] C6:L u ≤L max ,
[0200] In order to comply with the International Telecommunication Union Radio Regulations, the equivalent power flux density needs to be dynamically controlled to ensure that the interference from the LEO system to the geostationary orbit GEO network remains within the international standard range, which is expressed as:
[0201]
[0202] Among them, P i (t) is the transmission power of the i-th LEO satellite, G i (θ i (t)) is the antenna gain with respect to the off-axis angle θ i Function, G GEO (φ i (t)) is the GEO ground station antenna gain with respect to the off-axis angle φ i function, d i (t) is the Euclidean distance between the LEO satellite and the GEO ground station.
[0203] S6. A multi-dimensional resource collaborative optimization framework based on deep reinforcement learning is used to reconstruct the multi-objective optimization problem model, including designing the state space, action space and reward function of the target optimization reinforcement learning network, and using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network. Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0204] Optionally, the S6 specifically includes:
[0205] Real-time collection of network state parameters to construct a composite state space, where the state vector state t , including: beam real-time business requirements D b (t), user channel gain h b,u (t), user queue length Q u (t), beam interference level I b (t), user switching indication matrix ho u (t);
[0206] Generate a hybrid action space where the action vector act t Including: sub-channel allocation decision matrix w b,k,t ∈{0,1}, indicating whether subchannel k is allocated to beam b in time slot t, and the power adjustment decision matrix ΔP b,t ∈[-P step ,P step ], represents the power change step;
[0207] The reward function feeds back the system utility and guides the policy gradient update, which is defined as:
[0208]
[0209] The proximal policy optimization algorithm is used to update the executor network and the evaluator network of the target optimization reinforcement learning network. In the executor network, the subchannel allocation adopts the Bernoulli strategy and outputs the binary probability. The power adjustment adopts the Gaussian strategy and outputs the mean and variance. In the evaluator network, the state value function V(state t ), trained with temporal difference targets:
[0210]
[0211] Where γ is the discount factor;
[0212] The actor network updates the policy by clipping the alternative objective function:
[0213]
[0214] in, is the strategy probability ratio, is the advantage function,∈control strategy update amplitude;
[0215] Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0216] The embodiment of the present invention constructs a multi-dimensional resource collaborative optimization framework based on deep reinforcement learning. To address the problem that traditional centralized scheduling architectures are difficult to handle high-dimensional optimization tasks, the embodiment of the present invention adopts a proximal policy optimization algorithm to design a composite state space including beam activation state, channel gain and user queue length, as well as a hybrid action space for sub-channel allocation, power adjustment and switching triggering. Through the executor-evaluator dual network architecture, the policy network autonomously generates decisions and introduces a balancing mechanism between the logarithmic utility function and the interference penalty term in the reward function to achieve collaborative optimization of system throughput improvement and fairness among users, breaking through the local optimal limitation of traditional heuristic algorithms.
[0217] like Figure 3 As shown, an embodiment of the present invention further provides a low-orbit satellite resource allocation device based on beam and frequency domain coordinated optimization, the device comprising:
[0218] The acquisition and partitioning module 310 is used to obtain satellite system status information and ground user service demand information. Based on the ground user distribution and traffic demand, a dynamic soft frequency reuse scheme is used to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference.
[0219] The spatial clustering module 320 is used to spatially cluster all ground users according to their geographical locations and traffic demands to obtain multiple user clusters, and dynamically divide coverage cells based on the user clusters to achieve balanced service;
[0220] A generation module 330 is configured to select user clusters with high traffic volume according to priority based on dynamic cell division, generate a beam activation set, and update a beam hopping matrix to implement dynamic beam scheduling;
[0221] Establishing module 340 for establishing a channel model between the user and the beam, calculating the channel gain, path loss and signal-to-interference-and-noise ratio, and establishing an inter-satellite handover matrix to cope with link interruption;
[0222] A construction module 350 is used to construct a system utility target based on the established channel model, taking the system utility target as the main factor, comprehensively considering interference, switching overhead and delay penalty, constructing a multi-objective optimization problem model, and adding multiple constraints;
[0223] Reconstruction module 360 is used to reconstruct the multi-objective optimization problem model based on a multi-dimensional resource collaborative optimization framework of deep reinforcement learning, including designing the state space, action space and reward function of the target optimization reinforcement learning network, using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network, and finally the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
[0224] An embodiment of the present invention provides a low-orbit satellite resource allocation device based on beam and frequency domain collaborative optimization. Its functional structure corresponds to a low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization provided by an embodiment of the present invention, and will not be repeated here.
[0225] Figure 4 It is a structural diagram of an electronic device 400 provided in an embodiment of the present invention. The electronic device 400 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 401 and one or more memories 402, wherein the memory 402 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 401 to implement the steps of the above-mentioned low-orbit satellite resource allocation method based on coordinated optimization of beam and frequency domain.
[0226] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The instructions are executable by a processor in a terminal to implement the above-described method for allocating low-orbit satellite resources based on beam and frequency domain coordinated optimization. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0227] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0228] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization, characterized in that: The method comprises: S1. Obtain satellite system status information and ground user service demand information. Based on the ground user distribution and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference. S2: All ground users are spatially clustered according to their geographic location and traffic demand to obtain multiple user clusters. Based on these user clusters, coverage cells are dynamically divided to achieve balanced service. S3: Based on dynamic cell division, select user clusters with high traffic volume according to priority, generate beam activation sets, and update the beam hopping matrix to achieve dynamic beam scheduling; S4: Build a channel model between the user and the beam, calculate the channel gain, path loss, and signal-to-interference-and-noise ratio, and establish an inter-satellite handover matrix to handle link interruptions. S5. Based on the established channel model, construct the system utility goal. Taking the system utility goal as the main consideration, comprehensively consider interference, switching overhead and delay penalty, build a multi-objective optimization problem model, and add various constraints. S6. A multi-dimensional resource collaborative optimization framework based on deep reinforcement learning is used to reconstruct the multi-objective optimization problem model, including designing the state space, action space and reward function of the target optimization reinforcement learning network, and using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network. Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
2. The method according to claim 1, characterized in that Said S1 specifically includes: The ground control center periodically collects the status information of LEO satellites, including: position trajectory, the number of active beams N of the satellite currently b , beam activation matrix A, the upper limit of the transmit power of each beam P max , bandwidth W and power resource occupancy state matrix P; At the same time, collect demand information of ground user terminals, including: real-time location collection Flow density D u , communication delay tolerance L u ; Based on the distribution of ground users and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a central sub-band and an edge sub-band. The central sub-band is dedicated to serving high-demand users in the center of the beam, and the edge sub-band is reused by non-adjacent beams. The ratio of the two is adaptively adjusted according to the real-time user distribution. For beam b, the edge sub-band ratio is Dynamically adjust the spatial density of edge users to balance the contradiction between spectrum reuse efficiency and co-channel interference suppression. Edge users are defined as users located outside half the half-power beamwidth of the beam. The edge subband ratio is defined as: The power shaping algorithm based on interference perception calculates the equivalent interference strength of adjacent beams in real time, and uses the inverse proportional attenuation function to dynamically adjust the transmission power of the edge sub-channel, so that the system can maintain high spectrum utilization while reducing the co-channel interference level at the edge of the beam to below the threshold. Through the joint optimization of the power domain and the frequency domain, it breaks through the limitation of the traditional fixed multiplexing strategy on the concurrent service capability of the system. The power shaping algorithm based on interference perception calculates the edge sub-channel The power is: Among them, P b,t Indicates the initial transmit power of the beam subchannel, the interference term I j,k,t Indicates the interference of this beam to other activated beams, I b,k,t represents the co-channel interference from other active beams, N0 is the noise power, W k Corresponds to the bandwidth allocated to subchannel k.
3. The method according to claim 1, characterized in that Said S2 specifically includes: Assume that the total number of users is N, the clustering result is K clusters, and each user cluster is recorded as According to the number of active beams in the current satellite time slot and ground active user collection A variable grid strategy is used to initially screen users: for any two ground users u u with u v , if the spatial angular distance β u,v Satisfy β u,v ≤Δθ -3dB , then these two users are classified into the same user cluster, where Δθ -3dB represents the half-power beamwidth, which is determined by the variable grid strategy and used as the threshold control parameter for angular domain clustering; Next, the P-center minimum cover optimization framework based on angular distance is used to achieve iterative refinement of user partitioning across clusters, which is formally expressed as the following optimization objective: in, represents the geographic center of beam b, represents the geographic center of user u, represents user u and beam center The angular distance between the user Assigned to the number of currently active beams In the clusters, the maximum coverage radius of all beams is minimized. In the initial stage, randomly select The user is used as the initial cluster center, and the remaining users are assigned to the nearest cluster center in turn to form the initial cluster grouping. For any cluster The coverage radius is defined as the maximum distance from the user in the cluster to the center, which is expressed as follows: Subsequently, the extreme cluster with the largest radius is identified The corresponding coverage radius is recorded as Add it to the subset As the optimization target of the current iteration, for each Try to migrate it to the adjacent cluster in turn And judge whether the global maximum radius can be reduced. The specific judgment conditions are: If the condition is met, user u is removed from the original cluster Adjust to Cluster Iteratively identify extreme clusters with the largest radius Until the maximum cluster coverage radius cannot be further reduced until.
4. The method according to claim 1, wherein Said S3 specifically includes: First, calculate the weight of the traffic to be processed, the cell c of a certain user cluster j The priority weight calculation method at time slot t is: Among them, D u is the flow density, τ is the coefficient, It means giving priority to underserved users; Next, select the top N servers with the highest volume of business to be processed. b The cell of the user cluster generates an active beam set To ensure priority service in high-demand areas, a time slot beam activation framework is adopted. The system timeline is divided into N time slots with fixed durations, each of which lasts T s milliseconds, let A = [a m,t ] B×N represents the beam activation matrix, which is defined as: The beam activation sequence is based on beam b m The cumulative flow demand D at time slot t b (t) Dynamic adjustment, cumulative flow demand D b The expression of (t) is: in, is the set of users served by beam b in time slot t.
5. The method according to claim 1, wherein Said S4 specifically includes: First, the channel is modeled, user u n Receive beam b at time slot t m The gain is calculated as: Among them, α m,n (t) = θ m,n (t) / Δθ -3dB is the normalized off-axis ratio, which represents the angular offset of the user relative to the half-power width of the beam, θ m,n (t) is beam b m With user u n The angular deviation of the satellite is calculated as follows: in, is the geocentric coordinate of satellite s at time t, is the user position coordinate, η is the antenna aperture efficiency factor, J ω (·) is the first kind ω-order Bessel function; Secondly, calculate the path loss and channel gain, user u n With beam b m The time-varying distance is defined as: Among them, h s is the satellite orbit height, R e is the average radius of the Earth’s equator, and the free space path loss is calculated as: Where λ represents the wavelength, so the channel gain calculation formula is defined as: h m,n (t)=G m,n (t)L m,n (t)κ(t), Where κ(t) is the additional loss factor caused by atmospheric absorption and shadowing; Next, the signal to interference noise ratio is calculated. When adjacent beams share spectrum resources, user u n The signal to interference and noise ratio is modeled as: Among them, P m (t) is beam b m The transmission power, Σ j≠m P j (t)h j,n (t) represents the total interference caused by frequency reuse; The on-board processing and forwarding equipment is used to trigger the switching autonomously, and a binary switching matrix is generated based on real-time off-axis angle monitoring and minimum circle coverage criteria. for: Among them, θ b,max is the maximum coverage angle of the satellite beam.
6. The method according to claim 1, characterized in that Said S5 specifically includes: First, according to Shannon's formula, the instantaneous rate of user u in time slot t is defined as: r u,k,t =W k log2(1+γ u,k,t ) Where W k Corresponding to the bandwidth allocated to subchannel k, the total data rate of user u during the beam hopping period of T time slots is: Among them, w u,k,t ∈{0,1} is the subchannel allocation matrix for user u; Furthermore, to ensure balanced throughput among users with different channel conditions and requirements, the system utility is defined as: Taking the system utility goal as the main objective, the multi-objective optimization problem is constructed by integrating interference, switching overhead and delay penalty, which can be expressed as: Where A is the beam activation matrix, W = [w u,k,t ] is the subchannel allocation matrix, P = [P b,k,t ] is the power allocation matrix, H=[ho u,t ] is the switching decision matrix, the interference term I b,k,t =∑ j≠m P j (t)h j,n (t), represents the co-channel interference from other activated beams, and the total switching overhead is Delay L u =Q u / R u , is the queue length Q u Ratio to data rate, Q u By the flow density D u (t) is accumulated, and the weights λ, μ and ν are used to balance these terms; Furthermore, to ensure that the number of beams activated simultaneously by each satellite s in any time slot t does not exceed the maximum number of beams that can be activated simultaneously under the interference limit N bmax , subject to the following constraints: Subchannel allocation follows the orthogonal allocation principle to avoid intra-beam interference. Subchannel k in beam b is only available when the beam is activated (a b,t =1) can be allocated to the user, expressed as: To mitigate co-channel interference in frequency reuse scenarios, edge subbands are subject to spatial isolation constraints to prevent adjacent beams from using the same edge subchannel simultaneously in the same time slot. If beam b is assigned an edge subchannel k∈K edge , whose spatially adjacent beams The same subchannel should not be activated in the same time slot t, which is expressed as: in, represents the set of beams adjacent to beam b. This constraint ensures that edge subbands are reused only in non-adjacent beams, balancing spectrum efficiency and interference suppression. The total transmit power of all beams and subchannels must be lower than the satellite's maximum power budget P max , to prevent the power amplifier from saturating and ensure energy-efficient operation, expressed as: To meet the user-specific service level agreement requirements, each user u must achieve a minimum data rate R min,u , to ensure basic service quality, expressed as: For delay-sensitive applications, the end-to-end delay L u Must not exceed the threshold L max , expressed as: In order to comply with the International Telecommunication Union Radio Regulations, the equivalent power flux density needs to be dynamically controlled to ensure that the interference from the LEO system to the geostationary orbit GEO network remains within the international standard range, which is expressed as: Among them, P i (t) is the transmission power of the i-th LEO satellite, G i (θ i (t)) is the antenna gain with respect to the off-axis angle θ i Function, G GEO (φ i (t)) is the GEO ground station antenna gain with respect to the off-axis angle φ i function, d i (t) is the Euclidean distance between the LEO satellite and the GEO ground station.
7. The method according to claim 1, characterized in that Said S6 specifically includes: Real-time collection of network state parameters to construct a composite state space, where the state vector state t , including: beam real-time business requirements D b (t), user channel gain h b,u (t), user queue length Q u (t), beam interference level I b (t), user switching indication matrix ho u (t); Generate a hybrid action space where the action vector act t Including: sub-channel allocation decision matrix w b,k,t ∈{0,1}, indicating whether subchannel k is allocated to beam b in time slot t, and the power adjustment decision matrix ΔP b,t ∈[-P step ,P step ], represents the power change step; The reward function feeds back the system utility and guides the policy gradient update, which is defined as: The proximal policy optimization algorithm is used to update the executor network and the evaluator network of the target optimization reinforcement learning network. In the executor network, the subchannel allocation adopts the Bernoulli strategy and outputs the binary probability. The power adjustment adopts the Gaussian strategy and outputs the mean and variance. In the evaluator network, the state value function V(state t ), trained with temporal difference targets: Where γ is the discount factor; The actor network updates the policy by clipping the alternative objective function: in, is the strategy probability ratio, is the advantage function,∈control strategy update amplitude; Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
8. A low-orbit satellite resource allocation device based on beam and frequency domain collaborative optimization, characterized in that: The device comprises: The acquisition and division module is used to obtain satellite system status information and ground user service demand information. Based on the distribution of ground users and traffic demand, a dynamic soft frequency reuse scheme is adopted to divide the overall bandwidth of the user area into a center sub-band and an edge sub-band. The center sub-band is dedicated to serving high-demand users at the center of the beam, and the edge sub-band is multiplexed with non-adjacent beams. The transmit power is dynamically attenuated to suppress co-channel interference. The spatial clustering module is used to spatially cluster all ground users according to their geographical location and traffic demand to obtain multiple user clusters. The coverage cells are dynamically divided based on the user clusters to achieve balanced service. The generation module is used to select user clusters with high traffic volume according to priority based on dynamic cell division, generate beam activation sets and update the beam hopping matrix to achieve dynamic beam scheduling; A module is used to build the channel model between users and beams, calculate channel gain, path loss and signal-to-interference-and-noise ratio, and establish the inter-satellite handover matrix to cope with link interruptions; The construction module is used to build the system utility target based on the established channel model. Taking the system utility target as the main factor, it integrates interference, switching overhead and delay penalty to build a multi-objective optimization problem model and add various constraints. A reconstruction module is used to reconstruct the multi-objective optimization problem model based on a multi-dimensional resource collaborative optimization framework of deep reinforcement learning, including designing the state space, action space and reward function of the target optimization reinforcement learning network, and using the proximal policy optimization algorithm to update the executor network and evaluator network of the target optimization reinforcement learning network. Finally, the executor network outputs the solution of the multi-objective optimization problem model, including the optimal bandwidth and power allocation results.
9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that: The at least one instruction is loaded and executed by the processor to implement the low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization as described in any one of claims 1-7.
10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the low-orbit satellite resource allocation method based on beam and frequency domain collaborative optimization as described in any one of claims 1-7.
Citation Information
Cited By
Internet of Things terminal remote control method based on low earth orbit satellite communication
CN120825219A
Method for adjusting beam coverage area of low-altitude satellite group
CN121036841A
Multi-beam coherent high-reliability merging method and device, electronic equipment and storage medium
CN121077544A
Multi-satellite hopping beam scheduling method for random access
CN121462066A
Multi-user resource allocation method and device, satellite base station and storage medium
CN121547841A