Layered beam coverage and bandwidth allocation method for large-scale flexible HTS

CN122554858APending Publication Date: 2026-08-11BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0009]为了解决现有技术存在的计算复杂度过高、局限于小规模系统、资源维度受限、开销大、在从延迟信息中学习时效率低下、未能在减少全局重构频率的同时保证服务连续性的问题,本发明提供一种面向大规模灵活HTS的分层波束覆盖与带宽分配方法

Benefits of technology

[0059] 1. This invention proposes a method to construct a Markov Decision Process (MDP) guided by the objective function of the optimization problem. Based on the different time scales and physical characteristics of beam layout configuration and bandwidth allocation, the multidimensional resource optimization problem (beam pointing, beam width, beam switching, and bandwidth allocation) that is difficult to solve jointly is decomposed into two independently optimizable sub-problems through time scale decoupling, thereby realizing spatial domain beam layout configuration at a large time scale and frequency domain bandwidth allocation at a small time scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554858A_ABST
    Figure CN122554858A_ABST
Patent Text Reader

Abstract

This invention presents a hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS (High-Throughput Satellite Telecommunications), belonging to the high-throughput satellite field. The method includes: decoupling the multidimensional resource optimization problem into two independently optimizable sub-problems; modeling the large-timescale sub-problems as a Markov decision process model by designing a PPO (Progressive Point of Interest) reinforcement learning beam layout agent based on a sequential collaborative decision framework; performing deterministic refinement based on the initial beam layout using a constraint-aware capacity-based clustering beam fine-tuning mechanism, identifying and shutting down inefficient beams, and outputting the refined beam layout configuration; adaptively selecting to maintain the current beam layout configuration, trigger incremental repair, or trigger global reconstruction through elastic coverage reconstruction control logic; and decomposing the small-timescale sub-problems into three coupled stages using a fast, fair bandwidth allocation algorithm to maximize demand satisfaction fairness. This invention can efficiently match onboard multidimensional resources with user needs, achieving load-balanced beam layout and fair bandwidth allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-throughput satellite technology, specifically relating to a layered beam coverage and bandwidth allocation method for large-scale flexible HTS. Background Technology

[0002] As 6G networks evolve towards integrated air-space-ground / non-terrestrial networks (NTN), the demand for broadband connectivity for airborne users such as civil aviation, drones, and aerial platforms is rapidly increasing. High-throughput satellites (HTS) can significantly improve capacity through multi-beam coverage and spectrum reuse; while flexible HTSs employing digital transparent relay payloads and digital beamforming (DBF) can perform reconfigurable resource configuration in the spatial and frequency domains, such as adjustable beam pointing, adjustable beamwidth, switchable beams, and on-demand bandwidth allocation.

[0003] However, in large-scale flexible HTS systems (with hundreds of beams), the distribution of users in the air exhibits strong non-uniformity and strong time-varying characteristics, making it difficult for traditional fixed grid beams and static frequency reuse to match dynamic traffic. Meanwhile, reconfiguration of spatial domain beam layouts typically involves high on-board computing and signaling overhead, and excessively frequent beam reconfiguration can lead to service interruption risks; while frequency domain bandwidth adjustment overhead is relatively low and should be able to track short-term traffic fluctuations more quickly. Therefore, there is a significant "configuration timescale mismatch" problem between spatial and frequency domain resources: beam layouts are more suitable for stable optimization on large timescales, while bandwidth allocation is more suitable for rapid adaptation on small timescales.

[0004] In existing technologies, methods based on convex optimization, heuristic search, and clustering are effective in small-scale systems or low-dimensional resource scenarios. However, when the number of beams increases to hundreds and simultaneous optimization of beam pointing / width / switching is required, the decision space experiences combinatorial explosion, making it difficult to meet the complexity requirements of online decision-making. Existing technologies mainly suffer from the following problems:

[0005] 1. Existing beam layout methods based on convex optimization and heuristic search (such as differential convex programming, continuous convex approximation, deterministic annealing, etc.) have computational complexity that increases exponentially with the resource dimension and the number of beams in the joint optimization of high-dimensional resources. Moreover, most of them run offline and are difficult to solve in real time in large-scale flexible HTS systems, and cannot adapt to the rapidly changing traffic distribution in dynamic satellite-aviation networks.

[0006] 2. Existing machine learning-based beam configuration methods (such as weighted K-means clustering, deep neural network approximation, and deep reinforcement learning) have shown better adaptability in dynamic environments, but are mostly limited to small-scale systems (no more than a hundred beams). Furthermore, constrained by limited resources, they fail to fully utilize the different temporal scale characteristics of spatial and frequency resources, fail to effectively decouple large-scale optimization of beam layout from small-scale optimization of bandwidth allocation, and fail to simultaneously optimize multi-dimensional resources such as beam pointing, beamwidth, beam switching, and bandwidth allocation. In addition, the stochastic strategies employed by deep reinforcement learning-based beam configuration methods may lead to coverage gaps or inefficient beams, and they lack deterministic error correction mechanisms to ensure full coverage and load balancing.

[0007] 3. Existing methods based on multi-agent deep reinforcement learning, such as the multi-agent proximal policy optimization (MAPPO) algorithm under the centralized training and decentralized execution (CTDE) framework, although reducing the computational burden, generate a lot of overhead in power supply link data transmission and are inefficient when learning from latency information, making it difficult to achieve efficient online decision-making in large-scale flexible HTS systems.

[0008] 4. Existing technologies generally lack explicit control over flexible coverage and fail to control reserved antenna feed resources for incremental repair through beam switching, resulting in frequent global reconfiguration and the risk of service interruption. Summary of the Invention

[0009] To address the shortcomings of existing technologies, such as excessive computational complexity, limitation to small-scale systems, resource constraints, high overhead, inefficiency in learning from latency information, and failure to ensure service continuity while reducing global reconfiguration frequency, this invention provides a hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS. This invention can efficiently match multi-dimensional onboard resources with user needs, achieving load-balanced beam layout and fair bandwidth allocation.

[0010] The technical solution adopted by this invention to solve the technical problem is as follows:

[0011] This invention provides a hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS, which mainly includes the following steps:

[0012] Step 1: Construct the system model, coverage model, and communication model;

[0013] Step 2: Abstract the beam interaction into a beam overlap coupling matrix and a sidelobe gain coupling matrix to construct an inter-beam interference map;

[0014] Step 3: Based on the different time scales and physical characteristics of beam layout configuration and bandwidth allocation, the multidimensional resource optimization problem is decoupled into two independently optimizable sub-problems at large and small time scales, and corresponding constraints are constructed.

[0015] Step 4: By designing a PPO reinforcement learning beam layout agent based on a sequential collaborative decision-making framework, the large-scale sub-problems are remodeled into Markov decision process models.

[0016] Step 5: Based on the initial beam layout, a constraint-aware capacity-based clustering beam fine-tuning mechanism is used to perform deterministic refinement, identify and shut down inefficient beams, use the residual beams to cover missed users through constructive clustering methods, and output the final refined beam layout configuration.

[0017] Step 6: Based on the beam load variance increment and coverage index, the elastic coverage reconfiguration control logic adaptively selects to maintain the current beam layout configuration, trigger incremental repair, or trigger global reconfiguration operation to control the reconfiguration frequency of the beamformer.

[0018] Step 7: The fast fairness-satisfying bandwidth allocation algorithm decomposes the small-time-scale subproblem into three coupled stages: co-channel interference-aware frequency group partitioning, marginal utility-driven frequency group dimension adjustment, and sidelobe-aware frequency slot alignment, maximizing the fairness of demand satisfaction.

[0019] Furthermore, the large-timescale subproblem is formalized as follows:

[0020] ;

[0021] The constraints include:

[0022] (13a) Uniqueness of user connection ;

[0023] (13b) Beamwidth Feasible Domain ;

[0024] (13c) Switching binary constraint ;

[0025] (13d) Beam pointing uniqueness ;

[0026] (13e) Beam overlap penalty (soft constraint) ;

[0027] (13f) Empty coverage beam penalty (soft constraint) ;

[0028] in, For beam direction, For beamwidth selection, For switching matrix, For the beam in time slot t The direction it points to. For the beam in time slot t beamwidth, For the beam in time slot t Switching modes This represents the trade-off between coverage and load balancing. Let be the coverage of time slot t. For users With beam Related indicator variables, Let T be the set of time slots. For air users, Total number of air users, It is a flexible point beam set. The total number of flexible point beams, For inefficient beamforming, The Gaussian load balancing utility function is... For a set of selectable beamwidths, Q represents the total number of candidate beam positions. For decision-making instructions, Candidate beam position The coordinates.

[0029] Furthermore, the small-timescale subproblem is formalized as follows:

[0030] ;

[0031] The constraints include:

[0032] (15a) Total power constraint ;

[0033] (15b) Constraints on different frequency groups of adjacent beams ;

[0034] (15c) Single frequency group constraint ;

[0035] (15d) Resource Constraints ;

[0036] in, For beam-level bandwidth allocation, To ensure fairness in demand, Orthogonal frequency slots On the transmit power, This represents the total satellite launch power. For beam Does it belong to frequency group h at time t? and Beams The start and end orthogonal frequency slots, For frequency group Frequency resources.

[0037] Furthermore, the PPO reinforcement learning beam layout agent based on the sequential collaborative decision-making framework incorporates a Transformer encoder-based Actor network, a mask-constrained action probability adjustment mechanism, and a Critic network. The Transformer encoder-based Actor network captures the spatial dependencies between global candidate beam positions through multi-head self-attention; the mask-constrained action probability adjustment mechanism dynamically masks inactive actions; and the Critic network guides policy updates by evaluating state values.

[0038] Furthermore, the Markov decision process model is defined as follows:

[0039] state space : Includes candidate beam position occupancy status and potential traffic value status That is, the state space of BL-Agent ;

[0040] Action space : BL-Agent's action vector It is composed of the Cartesian product flattening of the beam direction and width space. For the beam in time slot t The direction it points to. For the beam in time slot t beamwidth;

[0041] Reward function: The composite reward consists of local and global components; local components Excitation single-beam load efficiency and constraints satisfy: when hour ;when hour ;otherwise Among them, the penalty constant , Represents the set of positive real numbers; global reward Compound rewards , This represents the trade-off between coverage and load balancing. The Gaussian load balancing utility function is... For a set of selectable beamwidths, It is a flexible point beam set. The total number of flexible point beams, For decision-making instructions, Candidate beam position coordinates Let be the coverage of time slot t. For users With beam The associated indicator variable, where T is the time slot, For air users, This represents the total number of users on the air.

[0042] Furthermore, the sequential collaborative decision-making process hierarchically divides the decision-making process into macrosteps and microsteps, with the macrosteps corresponding to physical time slot conversions. Further subdivided into a variable number of microsteps The index is Each microstep encapsulates the observation-decision-reward cycle of a specific beam.

[0043] Furthermore, the sequential cooperative decision-making transforms the joint action space into a solvable sequential Markov decision process through a probabilistic chain, ensuring the convergence of the BL-Agent; the objective is to find the policy network policy of the BL-Agent. maximize , To conform to the strategy A series of decisions Seeking expectations, As a discount factor, , For the reward function, This is the global state. For state space, For joint actions; sequential collaborative decision-making applies the probabilistic chain rule to decompose the joint strategy into... The generalized advantage is estimated as follows:

[0044] ;

[0045] ;

[0046] in, For Hongbu The The generalized advantage estimate is calculated in microsteps, where v is the smoothing parameter for the generalized advantage estimate, typically set to 0.95. For microsteps TD error (Temporal Difference Error) For microsteps TD error (Temporal Difference Error) The target value output by the Critic network. Here are the parameters of the Critic network, and the Critic loss function is... , To represent the function for finding the expectation within the macrostep, This is for estimating the actual value function within the macrostep.

[0047] Furthermore, in the constraint-aware capacity-based clustering beam fine-tuning mechanism, inefficient beam sets are first identified. Inefficient beamforming with empty coverage and inefficient beam sets with insufficient load and narrow beams The inefficient beam set is obtained by turning off inefficient beams, and the set of uncovered users is defined. Then, a constructive method is used to iteratively generate users covered by beams but not covered. In each iteration:

[0048] (1) Seed users are selected based on the maximum-minimum distance criterion to maximize spatial separation;

[0049] (2) cluster Start with seed users and gradually expand;

[0050] (3) The iteration terminates when all users are covered or The residual beam is turned off to preserve antenna feed resources as a strategic reserve for resilient coverage. This represents the total number of flexible point beams.

[0051] Furthermore, the elastic coverage reconstruction control logic is as follows:

[0052] (1) Maintain: When coverage Exceeding the preset coverage threshold And beam load variance increment Below the load variance increment threshold At the same time, maintain the current beam layout configuration to prioritize time stability;

[0053] (2) Incremental repair: when only coverage When the antenna feed resources are available and the current level is low, only the constraint-aware capacity clustering beam fine-tuning mechanism is triggered for incremental repair to avoid global reconstruction.

[0054] (3) Global Restructuring: When When the antenna feed is exhausted and coverage metrics are violated, global beam layout reconfiguration is triggered.

[0055] Furthermore, in the co-channel interference sensing frequency group partitioning stage: based on the beam overlap coupling matrix and inter-beam interference map, a vertex coloring strategy based on saturation sorting is used to assign adjacent beams to different frequency groups, thus obtaining the frequency groups. Beams in the middle;

[0056] The marginal utility-driven frequency group dimension adjustment phase dynamically adjusts the bandwidth of each frequency group to maximize global fairness utility.

[0057] The sidelobe sensing frequency slot alignment stage: in the same frequency group First, the beams are sorted in descending order of total sidelobe coupling strength, and the beams with the total sidelobe coupling strength are allocated preferentially. Among all feasible starting positions, the position that maximizes the weighted separation metric is selected.

[0058] The beneficial effects of this invention are:

[0059] 1. This invention proposes a method to construct a Markov Decision Process (MDP) guided by the objective function of the optimization problem. Based on the different time scales and physical characteristics of beam layout configuration and bandwidth allocation, the multidimensional resource optimization problem (beam pointing, beam width, beam switching, and bandwidth allocation) that is difficult to solve jointly is decomposed into two independently optimizable sub-problems through time scale decoupling, thereby realizing spatial domain beam layout configuration at a large time scale and frequency domain bandwidth allocation at a small time scale.

[0060] 2. This invention designs a PPO reinforcement learning beam placement agent (BL-Agent) based on the Sequential Cooperative Decision (SCD) framework, which integrates a Transformer encoder (TE)-based Actor network, a mask constraint action probability adjustment (MCAPA) mechanism, and a Critic network. The joint action space is transformed into a solvable Markov decision process model through probabilistic chain decomposition, realizing efficient online placement decision-making for hundreds of beams in a large-scale HTS system.

[0061] 3. This invention proposes a constraint-aware capacity clustering beam tuning (CACC-BR) mechanism as a deterministic second stage. It selects seed users through the maximum-minimum distance criterion and iteratively constructs capacity clusters to eliminate coverage holes and inefficient beams caused by the inherent randomness of deep reinforcement learning, ensuring full coverage and load balance, while closing redundant beams and reserving antenna feed resources.

[0062] 4. This invention designs an elastic coverage reconfiguration control logic. Based on the beam load variance increment and coverage index, it adaptively selects to maintain the current beam layout configuration, trigger incremental repair (only execute CACC-BR), or trigger global reconfiguration (execute the two-stage complete algorithm) operation, limiting global reconfiguration to about 15% of time slots, significantly reducing service interruption overhead, reducing reconfiguration frequency and service interruption.

[0063] 5. This invention proposes a Fast Satisfaction Fairness Bandwidth Allocation (RSF-BA) algorithm, which decomposes frequency domain resource allocation into three coupled stages: co-channel interference-aware frequency group partitioning, marginal utility-driven frequency group dimension adjustment, and sidelobe-aware frequency slot alignment. Based on beam layout, it achieves efficient allocation of frequency domain resources and maximizes the fairness of demand satisfaction. Attached Figure Description

[0064] Figure 1 The flowchart illustrates a layered beam coverage and bandwidth allocation method for large-scale flexible HTS provided by this invention.

[0065] Figure 2 A flowchart for refactoring the control logic for flexible coverage. Detailed Implementation

[0066] In a first aspect, the present invention provides a hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS.

[0067] like Figure 1 As shown, the hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS provided by this invention is mainly implemented by the following steps:

[0068] Step S1: System Model and Considered Scenarios;

[0069] Consider a large-scale flexible HTS system that serves There are three non-uniformly distributed dynamic air users, denoted as [a_n]. , The total number of users in the air is divided into time slot sets. Each time slot The duration of the equal length is Due to the duration of the time slot The environment is small enough that it is assumed to be quasi-static within each time slot, but dynamically changing between different time slots. Users in time slot t... The traffic demand is .

[0070] Large-scale flexible HTS systems employ flexible point beamforming. J represents the total number of flexible point beams, and each beam in time slot t represents... It is equipped with an adjustable pointing direction. Beamwidth Switch mode and bandwidth allocation To adapt to dynamic and non-uniformly distributed traffic demands.

[0071] Total downlink bandwidth resources of the system In each time slot Adaptively divided into non-overlapping and unequal-sized frequency groups (FGs) Under single-polarization conditions, frequency group frequency resources Divided into bandwidth The orthogonal frequency slots (FS), and .

[0072] Assigned to beam resources It cannot span multiple frequency groups, i.e. ,in and Beams The start and end orthogonal frequency slots, For time slots Active beam set, Indicator Beam Is it assigned to a frequency group? .

[0073] Spatially separated beams can reuse the same frequency group, but due to the presence of sidelobes, co-channel interference (CCI) may still exist between beams reusing the same orthogonal frequency slots, and adjacent beams need to be assigned to different frequency groups. Users share the frequency resources of the same beam through Time Division Multiple Access (TDMA).

[0074] Step S2: Construct the coverage model;

[0075] Using the satellite angular coordinate system, the beam is a perfect circle on the corresponding UV plane. Considering its location at altitude... The geostationary orbit HTS, whose geographical coordinates are determined by longitude and latitude Specify, i.e. Due to the satellite's orbital altitude Much higher than the aircraft's flight altitude, ignoring differences in aircraft flight altitude. (User) In the time slot The geographical coordinates are .

[0076] set up For Earth's radius, user The specific formula for calculating the geographical location coordinates of the satellite center in the Cartesian coordinate system is as follows:

[0077] (1);

[0078] (2);

[0079] (3);

[0080] The three axes of the Cartesian coordinate system at the satellite center are elevation, azimuth, and z-axis. These are the coordinate values ​​of the azimuth axis. These are the coordinate values ​​of the elevation axis. This represents the z-axis coordinate. A visual illustration can be found in the image below.

[0081] user The geographical coordinates on the UV plane are azimuth angle Angle of elevation Beam center coordinates The same calculation applies.

[0082] The service area is discretized into a set of candidate beam locations (BPs) of a uniform ground grid. Q represents the total number of candidate beam positions, and the candidate beam positions are... The coordinates are Beam The coverage configuration mainly includes:

[0083] (i) Beamwidth Determine the beam Coverage radius in the UV plane , A set of selectable beamwidths;

[0084] (ii) Beam switch mode ;

[0085] (iii) Beam pointing In the Large Timescale Beam Layout Algorithm (ACLB-BL), the first stage starts from... Selected (Decision Instructions) The second stage can be adjusted freely.

[0086] Using the minimum distance beam selection principle, the beam-user association vector is defined as follows:

[0087] (4);

[0088] (5);

[0089] in, For users With beam Related indicator variables, For users With beam The two-dimensional Euclidean distance centered in the UV plane. Beam. In the time slot The user set of the service is .

[0090] Define candidate beam position occupancy vector ,in .

[0091] Based on the unconnected user traffic demand at different beamwidths at each candidate beam location, a Potential Traffic Value (PTV) matrix is ​​defined. , :

[0092] (6);

[0093] in, For the potential traffic value (PTV) matrix elements, i.e., candidate beam positions q in the beamwidth The potential traffic value below; It is a vector consisting entirely of 1s; Beamwidth A defined coverage radius; For users in time slot t Traffic demand.

[0094] Step S3: Construct the communication model;

[0095] Because the aircraft cruises at a high altitude, modeling is performed from satellite via beam. To users In the quadrature frequency slot The dynamic line-of-sight (LoS) downlink channel is:

[0096] (7);

[0097] in, For the user's receiving antenna gain (assuming an omnidirectional antenna); This is a large-scale decay; The transmit antenna gain is modeled as follows:

[0098] (8);

[0099] (9);

[0100] in, For maximum transmit antenna gain, It is a first-order Bessel function of the first kind. It is a third-order Bessel function of the first kind. For users With beam The off-axis angle between them.

[0101] HTS employs a unified DBF architecture, following the system-level calibration method in 3GPP TR38.821. The SINR of each beam is approximated by the average SINR of the users served by that beam. A uniform power spectral density is used, based on the equivalent isotropic radiated power (EIRP). and maximum transmit antenna gain orthogonal frequency slots The transmit power is Then the user In the quadrature frequency slot The SINR on the above is:

[0102] (10);

[0103] in, For noise power density, Indicator Beam and beam Does the frequency overlap? To allocate to beam Resources To allocate to beam Resources.

[0104] Beam In the quadrature frequency slot Average spectral efficiency on beam In the time slot The user set of the service is Find the average:

[0105] (11);

[0106] in, For beam In the quadrature frequency slot Average spectral efficiency over [the beam]. In the time slot The channel capacity is .

[0107] Step S4: Spatial coupling between beams and interference;

[0108] To facilitate bandwidth allocation for co-channel interference sensing, beam interaction is abstracted into a beam overlap coupling matrix (OCM). And the sidelobe gain coupling matrix (GCM) .

[0109] For active beams : , For beam With beam The two-dimensional Euclidean distance centered on the UV plane. For beam The coverage radius of the UV plane, For beam The coverage radius in the UV plane. Based on this, an inter-beam interference map (IBIG) is constructed. edge set beam The degree is .

[0110] In the sidelobe gain coupling matrix , Let be the normalized transmit antenna gain function in equation (8). As a dimensionless coupling coefficient, it ranks the potential interference intensity between beams.

[0111] Step S5: Problem Modeling;

[0112] Based on the different time scales and physical characteristics of beam layout configuration and bandwidth allocation, the multidimensional resource optimization problem (beam pointing, beamwidth, beam switching, and bandwidth allocation), which is difficult to solve jointly, is decoupled into two sub-problems.

[0113] S501: Large timescale subproblem (P1);

[0114] Jointly determine beam pointing Beamwidth selection and switch matrix This aims to achieve long-term load balancing, meet full coverage requirements, and provide elastic coverage. The Gaussian load balancing utility function is defined as follows:

[0115] (12);

[0116] in, For beam The available rate at time t, The average load in time slot t, For deviation tolerance, For the reconstruction cost indicator function, Let t be the coverage rate of time slot t.

[0117] The large timescale subproblem (P1) is formalized as follows:

[0118] (13);

[0119] The constraints include:

[0120] (13a) Uniqueness of user connection ;

[0121] (13b) Beamwidth Feasible Domain ;

[0122] (13c) Switching binary constraint ;

[0123] (13d) Beam pointing uniqueness ;

[0124] (13e) Beam overlap penalty (soft constraint) ;

[0125] (13f) Empty coverage beam penalty (soft constraint) .

[0126] S502: Small timescale subproblems (P2);

[0127] Optimizing beam-level bandwidth allocation under co-channel interference and payload resource constraints To achieve fairness in meeting demand. Define the beam. Demand satisfaction rate ,in The fairness of demand fulfillment is quantified using the Jain Fairness Index as follows:

[0128] (14);

[0129] The small timescale subproblem (P2) is formalized as follows:

[0130] (15);

[0131] The constraints include:

[0132] (15a) Total power constraint , This refers to the total launch power of the satellite.

[0133] (15b) Constraints on different frequency groups of adjacent beams , For beam Does it belong to frequency group h at time t?

[0134] (15c) Single frequency group constraint ;

[0135] (15d) Resource Constraints .

[0136] Step S6: Large Timescale Beam Layout Algorithm (ACLB-BL) – First Stage;

[0137] Since the large-scale subproblem (P1) involves long-term cumulative objectives in a dynamic environment, it is remodeled as a Markov Decision Process (MDP). A PPO reinforcement learning beamforming agent (BL-Agent) based on the Sequential Cooperative Decision (SCD) framework is designed and deployed in an on-board digital beamformer. Under the SCD framework, the BL-Agent is sequentially configured as follows: Each beam generates pointing and width decisions, enabling each beam to acquire global state.

[0138] The PPO reinforcement learning beam layout agent (BL-Agent) based on the Sequential Cooperative Decision (SCD) framework incorporates a Transformer encoder, an Actor network, and a Critic network.

[0139] S601: The Markov Decision Process (MDP) model is defined as follows:

[0140] state space : Includes candidate beam position occupancy status and potential traffic value status That is, the state space of BL-Agent .

[0141] Action space : BL-Agent's action vector It is composed of the Cartesian product flattening of the beam direction and width space.

[0142] Reward function: The composite reward consists of local and global components. Local components Excitation single-beam load efficiency and constraints satisfy: when hour ;when hour ;otherwise Among them, the penalty constant , Represents the set of positive real numbers; global reward Compound rewards , This represents the trade-off between coverage and load balancing.

[0143] S602: The Sequential Cooperative Decision (SCD) framework is defined as follows:

[0144] Sequential collaborative decision-making hierarchically divides decisions into macrosteps and microsteps, making large-scale HTS optimization computationally feasible. Macrosteps correspond to physical time-slot transitions. Further subdivided into a variable number of microsteps (index is) Each microstep encapsulates the observation-decision-reward cycle for a specific beam.

[0145] In Microstep beam Invoke BL-Agent to interact with the environment, based on the current policy. Generate discrete classification distribution Sampling action and collect transfer , For beam In the first of Hongbu t The observation status in a microstep, For beam In the first of Hongbu t A tiny step, For beam In the first of Hongbu t The reward for each microstep. For beam In the first of Hongbu t The observation status of each microstep.

[0146] After a valid action is executed, the environment is virtually updated, and the state is refreshed in real time. ( For the first step of macro step t The global wave position occupancy matrix in microsteps For the first step of macro step t (The global beam rate requirement matrix for each microstep) and update the dynamic action mask to support subsequent beams in microsteps. The decision-making process involves sequential collaborative decision-making iteratively until all beam configurations are complete, yielding the macrostep trajectory. Invalid actions require the beam to be re-evaluated in subsequent microsteps until a valid action is achieved. , For beam In the first of Hongbu t The observation status in a microstep, For beam In the first of Hongbu t A tiny step, For beam In the first of Hongbu t The reward for each microstep. For beam In the first of Hongbu t The observation status of each microstep.

[0147] S603: Convergence of the SCD framework (Theorem 1);

[0148] Sequential cooperative decision-making transforms the joint action space into a solvable sequential Markov decision process through a probabilistic chain, ensuring the convergence of the BL-Agent. The objective is to find the policy network (parameterized as θ) of the BL-Agent. maximize ,in, To conform to the strategy A series of decisions Seeking expectations, As a discount factor, For the reward function, This is the global state. For joint actions. Sequential collaborative decision-making applies the probabilistic chain rule to decompose the joint strategy into... .

[0149] Generalized advantage estimation (GAE) is as follows:

[0150] (16);

[0151] (17);

[0152] in, For Hongbu The The generalized advantage estimate is calculated in microsteps, where v is the smoothing parameter for the generalized advantage estimate, typically set to 0.95. For microsteps TD error (Temporal Difference Error) For microsteps TD error (Temporal Difference Error) The target value output by the Critic network. Here are the parameters of the Critic network, and the Critic loss function is... , To represent the function for finding the expectation within the macrostep, This is for estimating the actual value function within the macrostep.

[0153] S604: Transformer encoder (TE) based Actor network (TE-Actor);

[0154] The state vector Tokenization into candidate beam position-level feature vector sequences Each token is embedded into a high-dimensional space using an MLP. , For the embedding dimension. After adding trainable positional encoding, input to the multi-head self-attention (MHSA) module.

[0155] For the Each layer's individual attention head computes the query, key, and value matrices separately:

[0156] (18);

[0157] (19);

[0158] (20);

[0159] in, , , These are query, key, and value weight matrices, respectively. For the number of attention heads, For input to the first The output of the layer above the MHSA module. Then the attention weights are:

[0160] (twenty one);

[0161] through After multi-head self-attention (MHSA), the data is aggregated into a global representation through attention pooling, and the MLP classifier is mapped to action logits. s is a simplified representation of the observed state (meaning that the output y is determined by the state s and the network parameters), and θ is the policy network parameter of the BL-Agent.

[0162] S605: Mask-Constrained Action Probability Adjustment (MCAPA) mechanism;

[0163] Constructing dynamic action masks using historical decision information:

[0164] (twenty two);

[0165] in, For the first step of macro step t The global position occupancy matrix in microsteps, and the penalty constant. , For Kronecker product, For length A vector of all 1s. This dynamic action mask is applied to logits before softmax, and the action probability distribution is:

[0166] (twenty three);

[0167] in, For the mask applied to action a, For the mask applied to action a', The neural network outputs a probability vector for action a. The neural network outputs a probability vector for action a'. When the penalty function... When the probability of an inability to act approaches 0, the softmax calculation range expands from... Reduced to a subset of effective actions .

[0168] The loss function for the Actor policy is:

[0169] (twenty four);

[0170] The cropping range is: Ratio of new to old strategies .

[0171] The computational complexity of the first stage is mainly driven by the forward propagation of the TE-Actor, which is... .

[0172] Step S7: Large Timescale Beam Layout Algorithm – Second Stage (CACC-BR);

[0173] Based on the initial beam layout in the first phase, a constraint-aware capacity-based clustering beam tuning (CACC-BR) mechanism is used to perform deterministic refinement. This identifies and shuts down inefficient beams, utilizes residual beams to cover missed users through constructive clustering, and outputs the refined final beam layout configuration. Seed users are selected using the maximum-minimum distance criterion, and capacity-based clustering is iteratively constructed to eliminate coverage holes and inefficient beams caused by the inherent randomness of DRL, ensuring full coverage and load balancing, while simultaneously disabling redundant beams and reserving antenna feed resources.

[0174] Specifically, the first step is to identify inefficient beam sets. Among them, empty coverage beam Narrow beam with insufficient load , , , Minimum beamwidth limitation.

[0175] Inefficient beams are turned off to obtain inefficient beam sets. Define the set of users not covered .

[0176] The second phase (CACC-BR) uses a constructive approach to iteratively generate beam coverage for uncovered and uncovered users. In each iteration:

[0177] (1) Seed user selection is based on the maximum-minimum distance criterion to maximize spatial separation: .

[0178] (2) cluster Starting with seed users and gradually expanding. Specifically, set... , , Clusters Center, beamwidth, and coverage radius. (User) Included in cluster Beamwidth constraint must be met simultaneously. , Maximum beamwidth limitation, load balancing constraint Spatial isolation constraints , This is the beam spacing constraint coefficient.

[0179] (3) The iteration terminates when all users are covered or The residual beam is turned off to preserve antenna feed resources as a strategic reserve for resilient coverage.

[0180] The computational complexity of the second stage (CACC-BR) is Because the first phase provided a near-optimal beam configuration, and Typically small in size and highly efficient in computation.

[0181] Step S8: Elastic Overlay Reconstruction Strategy;

[0182] Large temporal beam layout algorithm (ACLB-BL) adaptively determines reconstruction depth based on observable states, such as Figure 2 As shown, its decision-making logic is as follows:

[0183] (1) Retain: When coverage Exceeding the preset coverage threshold And beam load variance increment Below the load variance increment threshold At the same time, the current beam configuration will be maintained to prioritize time stability. , It is the variance function.

[0184] (2) Incremental Repair: When only coverage is covered Decreasing and having reserve antenna feed resources ( When this occurs, only the second phase (CACC-BR) is triggered for incremental repair, avoiding global reconstruction.

[0185] (3) Global Reconfiguration: When Or the antenna feed is exhausted. Furthermore, when the coverage metric is violated, the full execution of the first and second phases of the Large Timescale Beam Layout Algorithm (ACLB-BL) is triggered to reconstruct the global beam layout.

[0186] Based on the aforementioned flexible coverage reconstruction strategy, global reconstruction accounts for only about 15% of the time slots, incremental repair accounts for 57% of the time slots, and 28% of the time slots remain unchanged. Throughout the entire simulation period, the coverage rate remains close to 100%, and the beam load variance is maintained within a moderate range.

[0187] Step S9: Small Time Scale Bandwidth Allocation Algorithm (RSF-BA);

[0188] Given a load balancing beam layout The Small Time Scale Bandwidth Allocation Algorithm (RSF-BA) decomposes the small time scale subproblem (P2) into three coupled stages:

[0189] (1) Co-channel interference sensing frequency group partitioning;

[0190] Based on beam overlap coupling matrix Inter-beam interference diagram The vertex coloring strategy based on saturation sorting is used to assign adjacent beams to different frequency groups, eliminating hard co-channel interference within the frequency groups and satisfying constraint (15b), thus obtaining the frequency groups. beam in .

[0191] (2) Marginal utility drives frequency group dimension adjustment;

[0192] Dynamically adjust the bandwidth of each frequency group In order to maximize the utility of global fairness ,in, Let the frequency group be... The satisfaction rate ,in, For bandwidth The achievable rate, .

[0193] The frequency group dimension adjustment subproblem (P3) is as follows: Subject to constraints and , Let be the set of frequency slots assigned to beam k at time t.

[0194] Orthogonal frequency slots are allocated iteratively using marginal utility. Specifically, for each frequency group, each step... and candidate orthogonal frequency slots Evaluate marginal utility:

[0195] (25);

[0196] choose ,renew , ,in, Let be the set of remaining frequency slots that have not yet been assigned to any frequency group at time t. To allocate frequency groups, To be allocated to frequency slot, For frequency group A set of frequency slots.

[0197] (3) Sidelobe sensing frequency slot alignment (SAFSA);

[0198] In the same frequency group First, arrange them in descending order of total sidelobe coupling strength:

[0199] (26);

[0200] in, For beam The total sidelobe coupling strength. Beam priority allocation of the total sidelobe coupling strength, with its initial orthogonal frequency slot being... ( For beam The bandwidth requirement without co-channel interference is calculated using SNR. Among all feasible starting positions, the position that maximizes the weighted separation metric is selected, and its mathematical expression is as follows:

[0201] (27);

[0202] in, Let t be the starting frequency slot assigned to beam i at time t (this slot is within frequency group h). This refers to a set of beams for which orthogonal frequency slots have been assigned. This criterion explicitly forces a larger orthogonal frequency slot distance between beam pairs with strong sidelobe coupling.

[0203] The total complexity of the small timescale bandwidth allocation algorithm (RSF-BA) is: , H represents the starting frequency slot of frequency group h at time t, and H represents the number of frequency groups.

[0204] Secondly, the present invention provides an intelligent digital beamforming device.

[0205] The present invention provides an intelligent digital beamforming device deployed in the payload of a flexible high-throughput satellite, which mainly comprises the following modules:

[0206] (1) Environmental perception module: responsible for receiving and processing observation information such as ground user location and traffic demand, and constructing candidate beam position occupancy status. and potential traffic value status This forms the state input for the BL-Agent.

[0207] (2) PPO reinforcement learning beam layout agent (BL-Agent) based on the Sequential Cooperative Decision (SCD) framework: It incorporates an Actor network based on a Transformer encoder and an MLP-Critic network, which are sequentially configured under the SCD framework as follows: Each beam generation direction and width decision is made. The Actor network based on the Transformer encoder captures the spatial dependence between global candidate beam positions through multi-head self-attention (MHSA), and the Mask-Constrained Action Probability Adjustment (MCAPA) mechanism uses dynamic masking to shield inactive actions; the Critic network guides policy updates by evaluating state values.

[0208] (3) Beam Refinement Module: Execute the second phase (CACC-BR), identify and shut down inefficient beams, use the residual beams to cover missed users through constructive clustering, and output the final refined beam layout configuration. .

[0209] (4) Elastic Reconfiguration Control Module: Based on coverage Load variance increment and reserve resource status It adaptively determines to perform one of the following operations: hold, incremental repair, or global reconstruction, controlling the reconstruction frequency of the beamformer.

[0210] (5) Bandwidth allocation module: Executes the small timescale bandwidth allocation algorithm (RSF-BA) based on the beam overlap coupling matrix output by the beam layout. and sidelobe gain coupling matrix The process sequentially completes the partitioning of co-channel interference sensing frequency groups, dimensional adjustment of utility-driven frequency groups, and slot alignment of sidelobe sensing frequency groups, outputting the final bandwidth allocation scheme. .

[0211] (6) Beamforming execution module: Receives the configuration parameters of the above modules, generates corresponding beamforming weights through digital signal processing, controls the antenna array to form a specified number, direction and width of flexible beams, and maps the allocated bandwidth resources to each beam.

[0212] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS, characterized in that, Includes the following steps: Step 1: Construct the system model, coverage model, and communication model; Step 2: Abstract the beam interaction into a beam overlap coupling matrix and a sidelobe gain coupling matrix to construct an inter-beam interference map; Step 3: Based on the different time scales and physical characteristics of beam layout configuration and bandwidth allocation, the multidimensional resource optimization problem is decoupled into two independently optimizable sub-problems at large and small time scales, and corresponding constraints are constructed. Step 4: By designing a PPO reinforcement learning beam layout agent based on a sequential collaborative decision-making framework, the large-scale subproblem is remodeled into a Markov decision process model. Step 5: Based on the initial beam layout, a constraint-aware capacity-based clustering beam fine-tuning mechanism is used to perform deterministic refinement, identify and shut down inefficient beams, use the residual beams to cover missed users through constructive clustering methods, and output the final refined beam layout configuration. Step 6: Based on the beam load variance increment and coverage index, the elastic coverage reconfiguration control logic adaptively selects to maintain the current beam layout configuration, trigger incremental repair, or trigger global reconfiguration operation to control the reconfiguration frequency of the beamformer. Step 7: The fast fairness-satisfying bandwidth allocation algorithm decomposes the small-time-scale subproblem into three coupled stages: co-channel interference-aware frequency group partitioning, marginal utility-driven frequency group dimension adjustment, and sidelobe-aware frequency slot alignment, maximizing the fairness of demand satisfaction.

2. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The large-timescale subproblem is formalized as follows: ; The constraints include: (13a) Uniqueness of user connection ; (13b) Beamwidth Feasible Domain ; (13c) Switching binary constraint ; (13d) Beam pointing uniqueness ; (13e) Beam overlap penalty (soft constraint) ; (13f) Empty coverage beam penalty (soft constraint) ; in, For beam direction, For beamwidth selection, For switching matrix, For the beam in time slot t The direction it points to. For the beam in time slot t beamwidth, For the beam in time slot t Switching modes This represents the trade-off between coverage and load balancing. Let be the coverage of time slot t. For users With beam Related indicator variables, Let T be the set of time slots. For air users, Total number of air users, It is a flexible point beam set. The total number of flexible point beams, For inefficient beamforming, The Gaussian load balancing utility function is... For a set of selectable beamwidths, Q represents the total number of candidate beam positions. For decision-making instructions, Candidate beam position The coordinates.

3. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The small-timescale subproblem is formalized as follows: ; The constraints include: (15a) Total power constraint ; (15b) Constraints on different frequency groups of adjacent beams ; (15c) Single frequency group constraint ; (15d) Resource Constraints ; in, For beam-level bandwidth allocation, To ensure fairness in demand, Orthogonal frequency slots On the transmit power, This represents the total satellite launch power. For beam Does it belong to frequency group h at time t? and Beams The start and end orthogonal frequency slots, For frequency group Frequency resources.

4. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The PPO reinforcement learning beam layout agent based on the sequential collaborative decision-making framework incorporates a Transformer encoder-based Actor network, a mask-constrained action probability adjustment mechanism, and a Critic network. The Transformer encoder-based Actor network captures the spatial dependencies between global candidate beam positions through multi-head self-attention; the mask-constrained action probability adjustment mechanism uses dynamic masking to shield inactive actions; and the Critic network guides policy updates by evaluating state values.

5. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The Markov decision process model is defined as follows: state space : Includes candidate beam position occupancy status and potential traffic value status That is, the state space of BL-Agent ; Action space : BL-Agent's action vector It is composed of the Cartesian product flattening of the beam direction and width space. For the beam in time slot t The direction it points to. For the beam in time slot t beamwidth; Reward function: The composite reward consists of local and global components; local components Excitation single-beam load efficiency and constraints satisfy: when hour ;when hour ;otherwise Among them, the penalty constant , Represents the set of positive real numbers; global reward Compound rewards , This represents the trade-off between coverage and load balancing. The Gaussian load balancing utility function is... For a set of selectable beamwidths, It is a flexible point beam set. The total number of flexible point beams, For decision-making instructions, Candidate beam position coordinates Let be the coverage of time slot t. For users With beam The associated indicator variable, where T is the time slot, For air users, This represents the total number of users on the air.

6. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The sequential collaborative decision-making process hierarchically divides decisions into macrosteps and microsteps, with macrosteps corresponding to physical time slot conversions. Further subdivided into a variable number of microsteps The index is Each microstep encapsulates the observation-decision-reward cycle of a specific beam.

7. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The sequential cooperative decision-making process transforms the joint action space into a solvable sequential Markov decision process through a probabilistic chain, ensuring the convergence of the BL-Agent; the objective is to find the policy network policy of the BL-Agent. maximize , To conform to the strategy A series of decisions Seeking expectations, As a discount factor, , For the reward function, This is the global state. For state space, For joint actions; sequential collaborative decision-making applies the probabilistic chain rule to decompose the joint strategy into... The generalized advantage is estimated as follows: ; ; in, For Hongbu The The generalized advantage estimate is given by n microsteps, where v is the smoothing parameter for the generalized advantage estimate. For microsteps TD error, For microsteps TD error, The target value output by the Critic network. Here are the parameters of the Critic network, and the Critic loss function is... , To represent the function for finding the expectation within the macrostep, This is for estimating the actual value function within the macrostep.

8. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, In the constrained-aware capacity-based clustering beam fine-tuning mechanism, inefficient beam sets are first identified. Inefficient beamforming with empty coverage and inefficient beam sets with insufficient load and narrow beams The inefficient beam set is obtained by turning off inefficient beams, and the set of uncovered users is defined. Then, a constructive method is used to iteratively generate users covered by beams but not covered. In each iteration: (1) Seed users are selected based on the maximum-minimum distance criterion to maximize spatial separation; (2) cluster Start with seed users and gradually expand; (3) The iteration terminates when all users are covered or The residual beam is turned off to preserve antenna feed resources as a strategic reserve for resilient coverage. This represents the total number of flexible point beams.

9. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The elastic coverage reconstruction control logic is as follows: (1) Maintain: When coverage Exceeding the preset coverage threshold And beam load variance increment Below the load variance increment threshold At the same time, maintain the current beam layout configuration to prioritize time stability; (2) Incremental repair: when only coverage When the antenna feed resources are available and the current level is low, only the constraint-aware capacity clustering beam fine-tuning mechanism is triggered for incremental repair to avoid global reconstruction. (3) Global Restructuring: When When the antenna feed is exhausted and coverage metrics are violated, global beam layout reconfiguration is triggered.

10. The hierarchical beam coverage and bandwidth allocation method for large-scale flexible HTS according to claim 1, characterized in that, The co-channel interference sensing frequency group partitioning stage: Based on the beam overlap coupling matrix and inter-beam interference map, a vertex coloring strategy using saturation sorting is used to assign adjacent beams to different frequency groups, resulting in frequency groups. Beams in the middle; The marginal utility-driven frequency group dimension adjustment phase dynamically adjusts the bandwidth of each frequency group to maximize global fairness utility. The sidelobe sensing frequency slot alignment stage: in the same frequency group First, the beams are sorted in descending order of total sidelobe coupling strength, and the beams with the total sidelobe coupling strength are allocated preferentially. Among all feasible starting positions, the position that maximizes the weighted separation metric is selected.