Resource allocation optimization method for ground and non-ground networks
By using an active RIS-assisted NOMA downlink transmission network and combining drone trajectory and RIS phase shift optimization to optimize resource allocation, the problem of insufficient flexibility of drone base stations is solved, achieving optimal performance optimization for both terrestrial and non-terrestrial networks, and maximizing system speed and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, when non-terrestrial network platforms such as drones are used as base stations, their flexibility is not as good as that of reflective smart surfaces (RIS), making it difficult to achieve the best performance optimization of terrestrial and non-terrestrial networks, and lacking hybrid optimization methods.
An active RIS-assisted NOMA downlink transmission network is adopted, which combines UAV trajectory, RIS phase shift, power allocation and channel model. Through reinforcement learning, resource allocation is optimized, UAV trajectory and RIS configuration are optimized, interference is eliminated and the system rate is maximized.
Achieve maximum speed at optimal signal-to-noise ratio, eliminate interference from non-cooperative multi-point transmission base stations to edge users, and maximize user experience and network capacity.
Smart Images

Figure CN121793153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, specifically to a resource allocation optimization method for terrestrial and non-terrestrial networks. Background Technology
[0002] With the development of 6G communication technology, the integration of terrestrial and non-terrestrial (T-NT) communication networks has become a new trend. Non-terrestrial networks (NTNs) utilize satellites and high-altitude platforms (such as drones and stratospheric balloons) to achieve communication coverage, among which low-Earth orbit (LEO) satellites and geostationary orbit (GEO) satellites are important components.
[0003] Currently, when non-terrestrial network platforms such as drones are used as base stations, their flexibility is not as good as that of Reflective Smart Surfaces (RIS). RIS can dynamically configure reflection characteristics according to channel conditions to improve network capacity. However, existing research on RIS and drone systems is mostly focused on continuous or discrete action spaces, lacking hybrid optimization methods that combine the two, making it difficult to achieve optimal performance. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a resource allocation optimization method for terrestrial and non-terrestrial networks, solving the problems mentioned in the background section.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a resource allocation optimization method for terrestrial and non-terrestrial networks, comprising the following steps: S1. Scene Modeling: For a downlink transmission network with active RIS-assisted NOMA deployed by drones, distributed in M circular grids, each grid center is equipped with a single-antenna base station to serve central users and edge users, the system includes ground RIS fixed on buildings and airborne RIS located on drones, the drones hover in specific areas, their positions change with time slots and they avoid no-fly zones; S2, Channel Model: Line-of-sight (LOS) and non-line-of-sight (NLOS) paths were used, and Rayleigh and Rice fading models were combined to simulate the channel between the base station and the user. With users Between these two points, the channel in time slot t is represented as: S3, Active RIS Configuration: An active RIS is also known as a reactive RIS. An active RIS amplifies the incident signal using reflective amplifiers. These amplifiers are powered by the power supply of the active RIS components. The output signal of an active RIS with K reflective units can be expressed as: S4, NOMA power allocation: Each The system receives direct signals from the base station, indirect reflected links from the RIS, interference from base stations serving users in another area, and inherent noise from the active RIS. Ultimately, The received signal is represented as: S5. Optimization Objective: To maximize the sum of the system's rates over t time slots, there are three key control variables: the UAV trajectory is represented as... , defined as { }, RIS phase shift , defined as { [t], }, the power allocation factor Λ, is defined as { , m}, and the RIS amplification matrix P, are defined as { , The optimization objective for k} is as follows: S6, Solving using reinforcement learning: Based on the Markov decision process framework, the state space, action space, and reward function are defined. The PPO algorithm is used for iterative optimization. Combining discrete and continuous action spaces, the optimization strategy is estimated through the advantage function.
[0006] Furthermore, in step S1, the total flight time of the UAV is divided into t time slots, where t∈ ={1,..., Drones are prohibited from flying in no-fly zones, which are modeled as zones centered on various obstacles with a radius of [missing information]. A circular region, with obstacles o∈ ={1,2,...,O}.
[0007] Furthermore, in step S2, The reference path loss is at 1m. For large-scale loss, model as , where α is the path loss exponent. for and The distance between them, while small-scale loss uses Rayleigh fading with a coefficient of . ~ , It follows the Rayleigh distribution.
[0008] Furthermore, in step S3, It is a signal of expectation. It is dynamic noise. It is static noise. =diag( ,diag( ) performs a diagonalization operation on a matrix. It is the amplification matrix of the active RIS in time slot t.
[0009] Furthermore, each effective element of a RIS can be denoted as 1≤ ≤ ,in The maximum amplification capability that the component can provide. , representing the input signal of RIS, This represents the output signal of the RIS. =diag( ) , indicating that in time slot t, Indicates the magnification factor. , representing the phase shift of the k-th RIS element; Active RIS devices consume additional power and generate thermal noise when amplifying reflected signals. This noise consists of dynamic and static components. The amplitude of static noise is much smaller than that of dynamic noise and can be ignored. The dynamic noise variable is represented by v and modeled as follows: ,in Let represent a complex multivariate Gaussian distribution with mean µ and variance Σ, and parameters Σ. and Let them represent the K×K identity matrix and the K×1 zero vector, respectively; The base station serves both central and edge users, and its power is distributed between the two users, with edge users receiving a higher-power signal. The transmitted signal can be represented as: in and They are respectively and The received signal indicates that all base stations in the system have the same transmit power. , The power allocation factor is denoted as Power allocation factor Satisfies constraint 0.5 < The reason for the difference is that edge users are farther away from the base station and require higher power allocation.
[0010] Furthermore, in step S4, Signals indicating expectations This indicates interference between users. This indicates noise in an active RISC system. The user u represents Gaussian white noise with an expected value of 0 and a variance of . Gaussian distribution, from To users The channel is defined as: This includes direct and indirect links involving RIS. The subscript m usually represents a base station that directly transmits signals in this area, while j represents a base station in another area. The user receives signals broadcast by base stations in other areas.
[0011] Furthermore, the application of SIC (Superposition Coding) and continuous interference cancellation techniques enables multiple users to share the same time and frequency resources. for The achievable rate of the decoded signal is: in, = , , The decoding rate of its own signal is: The received signal is through and RIS (used above) The received (represented) data, taken together, is represented as: + in Represents the Hermitian transpose of the channel between RIS(R) and user e; The data rate can be expressed as: The combined rate of central users and edge users results in the following overall system rate: Regarding energy efficiency, a trade-off needs to be struck between network performance and energy costs. Energy utilization rate is expressed as follows: The total network power consumption includes the power consumption of the base station, active RIS, and drone. The total system speed is expressed as... The total power consumed in the network is expressed as: in This indicates the static power consumption and transmission power of the base station. RIS power consumption includes: for (R represents the amplification efficiency of RIS) for The input signal power of the k-th element for The phase shift of the k-th element of the panel, for The input signal of the kth element, for The circuit power consumption of the k-th element. This indicates the power consumed by the drone during hovering and movement.
[0012] Furthermore, in step S5, the constraints on the optimization objective are as follows: , , .
[0013] Furthermore, in step S6, the state space is defined to include the UAV position, power allocation factor, RIS amplification matrix, and total rate; the action space includes UAV movement, RIS phase shift, power allocation coefficient, and amplification matrix; and the reward function includes the total rate, distance incentive, and out-of-bounds penalty.
[0014] This invention provides a resource allocation optimization method for terrestrial and non-terrestrial networks, which has the following beneficial effects: 1. This resource allocation optimization method for terrestrial and non-terrestrial networks optimizes phase shift, UAV trajectory, NOMA power allocation factor and base station transmit power through reasonable modeling and reinforcement learning. It can achieve maximum data rate under optimal signal-to-noise ratio and optimize phase shift of terrestrial RIS to eliminate interference of non-cooperative multipoint transmission base stations to edge users, while maximizing user experience and network capacity at the edge. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the basic elements of a resource allocation optimization method for terrestrial and non-terrestrial networks according to the present invention. Detailed Implementation
[0016] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0017] like Figure 1 As shown, the present invention provides a technical solution: a resource allocation optimization method for terrestrial and non-terrestrial networks, comprising the following steps: S1. Scene Modeling: For a downlink transmission network with active RIS-assisted NOMA deployed by drones, distributed in M circular grids, each grid center is equipped with a single-antenna base station to serve central users and edge users, the system includes ground RIS fixed on buildings and airborne RIS located on drones, the drones hover in specific areas, their positions change with time slots and they avoid no-fly zones; Specifically, each grid m has a single-antenna base station at its center, where m ∈ ={1,2,...,M}, each in grid m Dual-user NOMA provides downlink channels to both center and edge users; The central user is located within the grid radius and is controlled by... Service, recorded as (c represents center), where c∈ ={1,2,...,C} represents the central user; Edge users are located outside the grid radius, denoted as (e stands for edge) by Provide services, where e∈ ={1,2,...,E} represents a marginal user; All users in the system can be represented as = ∪ Each user is denoted as u∈ They can only be peripheral users or central users; The total flight time of the drone is divided into t time slots, where t∈ ={1,..., Drones are prohibited from flying in no-fly zones, which are modeled as zones centered on various obstacles with a radius of [missing information]. A circular region, with obstacles o∈ ={1,2,...,O}; The system contains two active RIS: one fixed to all grids. (G stands for grid) On equidistant buildings, another drone is positioned. (U represents UAV, unmanned aerial vehicle) On the device, its position changes to attempt to provide the maximum total rate to users in the network. Both RIS serve all users in the network and are operated by a microcontroller to change the phase shift, R={ , }, equipped with k reflection units, k∈{1,2,...,K}; Assuming the signal power reflected multiple times by RIS is minimal, for all m∈ ,u∈ and o∈ Base station ,user and obstacles The positions (position, p) are respectively represented as , =( 0) and =( ),in and The heights of the base station and the obstacle are respectively. At time slot t, the position of the airborne RIS is represented as... Furthermore, it hovers over a specific area A, while the drone moves, changing its position in the xy plane while maintaining a fixed altitude. ; S2, Channel Model: Line-of-sight (LOS) and non-line-of-sight (NLOS) paths were used, and Rayleigh and Rice fading models were combined to simulate the channel between the base station and the user. With users Between these two points, the channel in time slot t is represented as: in The reference path loss is at 1m. For large-scale loss, model as , where α is the path loss exponent. for and The distance between them, while small-scale loss uses Rayleigh fading with a coefficient of . ~ , It is a Rayleigh distribution; S3, Active RIS Configuration: An active RIS is also known as a reactive RIS. An active RIS amplifies the incident signal using reflective amplifiers. These amplifiers are powered by the power supply of the active RIS components. The output signal of an active RIS with K reflective units can be expressed as: in It is a signal of expectation. It is dynamic noise. It is static noise. =diag( ,(diag( (This refers to performing a diagonalization operation on a matrix.) It is the amplification matrix of the active RIS in time slot t; Each effective element of a RIS can be represented as 1≤ ≤ ,in The maximum amplification capability that the component can provide. , representing the input signal of RIS, This represents the output signal of the RIS. =diag( ) , indicating that in time slot t, Indicates the magnification factor. , representing the phase shift of the k-th RIS element; Active RIS devices consume additional power and generate thermal noise when amplifying reflected signals. This noise consists of dynamic and static components. The amplitude of static noise is much smaller than that of dynamic noise and can be ignored. The dynamic noise variable is represented by v and modeled as follows: ,in Let represent a complex multivariate Gaussian distribution with mean µ and variance Σ, and parameters Σ. and Let them represent the K×K identity matrix and the K×1 zero vector, respectively; The base station serves both central and edge users, and its power is distributed between the two users, with edge users receiving a higher-power signal. The transmitted signal can be represented as: in and They are respectively and The received signal indicates that all base stations in the system have the same transmit power. , The power allocation factor is denoted as Power allocation factor Satisfies constraint 0.5 < The reason for the value being <1 is that edge users are farther away from the base station and require higher power allocation; S4, NOMA power allocation: Each The system receives direct signals from the base station, indirect reflected links from the RIS, interference from base stations serving users in another area, and inherent noise from the active RIS. Ultimately, The received signal is represented as: in Signals indicating expectations This indicates interference between users. This indicates noise in an active RISC system. The user u represents Gaussian white noise with an expected value of 0 and a variance of . Gaussian distribution, from To users The channel is defined as: This includes direct and indirect links involving RIS. The subscript m usually represents a base station that directly transmits signals in this area, while j represents a base station in another area. The user receives signals broadcast by base stations in other areas. The application of SIC (Superposition Coding) and continuous interference cancellation technology enables multiple users to share the same time and frequency resources. for The achievable rate of the decoded signal is: in, = , , The decoding rate of its own signal is: The received signal is through and RIS (used above) The received (represented) data, taken together, is represented as: + in Represents the Hermitian transpose of the channel between RIS(R) and user e; The data rate can be expressed as: The combined rate of central users and edge users results in the following overall system rate: Regarding energy efficiency, a trade-off needs to be struck between network performance and energy costs. Energy utilization rate is expressed as follows: The total network power consumption includes the power consumption of the base station, active RIS, and drone. The total system speed is expressed as... The total power consumed in the network is expressed as: in This indicates the static power consumption and transmission power of the base station. RIS power consumption includes: for (R represents the amplification efficiency of RIS) for The input signal power of the k-th element for The phase shift of the k-th element of the panel, for The input signal of the kth element, for The circuit power consumption of the k-th element. This indicates the power consumed by the drone during hovering and movement; S5. Optimization Objective: To maximize the sum of the system's rates over t time slots, there are three key control variables: the UAV trajectory is represented as... , defined as { }, RIS phase shift , defined as { [t], }, the power allocation factor Λ, is defined as { , m}, and the RIS amplification matrix P, are defined as { , The optimization objective for k} is as follows: Constraints: , , S6, Solving using reinforcement learning: Based on the Markov Decision Process (MDP) framework, the state space (including UAV position, power allocation factor, RIS amplification matrix and total rate), action space (including UAV movement, RIS phase shift, power allocation coefficient, amplification matrix) and reward function (including total rate, distance incentive and out-of-bounds penalty) are defined. The PPO algorithm is used for iterative optimization. Combining discrete and continuous action spaces, the optimization strategy is estimated through the advantage function. Specifically, we introduce MDP, which is defined by tuples (S, A, P, R, D). S (status) and A (action) represent the state space and action space, respectively; R (reward) is the reward function; D is a discount factor balancing the weights of current and future rewards, a value between 0 and 1. A discount factor close to 1 indicates that the agent values future rewards almost as much as immediate rewards, while a discount factor close to 0 indicates that the agent prioritizes immediate rewards; P (probability) represents the state transition probability. In each time slot t, the agent observes the current state. Choose an action based on its strategy. and transition to a new state. And receive the corresponding rewards; Spatial State: The state space in time slot t consists of the UAV's current position, the base station power allocation factor, the active RIS amplification matrix, and the total rate, respectively represented as: [t]、ζ、 [t] and R[t] = { [t]、 [t]、 The state space of {c,e} is expressed as: Action space: It consists of the UAV's movement in the xy plane, the amplification factor and phase shift of the active RIS, and the power allocation factor of the base station. The action space in time slot t contains the UAV's actions. {( 1,0),(1,0),(0, Let 1), (0,1), (0,0)} represent left, right, down, up, and hover respectively, with phase shift as... [t]={ [t], k}, power allocation coefficient ={ , m} and the amplification matrix of active RIS [t]={ , k}, the action space is represented as: The reward function ensures the maximum sum rate by punishing the behavior of the UAV going out of bounds, while ensuring the safety of the UAV and meeting the quality of service requirements. The reward function is defined as: where ζ[t]= ( (<threshold), where threshold is the threshold value, As an indicator function, ζ[t] is used to represent the distance incentive for keeping the UAV close to the user, where represents the distance between the UAV and the user. represents when goes out of the grid boundary, and the penalty coefficient given can be a constant. The indicator function represents whether it goes out of bounds, and C is a constant; Iterative method: Discrete (denoted as d) and continuous (denoted as c) are adopted, and the policy is denoted as Specifically as follows: Initialization parameters: Iterate for l = 1, 2,... L rounds: Obtain the initial state For time slots t = 0, 1,... T, run: Select continuous actions For each element n of the RIS, run: Energy efficiency optimization Based on the channel condition, update the gain of the k-th RIS element End the run for the RIS element Execute the action Calculate the reward according to the total rate after the action is completed Store the transition tuple ( ) in the cache End the execution of the epoch. End iteration Update strategy parameters , ; Using the value function V( (Representing the predictive value of a state) Calculate the advantage function estimate To optimize the strategy, for continuous actions, the usual Gaussian distribution prediction and sampling method can be used to convert them into discrete actions, and the function is defined as: in = Intuitively speaking, it's the probability ratio between the old and new strategies. It is a constant, typically ranging from 0.1 to 0.3, used to limit... Deviation range, Expressing expectations; Generalized advantage estimation (GAE) is used to calculate .
[0018] in V( )-V( ),if A positive value indicates that the action was better than expected, and the strategy update did not exceed [a certain threshold]. Otherwise, the advantage is negative, the action is worse than expected, and the strategy update should be no less than [amount missing]. , It is the discount factor, a constant, and T is the length of the episode, the total number of time slots.
[0019] Based on the above description, this invention optimizes phase shift, UAV trajectory, NOMA power allocation factor and base station transmit power through reasonable modeling and a reinforcement learning method. It can achieve maximum data rate at optimal signal-to-noise ratio and optimize phase shift of ground RIS to eliminate interference of non-cooperative multipoint transmission base stations to edge users, while maximizing user experience and network capacity at the edge.
[0020] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A resource allocation optimization method for terrestrial and non-terrestrial networks, characterized in that: The process includes the following steps: S1. Scene Modeling: For a downlink transmission network with active RIS-assisted NOMA deployed by drones, distributed in M circular grids, each grid center is equipped with a single-antenna base station to serve central users and edge users, the system includes ground RIS fixed on buildings and airborne RIS located on drones, the drones hover in specific areas, their positions change with time slots and they avoid no-fly zones; S2, Channel Model: Using line-of-sight and non-line-of-sight paths, and combining Rayleigh fading and Rice fading models, the channel between the base station and the user is simulated. With users Between these two points, the channel in time slot t is represented as: S3, Active RIS Configuration: An active RIS is also known as a reactive RIS. An active RIS amplifies the incident signal using reflective amplifiers. These amplifiers are powered by the power supply of the active RIS components. The output signal of an active RIS with K reflective units can be expressed as: S4, NOMA power allocation: Each The system receives direct signals from the base station, indirect reflected links from the RIS, interference from base stations serving users in another area, and inherent noise from the active RIS. Ultimately, The received signal is represented as: S5. Optimization Objective: To maximize the sum of the system's rates over t time slots, there are three key control variables: the UAV trajectory is represented as... , defined as { }, RIS phase shift , defined as { [t], }, the power allocation factor Λ, is defined as { , m}, and the RIS amplification matrix P, are defined as { , The optimization objective for k} is as follows: S6, Solving using reinforcement learning: Based on the Markov decision process framework, the state space, action space, and reward function are defined. The PPO algorithm is used for iterative optimization. Combining discrete and continuous action spaces, the optimization strategy is estimated through the advantage function.
2. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S1, the total flight time of the UAV is divided into t time slots, where t∈ ={1,..., Drones are prohibited from flying in no-fly zones, which are modeled as zones centered on various obstacles with a radius of [missing information]. A circular region, with obstacles o∈ ={1,2,...,O}.
3. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S2 The reference path loss is at 1m. For large-scale loss, modeled as , where α is the path loss exponent. for and The distance between them, while small-scale loss uses Rayleigh fading with a coefficient of . ~ , It follows the Rayleigh distribution.
4. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S3 It is a signal of expectation. It is dynamic noise. It is static noise. =diag( ,diag( ) performs a diagonalization operation on a matrix. It is the amplification matrix of the active RIS in time slot t.
5. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 4, characterized in that: Each effective element of a RIS can be represented as 1≤ ≤ ,in The maximum amplification capability that the component can provide. , representing the input signal of RIS, This represents the output signal of the RIS. =diag( ) , indicating that in time slot t, Indicates the magnification factor. , representing the phase shift of the k-th RIS element; Active RIS devices consume additional power and generate thermal noise when amplifying reflected signals. This noise consists of dynamic and static components. The amplitude of static noise is much smaller than that of dynamic noise and can be ignored. The dynamic noise variable is represented by v and modeled as follows: ,in Let represent a complex multivariate Gaussian distribution with mean µ and variance Σ, and parameters Σ. and Let them represent the K×K identity matrix and the K×1 zero vector, respectively; The base station serves both central and edge users, and its power is distributed between the two users, with edge users receiving a higher-power signal. The transmitted signal can be represented as: in and They are respectively and The received signal indicates that all base stations in the system have the same transmit power. , The power allocation factor is denoted as Power allocation factor Satisfies constraint 0.5 < The reason for the difference is that edge users are farther away from the base station and require higher power allocation.
6. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S4 Signals indicating expectations This indicates interference between users. This indicates noise in an active RISC system. The user u represents Gaussian white noise with an expected value of 0 and a variance of . Gaussian distribution, from To users The channel is defined as: This includes direct and indirect links involving RIS. The subscript m usually represents a base station that directly transmits signals in this area, while j represents a base station in another area. The user receives signals broadcast by base stations in other areas.
7. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 6, characterized in that: The application of SIC (Superposition Coding) and continuous interference cancellation technology enables multiple users to share the same time and frequency resources. for The achievable rate of the decoded signal is: in, = , , The decoding rate of its own signal is: The received signal is through and RIS (used above) The received (represented) data, taken together, is represented as: + in Represents the Hermitian transpose of the channel between RIS(R) and user e; The data rate can be expressed as: The combined rate of central users and edge users results in the following overall system rate: Regarding energy efficiency, a trade-off needs to be struck between network performance and energy costs. Energy utilization rate is expressed as follows: The total network power consumption includes the power consumption of the base station, active RIS, and drone. The total system speed is expressed as... The total power consumed in the network is expressed as: in This indicates the static power consumption and transmission power of the base station. RIS power consumption includes: for (R represents the amplification efficiency of RIS) for The input signal power of the k-th element for The phase shift of the k-th element of the panel, for The input signal of the kth element, for The circuit power consumption of the k-th element. This indicates the power consumed by the drone during hovering and movement.
8. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S5, the constraints on the optimization objective are as follows: , , 。 9. The resource allocation optimization method for terrestrial and non-terrestrial networks according to claim 1, characterized in that: In step S6, the state space is defined to include the UAV position, power allocation factor, RIS amplification matrix, and total rate; the action space includes UAV movement, RIS phase shift, power allocation coefficient, and amplification matrix; and the reward function includes the total rate, distance incentive, and out-of-bounds penalty.