Resource allocation method and allocation system for satellite-ground downlink, electronic equipment and medium

By jointly optimizing RIS's passive beamforming, satellite power allocation and NOMA user pairing, the problem of low resource allocation efficiency in satellite downlink NOMA communication network is solved, and higher data rates and longer service time are achieved.

CN120567237AActive Publication Date: 2025-08-29CHINA SATELLITE NETWORK EXPLORATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511044997.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-08-29
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

The prior art has not fully studied the satellite downlink NOMA communication network in satellite communication, and has not considered transmission power and user pairing optimization, resulting in low resource allocation efficiency and cannot provide resource allocation to ground users in real time and accurately.

Method used

By jointly optimizing passive beamforming of reconstructible intelligent surfaces (RIS), power allocation of satellites and NOMA user pairing, the effective data rate of the downlink NOMA system is maximized, and the resource allocation is used using deep reinforcement learning (DRL) method.

Benefits of technology

The average data rate and continuous service time of the satellite downlink are improved, and the efficiency and accuracy of resource allocation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567237A_ABST
    Figure CN120567237A_ABST
Patent Text Reader

Abstract

The invention discloses a satellite-ground downlink resource allocation method and system, electronic equipment and a medium. The method comprises the following steps: determining an agent state based on equivalent channels from a satellite to all users; obtaining power distribution in the user pair according to the power distribution proportionality coefficient of the user pair; the first agent obtains a user pairing matrix according to the decision factor output by the second agent; and maximizing the accumulated rewards corresponding to the first agent and the second agent to obtain the phase of each reflection unit in the optimal reconfigurable intelligent surface, the power distribution of the satellite and the user pairing matrix. According to the invention, resources can be effectively allocated to the user and the RIS in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of satellite communication technology, and in particular to a resource allocation method, allocation system, electronic equipment and medium for a satellite-to-ground downlink. Background Art

[0002] In satellite communications, how to effectively allocate onboard resources has become a key component of satellite resource management. As satellite resources become increasingly diverse, traditional satellite resource allocation methods are inefficient and unable to accurately allocate resources to ground users or related devices in real time.

[0003] Existing technical solutions often use traditional optimization methods such as fractional programming, semidefinite relaxation, and continuous convex approximation to optimize beamforming for large-scale satellite arrays and smart reflective metasurfaces to improve communication and perception performance. However, these methods do not consider the use of NOMA networks, do not address NOMA user grouping, and do not optimize transmit power. Furthermore, these optimization methods are relatively complex.

[0004] Satellite downlink NOMA communication networks have not been fully explored in existing technologies. Optimization of satellite transmit power and user pairing, two parameters crucial for improving NOMA network speeds, has not been considered. Furthermore, existing optimization methods are either heuristic or complex. Summary of the Invention

[0005] In view of this, the present application provides a resource allocation method, allocation system, electronic device and medium for a satellite-to-ground downlink, which comprehensively considers the joint optimization of the passive beamforming of RIS, the power allocation of the satellite and the NOMA user pairing in each time slot, that is, maximizes the effective data rate of the downlink NOMA system for the entire continuous service time, and obtains the optimal resource allocation result.

[0006] The present application discloses a method for allocating resources of a satellite-to-ground downlink, which includes: Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; Obtaining a power allocation within the user pair according to the power allocation ratio coefficient of the user pair; wherein the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; The first agent obtains a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results; The cumulative rewards corresponding to the first agent and the second agent are maximized to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the reward is associated with the effective data rate corresponding to all the users.

[0007] Furthermore, the determining of the agent state based on the equivalent channels from the satellite to all users includes: The state used by the agent in the current training time step is defined as the equivalent channel of the signal transmitted by the satellite to each of the users in the previous time slot; the training time step corresponds to the time slot; and a single communication cycle of the satellite is evenly divided to obtain multiple time slots.

[0008] Furthermore, before obtaining the power allocation within the user pair according to the power allocation ratio coefficient of the user pair, the method further includes: Taking the phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair as the action space of the first intelligent agent; Taking all possible values ​​of the decision factor corresponding to the user pairing matrix as the action space of the second agent; Rewards are set for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot.

[0009] Furthermore, the step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If the signal-to-interference-and-noise ratio (SINR) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is greater than a preset threshold, rewards are set for the first agent and the second agent in the current training time step according to the effective data rate corresponding to all the users in the current time slot; wherein the signal transmitted by the satellite received by each user is the signal transmitted by the satellite and transmitted through the equivalent channel to each user.

[0010] Furthermore, the step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If a signal-to-interference-plus-noise ratio (SIN) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is at least partially less than a preset threshold, a penalty coefficient is set based on the SIN corresponding to the signal transmitted by the satellite received by each user, and a reward is set for the first agent and the second agent in the current training time step based on the penalty coefficient and the effective data rate corresponding to all the users in the current time slot; the penalty coefficient is set based on the difference between the preset threshold and the SIN below the preset threshold; wherein the signal transmitted by the satellite received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user.

[0011] Furthermore, before the first agent obtains the user pairing matrix based on the decision factor output by the second agent, the method further includes: The second agent trains all possible user pairing results and outputs the decision factors; each decision factor corresponds to each user pairing result one by one.

[0012] Furthermore, obtaining the power allocation within the user pair according to the power allocation ratio coefficient of the user pair includes: determining a total power allocated by the satellite to each of the user pairs; According to the total power of each user pair and the power allocation ratio coefficient of each user pair, the power corresponding to each user in each user pair is obtained, that is, the power allocated by the satellite to each user in each user pair.

[0013] Furthermore, it also includes: In the current training time step, after the second neural network corresponding to the second agent calculates the reward corresponding to the action it performs, the second agent transitions to the next state; The second neural network corresponding to the second agent is trained and its network parameters are updated using the state corresponding to the second agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the second neural network is used to instruct the second agent to select and execute the corresponding action.

[0014] Furthermore, it also includes: In the current training time step, after the first neural network corresponding to the first agent calculates the reward corresponding to the action performed by the first agent, the first agent transitions to the next state; The first neural network corresponding to the first agent is trained and its network parameters are updated using the state corresponding to the first agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the first neural network is used to instruct the first agent to select and execute a corresponding action.

[0015] Furthermore, maximizing the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix includes: In the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

[0016] Furthermore, the method for obtaining the equivalent channel includes: obtaining the equivalent channel of the signal transmitted by the satellite to each of the users according to the reflection unit in the reconfigurable smart surface; The step of obtaining the equivalent channel of the signal transmitted by the satellite to each user according to the reflection unit in the reconfigurable smart surface comprises: Taking any one reflective unit in the reconfigurable smart surface as a reference, and obtaining a first component of a path loss of a link from the satellite to the reconfigurable smart surface according to a propagation direction of a signal transmitted from the satellite to the reconfigurable smart surface; Obtaining a second component of the path loss of a link from the reconfigurable smart surface to the user according to a propagation direction of a signal from the reconfigurable smart surface to the user after the signal transmitted by the satellite is reflected by the reconfigurable smart surface; The equivalent channel of the signal transmitted by the satellite to each of the users is obtained according to the first component, the distances between the first component and the satellite, the reconfigurable smart surface and the users.

[0017] Furthermore, obtaining the equivalent channel of the signal transmitted by the satellite to each of the users based on the first component, the first component, and the distances among the satellite, the reconfigurable smart surface, and the user includes: Obtaining a channel from the satellite to the reconfigurable smart surface according to the first component and a distance between the satellite and the reconfigurable smart surface; Obtaining a channel from the reconfigurable smart surface to the user according to the second component and a distance between the reconfigurable smart surface and the user; Obtaining a channel from the satellite to the user according to a distance between the satellite and the user; The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the channel from the satellite to the reconfigurable smart surface, the channel from the reconfigurable smart surface to the user, and the channel from the satellite to the user.

[0018] Furthermore, the method for obtaining the effective data rate includes: obtaining the effective data rate of the signal transmitted by the satellite to each of the users based on the equivalent channel of the signal transmitted by the satellite to each of the users; The obtaining, based on the equivalent channel of the signal transmitted by the satellite to each user, the effective data rate at which the satellite transmits the signal to each user, comprises: Obtaining a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user; the signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user; An effective data rate at which the satellite transmits data to the user is obtained according to a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

[0019] Furthermore, all users within the coverage of the satellite are divided into multiple user pairs; the pairing results of the multiple user pairs constitute the user pairing matrix; each user pair includes multiple users; and users in the same user pair use the same time and frequency band.

[0020] The present application also discloses a satellite-to-ground downlink resource allocation system, which includes: A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; A power allocation module, configured to obtain a power allocation within a user pair based on a power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; A user pairing module, configured for the first agent to obtain a user pairing matrix based on the decision factors output by the second agent; the user pairing matrix is ​​used to store all user pairing results; A maximization module is configured to maximize the cumulative rewards corresponding to the first agent and the second agent to obtain an optimal phase of each reflective unit in the reconfigurable smart surface, power allocation of the satellite, and user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all the users.

[0021] The present application also discloses an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the resource allocation method for the satellite-to-ground downlink described in any one of the above items is implemented.

[0022] The present application also discloses a computer-readable storage medium, which includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes any one of the above-mentioned satellite-to-ground downlink resource allocation methods.

[0023] Due to the adoption of the above technical solution, the present application has the following advantages: the present application maximizes the average sum rate of the entire continuous service time of the downlink NOMA system by jointly optimizing the passive beamforming of the RIS, the power allocation of the satellite and the NOMA user pairing in each time slot, thereby improving the average sum rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0025] Figure 1 This is a schematic diagram of RIS-assisted downlink non-orthogonal multiple access in an urban scenario according to an embodiment of the present application; Figure 2 This is a flow chart of a method for allocating resources for a satellite-to-ground downlink according to an embodiment of the present application; Figure 3 This is a schematic diagram of the cumulative reward convergence of an embodiment of the present application; Figure 4 Schematic diagram showing the effect of the number of RIS reflection units on the average sum rate according to an embodiment of the present application; Figure 5 This is a schematic diagram of the impact of system bandwidth on average sum rate in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The present application is further described with reference to the accompanying drawings and embodiments. The embodiments described are only a part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0027] A RIS-assisted downlink Non-Orthogonal Multiple Access (NOMA) system in an urban setting is considered. Both satellite-to-user and satellite-to-RIS-to-user links are considered. The average downlink data rate over the entire service duration of the NOMA system is maximized by jointly optimizing the RIS passive beamforming, satellite power allocation, and NOMA user pairing in each time slot. This joint optimization problem involves both high-dimensional continuous and discrete optimization variables and is a long-term effective data rate maximization problem. This problem requires reconfiguring the phase of each reflector in the RIS, the satellite power allocation, and the NOMA user pairing in each time slot, making it difficult to solve using traditional optimization methods.

[0028] The embodiments of the present application provide a resource allocation method, allocation system, electronic device, and medium for a satellite-to-ground downlink, which integrates reconfigurable intelligent surface-assisted satellite downlink communication and non-orthogonal multiple access technology. By jointly optimizing the passive beamforming of the RIS in each time slot, the power allocation of the satellite, and the NOMA user pairing, the average sum rate of the downlink NOMA system over the entire continuous service time (the average value of the sum of the effective data rates of all users corresponding to all time slots in a single satellite data transmission cycle, i.e., the formula below) is maximized. , t is the time slot, T is the total number of time slots, K is the total number of user pairs, k is the sequence number of user pairs, and n is the user sequence number in each user pair. is the actual intra-pair power allocation).

[0029] Due to environmental factors and other factors, the channel quality of direct communication links between satellites and ground users can be poor. To enhance the user experience and increase the average and rate of the system during continuous service, a reconfigurable intelligent surface (RIS) can be used to assist in satellite-to-ground communication services. RIS consists of a large number of low-cost passive reflectors, each of which can be individually controlled. By dynamically adjusting the phase and amplitude of the incident signal, RIS achieves precise signal control and optimization, thereby improving link reliability and spectral efficiency.

[0030] As shown in Figure 1, a RIS-assisted downlink non-orthogonal multiple access system in an urban scenario is considered. For example, in a three-dimensional Cartesian coordinate system, a The satellite provides communication services to 2K single-antenna ground users. are the coordinates of the satellite on the x-axis, y-axis, and z-axis respectively. The RIS provides users with auxiliary indirect link communication services. are the coordinate values ​​of RIS on the x-axis, y-axis, and z-axis respectively. reflection units, each column contains reflection units, each row contains Reflection units, each unit is half a wavelength apart in both horizontal and vertical directions , the reflection unit set is expressed as 2K ground users are paired to form K pairs of NOMA user pairs, and the user pair set is represented as , users in the same user pair use the same time / frequency resources. The nth user in the kth user pair is represented as ,in . The coordinates are expressed as .

[0031] The optimization problem involves both high-dimensional continuous and discrete optimization variables and is a long-term effective data rate maximization problem. This problem requires reconfiguring the phase of each reflector unit in the RIS, the satellite's power allocation, and the NOMA user pairing in each time slot, making it difficult to solve using traditional optimization methods. To provide long-term, stable communication services to ground users, an embodiment of the present application provides a resource allocation method for a satellite-to-ground downlink. This method considers maximizing the average sum rate by jointly optimizing the passive beamforming of the RIS, the satellite's power allocation, and the NOMA user pairing within a specific transmission period (e.g., the period during which the satellite transmits data). This optimization problem is solved using a deep reinforcement learning (DRL) method that optimizes hybrid actions. Among currently popular DRL algorithms, the PPO algorithm can optimize continuous actions, while the DQN algorithm can optimize discrete actions. A hybrid DRL framework, the PPO-DQN algorithm, can be employed. This framework constructs an MDP (Markov Decision Process) for the primary and mirror environments, thereby merging the two agents and achieving joint hybrid action control.

[0032] The embodiments of the present application can be applied to Figure 1 In the application scenario shown, see Figure 2 The satellite-to-ground downlink resource allocation method may include the following steps: Step 201: Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent.

[0033] In one embodiment of the present application, the state used by the agent in the current training time step is defined as the equivalent channel of the signal transmitted by the satellite to each user in the previous time slot; the training time step corresponds to the time slot; and a single communication cycle of the satellite is evenly divided to obtain multiple time slots.

[0034] For example, the agent rationally plans the RIS reflection phase, satellite power allocation, and NOMA user pairing based on the channel conditions of the users in the network. Therefore, the equivalent channel from the satellite to all users is described as the agent state. The state space can be expressed as , There are 2K dimensions in total, and each dimension is continuous. is the equivalent channel of the 2Kth ground user. The state of the PPO agent is expressed as , the DQN agent state is represented as , and satisfies The state used in time slot t, i.e., time step t in the MDP Defined as The equivalent channel calculated by the time slot is thus expressed as , Indicates the equivalent channel of the 2Kth terrestrial user at time slot t-1. Directly set to .

[0035] Step 202: Obtain the power allocation within the user pair according to the power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair.

[0036] In one embodiment of the present application, obtaining power allocation within a user pair based on the power allocation ratio coefficient of the user pair includes: determining the total power allocated by the satellite to each user pair; obtaining the power corresponding to each user in each user pair based on the total power of each user pair and the power allocation ratio coefficient of each user pair, that is, the power allocated by the satellite to each user in each user pair; Specifically, the goal of the embodiment of the present application is to optimize the phase of RIS by joint , Satellite power distribution and NOMA user pairing matrix To maximize the average sum rate of the system during the entire service duration. Since the equal power allocation method is used between user pairs, the total power of each NOMA user pair is determined. , is the total transmission power of the satellite, so when optimizing the internal power allocation, it is only necessary to confirm the proportional coefficient of the power allocation of the two users. , at this time the internal power distribution satisfies the following relationship:

[0037] in, is a set of T time slots.

[0038] Step 203: The first agent obtains a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results.

[0039] Optionally, the PPO agent interacts with the primary environment (the communication environment where the user is located) and makes decisions based on the decision factors output by the second agent. Mapping the actual intra-pair power distribution , Mapping the actual NOMA user pairing matrix. Therefore, the PPO agent is trained at the time step The joint action performed on the main environment can be expressed as , Represents the reflection unit in time slot t phase.

[0040] Step 204: Maximize the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the reward is associated with the effective data rate corresponding to all users.

[0041] In one embodiment of the present application, in the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

[0042] In one embodiment of the present application, before step 202, the following steps are further included: Step 21: The phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair are used as the action space of the first agent; all possible values ​​of the decision factor corresponding to the user pairing matrix are used as the action space of the second agent; Step 22: If the signal-to-interference-plus-noise ratio (SINR) corresponding to the satellite signals received by all users in the current time slot is greater than a preset threshold, rewards are set for the first and second agents in the current training time step, respectively, based on the effective data rates corresponding to all users in the current time slot. If the SINR corresponding to the satellite signals received by all users in the current time slot is at least partially less than the preset threshold, a penalty coefficient is set based on the SINR corresponding to the satellite signals received by each user, and rewards are set for the first and second agents in the current training time step based on the penalty coefficient and the effective data rates corresponding to all users in the current time slot. The penalty coefficient is set based on the difference between the preset threshold and the SNR below the preset threshold. The satellite signal received by each user is the signal transmitted by the satellite and transmitted to each user via an equivalent channel.

[0043] Optionally, step 21 may be implemented by the following method: For the PPO agent, the action space is designed to be the range of change of the reflection phase of RIS and the range of change of the internal power allocation ratio coefficient. Therefore, the continuous action space of the PPO agent is expressed as , is the reflection phase variation range of RIS, is the variation range of the internal power distribution ratio coefficient, have The action output by the PPO agent at training time step t can be expressed as .

[0044] For the DQN agent, the action space is designed as all possible values ​​of the NOMA pairing matrix decision factors. Therefore, the discrete action space is represented as , is An array of non-negative integers. The DQN agent is trained at time steps The output action can be expressed as , is the time step The pairing matrix decision factor under .

[0045] Optionally, step 22 may be implemented by the following method: The DRL algorithm trains the network output decision to maximize the cumulative reward, which is consistent with maximizing the average sum rate of the entire service duration, so its reward needs to include the sum rate. Reward function with mirror environment feedback The form is consistent, that is The reward at training time step t is set to the weighted sum rate of all users at time slot t, which is specifically expressed as:

[0046] in, is the normalized reward coefficient; This is a penalty coefficient set to ensure QoS and is defined as:

[0047] To ensure QoS, Should not be lower than the pre-set SINR threshold If the SINR of the received signal of any user in time slot t is greater than the pre-set threshold, no penalty is required. Otherwise, the threshold and below The accumulated value of the difference between the actual SINRs is used as a coefficient, and a penalty coefficient in the form of a decreasing exponential function is further set. Satisfied in any case .

[0048] In one embodiment of the present application, before step 203, the following may also be included: the second agent trains all possible user pairing results and outputs the decision factor.

[0049] Optionally, each decision factor corresponds to each user pairing result. Make a decision. The DQN algorithm can process a one-dimensional discrete action space, and the output action is a non-negative integer. Therefore, before training, all pairing matrices will be exhausted. , a total of Since the channel environment changes dynamically over time, the output of the DQN agent training is Mapped to the corresponding user pairing matrix .

[0050] In one embodiment of the present application, after the second neural network corresponding to the second agent calculates the reward corresponding to its action in the current training time step, the second agent transitions to the next state. The second neural network corresponding to the second agent is trained and its network parameters are updated using the state, user pairing matrix, and reward corresponding to the second agent in the current training time step, as well as the state corresponding to the next training time step. The second neural network is used to instruct the second agent to select and execute a corresponding action. This embodiment can be implemented as described in Table 1 below.

[0051] In one embodiment of the present application, after the first neural network corresponding to the first agent calculates the reward corresponding to the action it performs in the current training time step, the first agent transitions to the next state. The first neural network corresponding to the first agent is trained and its network parameters are updated using the state, user pairing matrix, and reward corresponding to the first agent in the current training time step, as well as the state corresponding to the next training time step. The first neural network is used to instruct the first agent to select and execute a corresponding action. This embodiment can be implemented as described in Table 1 below.

[0052] Table 1 Average and rate maximization of NOMA network assisted by PPO-DQN

[0053] In Table 1, is the mean of the action probability density function, is the variance of the action probability density function, X is the empirical batch size for training the network, and Y is the number of data reuses.

[0054] The specific simulation parameter settings are shown in Table 2.

[0055] Table 2 Simulation parameter settings

[0056] In one embodiment of the present application, a method for acquiring an equivalent channel includes: obtaining an equivalent channel of a signal transmitted by a satellite to each user based on a reflection unit in a reconfigurable smart surface.

[0057] In one embodiment of the present application, obtaining an equivalent channel from a satellite-transmitted signal to each user based on a reflective unit in a reconfigurable smart surface includes: Taking any reflection unit in the reconfigurable smart surface as a reference, the first component of the path loss of the link from the satellite to the reconfigurable smart surface is obtained according to the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface; the second component of the path loss of the link from the reconfigurable smart surface to the user is obtained according to the propagation direction of the signal after the reconfigurable smart surface reflects the signal transmitted by the satellite to the user; and the equivalent channel of the signal transmitted by the satellite to each user is obtained based on the first component, the second component, and the distances between the satellite, the reconfigurable smart surface, and the user.

[0058] In one embodiment of the present application, obtaining an equivalent channel from a signal transmitted by a satellite to each user based on the first component, the first component, and the distances between the satellite, the reconfigurable smart surface, and the user includes: According to the first component and the distance between the satellite and the reconfigurable smart surface, a channel from the satellite to the reconfigurable smart surface is obtained; according to the second component and the distance between the reconfigurable smart surface and the user, a channel from the reconfigurable smart surface to the user is obtained; according to the distance between the satellite and the user, a channel from the satellite to the user is obtained; and according to the channel from the satellite to the reconfigurable smart surface, the channel from the reconfigurable smart surface to the user, and the channel from the satellite to the user, an equivalent channel from the signal transmitted by the satellite to each user is obtained.

[0059] In one embodiment of the present application, the method for obtaining an equivalent channel can be implemented by the following method: For example, it is assumed that all channels follow a quasi-static block fading channel model, and the channel coefficients remain approximately constant within each channel coherent block and vary independently between blocks. For example, the transmission period is divided into T time slots Therefore, in order to adapt to the fading channel blocks, the reflection phase of the RIS, the power allocation of the satellites, and the NOMA user pairing must be reconfigured at the beginning of each time slot.

[0060] Time slot The reflection phase diagonal matrix of the inner RIS is , Indicates the dimension The complex field of can be expressed as:

[0061] in, Represents the reflection unit in time slot t To explore the performance limit of RIS, assume It can change continuously.

[0062] make Indicates time slot Next, the satellite to RIS channel, Indicates the RIS to the user in time slot t channel, let Indicates the satellite to user time slot t channel. And and All are Rician fading channels. is a Rayleigh fading channel, which can be expressed as:

[0063] in, Indicates reference distance Path loss at 、 and Respectively represent the distance from satellite to RIS, RIS to user Distance from satellite to user distance. 、 and represents the path loss exponent; and Both represent Rice factors. and Both represent LoS components; 、 and Both represent random NLoS components and obey complex Gaussian distribution .

[0064] Taking the first reflection unit in RIS as the reference, the coordinates are , then the first List The coordinates of each reflection unit can be expressed as:

[0065] The propagation direction of the satellite signal to the RIS can be expressed as:

[0066] No. List The phase of the first reflection unit relative to the first reflection unit can be expressed as:

[0067] The LoS component of the satellite-to-RIS link can thus be Expressed as:

[0068] Similarly, RIS to user The propagation direction of the signal can be expressed as:

[0069] No. List The phase of the first reflection unit relative to the first reflection unit can be expressed as:

[0070] therefore, It can be expressed as:

[0071] Therefore, at time slot t, the satellite to user The equivalent channel can be expressed as:

[0072] In one embodiment of the present application, the method for obtaining the effective data rate includes: obtaining the effective data rate of the signal sent by the satellite to each user based on the equivalent channel of the signal transmitted by the satellite to each user.

[0073] In one embodiment of the present application, obtaining the effective data rate at which the satellite transmits the signal to each user based on the equivalent channel of the signal transmitted by the satellite to each user includes: A signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user is obtained based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user. The signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user. An effective data rate at which the satellite transmits data to the user is obtained based on the signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

[0074] In one embodiment of the present application, the method for obtaining the effective data rate can be implemented by the following method: Based on the above equivalent channel acquisition process, the system bandwidth B is optionally evenly distributed to all NOMA user pairs, that is, there is no inter-NOMA pair interference in the system, only intra-pair interference. For the downlink NOMA scenario, it is assumed that in the user pair Satellites provide users with and The signals sent are and , satellites are allocated to users and The power is and , then the satellite is the user's The superimposed signal sent by all users in is:

[0075] Therefore, after transmission through the equivalent channel integrated by RIS, the user The received signal is:

[0076] in, Represents the signal the user expects to receive, is the signal of other users who are not expected, It's noise.

[0077] It's important to note that not all undesired signals cause interference, as Successive Interference Cancellation (SIC) allows users to partially cancel other users' signals. In downlink NOMA, after receiving the superimposed signals from the satellite, the receiving user uses SIC to decode the superimposed signals and cancel the interference. The core idea is to first detect the user signal with the highest power, then remove it from the superimposed signal, and then continue detecting the user signal with the second highest power.

[0078] In one possible implementation of this application, it is assumed that the user The first user has the worst channel condition. , the second user has good channel conditions ,Right now , and satisfies After SIC, the user The Signal to Interference plus Noise Ratio (SINR) of a user in the network can be calculated as:

[0079]

[0080] in, and is the user's received noise power in the sub-band, is the transmitting antenna gain, is the antenna gain at the receiving end. In order to ensure the communication QoS (Quality of Service), SINR should not be lower than the preset threshold. In summary, the effective data rate can be expressed as:

[0081] in, For bandwidth.

[0082] In one embodiment of the present application, the power allocation between NOMA user pairs adopts the equal power allocation method, so the embodiment of the present application will focus on the intra-pair power allocation, that is, the power allocation within the user pair. The NOMA user pairing matrix is ​​expressed as , Indicates the dimension The field of real numbers, whose elements are ,satisfy .like It means that user i and user j are paired successfully. This means that user i and user j are not in the same NOMA pair.

[0083] In one embodiment of the present application, the goal is to maximize the average sum rate of the downlink NOMA system over the entire service duration by jointly optimizing the passive beamforming of the RIS, the power allocation of the satellite, and the NOMA user pairing in each time slot. Therefore, the joint optimization problem (corresponding to maximizing the cumulative reward of the agent) can be expressed as:

[0084] The first constraint is the phase variation range of the RIS reflection unit. represents the total transmission power of the satellite, and the second constraint indicates that the power distribution between NOMA user pairs is evenly distributed. The third and fourth constraints are constraints on SIC decoding. The fifth to eighth constraints are constraints on the NOMA pairing matrix, which means that in a certain time slot, all users cannot be paired with themselves and can only be paired with one other user. The ninth constraint represents the effective rate obtained to meet QoS. This joint optimization problem has both high-dimensional continuous optimization variables and discrete optimization variables, and is a long-term effective data rate maximization problem, that is, the phase of each reflection unit of RIS, the power distribution of the satellite, and the NOMA user pairing need to be reconfigured in each time slot, which is difficult to solve using traditional optimization methods.

[0085] The above-mentioned embodiments of the present application integrate reconfigurable smart surface-assisted satellite downlink communication and non-orthogonal multiple access technology, and consider a reconfigurable smart surface-assisted satellite downlink NOMA communication network; at the same time, two communication links, satellite-user and satellite-RIS-user, are considered. RIS plays the role of assisting satellite-user communication in this scenario and enhancing the signal; based on the two communication links, satellite-user and satellite-RIS-user, the equivalent channel from satellite to user is expressed, and then the signal-to-interference-and-noise ratio of each user in the user pair after SIC is expressed, and then the effective data rate is expressed based on the Shannon formula; an optimization problem is formed with the goal of maximizing the average sum rate of the entire continuous service time of the downlink NOMA system, and the constraints include constraints on the phase shift of the RIS reflection unit, the power allocation of the satellite, and the NOMA user pairing.

[0086] In order to verify the effectiveness of the solution proposed in this application (Solution 1) in improving the effective data rate, the embodiment of this application compares the proposed algorithm with other traditional solutions, namely: (1) Solution 1: This is the solution proposed in the embodiment of the present application, which jointly optimizes the passive beamforming of RIS, the power allocation of satellites, and the NOMA user pairing based on the PPO-DQN algorithm.

[0087] (2) Solution 2: This solution uses the proposed PPO-DQN algorithm to optimize the passive beamforming of RIS and NOMA user pairing. Under the constraints, the power within the NOMA pair is randomly allocated.

[0088] (3) Scheme 3: In this scheme, the phase of the RIS reflector unit is randomly configured. In addition, the proposed PPO-DQN algorithm is used to optimize NOMA user pairing and power allocation within user pairs.

[0089] (4) Scheme 4: This scheme does not use RIS technology and only considers the direct link between the BS and the user. NOMA pairing uses the traditional near-far pairing scheme and adopts the PPO algorithm to optimize the power allocation within the NOMA user pair.

[0090] Figure 3 The cumulative reward curves of the four schemes are plotted under the parameter settings of 50 RIS reflection units, 1 GHz system bandwidth, and 60 W total satellite transmission power. Figure 3 We observed that the cumulative rewards for all schemes gradually converged to a stable state with increasing training rounds. After a certain number of training rounds, the cumulative rewards for all four schemes increased rapidly and gradually stabilized, as the agent learned a better action. Notably, the cumulative rewards for the proposed scheme 1 were significantly greater than those for the other traditional schemes, indicating that the proposed scheme 1 resulted in a higher average sum rate, demonstrating the effectiveness of the proposed algorithm.

[0091] like Figure 4 As shown in Figure 2, the simulation results demonstrate the impact of the number of RIS reflectors on the average sum rate. First, it can be observed that the average sum rates of the first three schemes increase with the increase in the number of RIS reflectors. This is because more RIS reflectors can reflect more signal paths and signal power, thereby increasing the data rate. More importantly, the curve of Scheme 1 is always higher than that of the other three schemes, indicating that the proposed PPO-DQN algorithm, which jointly optimizes RIS passive beamforming, satellite power allocation, and NOMA user pairing, can achieve the highest average sum rate compared to traditional schemes. Since Scheme 4 does not use RIS technology, the number of RIS reflectors has no impact on its performance, and the curve of this scheme is a straight line.

[0092] like Figure 5 As shown in Figure 2, simulation results demonstrate the impact of system bandwidth on the effective data rate. First, it can be seen that the performance of all schemes improves with increasing system bandwidth. This is because, according to Shannon's equation, bandwidth increases throughput. Second, Scheme 1 consistently outperforms the other schemes, demonstrating that the proposed PPO-DQN algorithm can effectively improve the average sum rate.

[0093] The embodiment of the present application further provides a satellite-to-ground downlink resource allocation system, which includes: A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; A power allocation module is used to obtain the power allocation within the user pair based on the power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; A user pairing module is used for the first agent to obtain a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results; The maximization module is configured to maximize the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all users.

[0094] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any of the above-mentioned satellite-to-ground downlink resource allocation methods is implemented.

[0095] An embodiment of the present application further provides a computer-readable storage medium, which includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes any of the above-mentioned satellite-to-ground downlink resource allocation methods.

[0096] Those skilled in the art should clearly understand that, for the convenience and brevity of description, the specific working processes of the resource allocation system, electronic device and computer-readable storage medium described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present application can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present application should be included in the scope of protection of the claims of the present application.

Claims

1. A method for allocating resources for a satellite-to-ground downlink, characterized in that: include: Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; Obtaining a power allocation within the user pair according to the power allocation ratio coefficient of the user pair; wherein the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; The first agent obtains a user pairing matrix based on the decision factor output by the second agent; The user pairing matrix is ​​used to store all user pairing results; The cumulative rewards corresponding to the first agent and the second agent are maximized to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the reward is associated with the effective data rate corresponding to all the users.

2. The satellite-to-ground downlink resource allocation method according to claim 1, characterized in that: The determining of the agent state based on the equivalent channels from the satellite to all users includes: The state used by the agent in the current training time step is defined as the equivalent channel of the signal transmitted by the satellite to each of the users in the previous time slot; the training time step corresponds to the time slot; and a single communication cycle of the satellite is evenly divided to obtain multiple time slots.

3. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: Before obtaining the power allocation within the user pair according to the power allocation ratio coefficient of the user pair, the method further includes: Taking the phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair as the action space of the first intelligent agent; Taking all possible values ​​of the decision factor corresponding to the user pairing matrix as the action space of the second agent; Rewards are set for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot.

4. The satellite-to-ground downlink resource allocation method according to claim 3, wherein: The step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If the signal-to-interference-and-noise ratio (SINR) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is greater than a preset threshold, rewards are set for the first agent and the second agent in the current training time step according to the effective data rate corresponding to all the users in the current time slot; wherein the signal transmitted by the satellite received by each user is the signal transmitted by the satellite and transmitted through the equivalent channel to each user.

5. The method for allocating satellite-to-ground downlink resources according to claim 3, wherein: The step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If a signal-to-interference-plus-noise ratio (SIN) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is at least partially less than a preset threshold, a penalty coefficient is set based on the SIN corresponding to the signal transmitted by the satellite received by each user, and a reward is set for the first agent and the second agent in the current training time step based on the penalty coefficient and the effective data rate corresponding to all the users in the current time slot; the penalty coefficient is set based on the difference between the preset threshold and the SIN below the preset threshold; wherein the signal transmitted by the satellite received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user.

6. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: Before the first agent obtains the user pairing matrix based on the decision factor output by the second agent, the method further includes: The second agent trains all possible user pairing results and outputs the decision factors; each decision factor corresponds to each user pairing result one by one.

7. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: The step of obtaining the power allocation within the user pair according to the power allocation ratio coefficient of the user pair includes: determining a total power allocated by the satellite to each of the user pairs; According to the total power of each user pair and the power allocation ratio coefficient of each user pair, the power corresponding to each user in each user pair is obtained, that is, the power allocated by the satellite to each user in each user pair.

8. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: Also includes: In the current training time step, after the second neural network corresponding to the second agent calculates the reward corresponding to the action it performs, the second agent transitions to the next state; Training the second neural network corresponding to the second agent and updating its network parameters using the state corresponding to the second agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; The second neural network is used to instruct the second agent to select and execute corresponding actions.

9. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: Also includes: In the current training time step, after the first neural network corresponding to the first agent calculates the reward corresponding to the action performed by the first agent, the first agent transitions to the next state; The first neural network corresponding to the first agent is trained and its network parameters are updated using the state corresponding to the first agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the first neural network is used to instruct the first agent to select and execute a corresponding action.

10. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: Maximizing the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix includes: In the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

11. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: The method for obtaining the equivalent channel comprises: obtaining the equivalent channel of the signal transmitted by the satellite to each of the users according to the reflection unit in the reconfigurable smart surface; The step of obtaining the equivalent channel of the signal transmitted by the satellite to each user according to the reflection unit in the reconfigurable smart surface comprises: Taking any one reflective unit in the reconfigurable smart surface as a reference, and obtaining a first component of a path loss of a link from the satellite to the reconfigurable smart surface according to a propagation direction of a signal transmitted from the satellite to the reconfigurable smart surface; Obtaining a second component of the path loss of a link from the reconfigurable smart surface to the user according to a propagation direction of a signal from the reconfigurable smart surface to the user after the signal transmitted by the satellite is reflected by the reconfigurable smart surface; The equivalent channel of the signal transmitted by the satellite to each of the users is obtained according to the first component, the distances between the first component and the satellite, the reconfigurable smart surface and the users.

12. The satellite-to-ground downlink resource allocation method according to claim 11, characterized in that: Obtaining the equivalent channel of the signal transmitted by the satellite to each of the users based on the first component, the first component, and the distances among the satellite, the reconfigurable smart surface, and the user, includes: Obtaining a channel from the satellite to the reconfigurable smart surface according to the first component and a distance between the satellite and the reconfigurable smart surface; Obtaining a channel from the reconfigurable smart surface to the user according to the second component and a distance between the reconfigurable smart surface and the user; Obtaining a channel from the satellite to the user according to a distance between the satellite and the user; The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the channel from the satellite to the reconfigurable smart surface, the channel from the reconfigurable smart surface to the user, and the channel from the satellite to the user.

13. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: The method for obtaining the effective data rate includes: obtaining the effective data rate of the signal transmitted by the satellite to each of the users according to the equivalent channel of the signal transmitted by the satellite to each of the users; The obtaining, based on the equivalent channel of the signal transmitted by the satellite to each user, the effective data rate at which the satellite transmits the signal to each user, comprises: Obtaining a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user; the signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user; An effective data rate at which the satellite transmits data to the user is obtained according to a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

14. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 7, characterized in that: All users within the satellite coverage area are divided into multiple user pairs; the pairing results of the multiple user pairs constitute the user pairing matrix; each user pair includes multiple users; users in the same user pair use the same time and frequency band.

15. A satellite-to-ground downlink resource allocation system, characterized in that: include: A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; A power allocation module, configured to obtain a power allocation within a user pair based on a power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; A user pairing module, configured for the first agent to obtain a user pairing matrix based on the decision factor output by the second agent; The user pairing matrix is ​​used to store all user pairing results; A maximization module is configured to maximize the cumulative rewards corresponding to the first agent and the second agent to obtain an optimal phase of each reflective unit in the reconfigurable smart surface, power allocation of the satellite, and user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all the users.

16. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the satellite-to-ground downlink resource allocation method according to any one of claims 1 to 14 is implemented.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer is enabled to execute the satellite-to-ground downlink resource allocation method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Satellite-ground convergence network resource allocation method based on DRL

    CN117220751A

  • Resource allocation method in D2D-NOMA uplink communication system based on RIS

    CN118158802A

  • Method for optimizing intelligent Internet of Vehicles communication system model based on double-relay RIS

    CN119865267A

  • Reconfigurable intelligent surface enabled leakage suppression and signal power maximization system and method

    WO2023037243A1