Resource allocation method, allocation system, electronic device and medium for satellite downlink

By jointly optimizing the passive beamforming of RIS, the power allocation of satellites and NOMA user pairing, and adopting the hybrid deep reinforcement learning algorithm PPO-DQN, the problem of low resource allocation efficiency in the satellite downlink NOMA communication network is solved, and the average sum rate of satellite communication is improved.

CN120567237BActive Publication Date: 2025-10-10CHINA SATELLITE NETWORK EXPLORATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511044997.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-10
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies in satellite communications have not fully studied the satellite downlink NOMA communication network and have not considered the optimization of transmission power and user pairing, resulting in low resource allocation efficiency and the inability to provide real-time and accurate resource allocation for ground users.

Method used

By jointly optimizing the passive beamforming of RIS, the power allocation of satellites and NOMA user pairing, and adopting the hybrid deep reinforcement learning algorithm PPO-DQN, the continuous service time and rate of the downlink NOMA system are maximized.

Benefits of technology

It improves the average sum rate of satellite downlink, achieves more efficient resource allocation, and enhances users' communication experience and the system's continuous service capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567237B_ABST
    Figure CN120567237B_ABST
Patent Text Reader

Abstract

The application discloses a resource allocation method and system for a satellite downlink, electronic equipment and a medium. The method comprises determining an agent state based on an equivalent channel from a satellite to all users; obtaining power allocation within a user pair according to a power allocation proportion coefficient of the user pair; obtaining a user pairing matrix according to a decision factor output by a second agent; and maximizing cumulative rewards corresponding to the first agent and the second agent to obtain an optimal phase of each reflecting unit in a reconfigurable intelligent surface, power allocation of the satellite and the user pairing matrix. The application can effectively allocate resources for users and RIS in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of satellite communication technology, and in particular to a resource allocation method, allocation system, electronic equipment and medium for a satellite-to-ground downlink. Background Art

[0002] In satellite communications, how to effectively allocate onboard resources has become a key component of satellite resource management. As satellite resources become increasingly diverse, traditional satellite resource allocation methods are inefficient and unable to accurately allocate resources to ground users or related devices in real time.

[0003] Existing technical solutions often use traditional optimization methods such as fractional programming, semidefinite relaxation, and continuous convex approximation to optimize beamforming for large-scale satellite arrays and smart reflective metasurfaces to improve communication and perception performance. However, these methods do not consider the use of NOMA networks, do not address NOMA user grouping, and do not optimize transmit power. Furthermore, these optimization methods are relatively complex.

[0004] Satellite downlink NOMA communication networks have not been fully explored in existing technologies. Optimization of satellite transmit power and user pairing, two parameters crucial for improving NOMA network speeds, has not been considered. Furthermore, existing optimization methods are either heuristic or complex. Summary of the Invention

[0005] In view of this, the present application provides a resource allocation method, allocation system, electronic device and medium for a satellite-to-ground downlink, which comprehensively considers the joint optimization of the passive beamforming of RIS, the power allocation of the satellite and the NOMA user pairing in each time slot, that is, maximizes the effective data rate of the downlink NOMA system for the entire continuous service time, and obtains the optimal resource allocation result.

[0006] The present application discloses a method for allocating resources of a satellite-to-ground downlink, which includes:

[0007] Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent;

[0008] Obtaining a power allocation within the user pair according to the power allocation ratio coefficient of the user pair; wherein the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair;

[0009] The first agent obtains a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results;

[0010] maximizing cumulative rewards corresponding to the first agent and the second agent to obtain an optimal phase of each reflective unit in the reconfigurable intelligent surface, power allocation of the satellite, and a user pairing matrix; wherein the reward is associated with effective data rates corresponding to all the users.

[0011] Further, the agent state is determined based on equivalent channels of the satellite to all the users, including:

[0012] The state used by the agent in a current training time step is defined as the equivalent channel of a signal transmitted by the satellite in a previous time slot to each of the users; the training time step corresponds to the time slot; a single communication period of the satellite is evenly divided to obtain a plurality of time slots.

[0013] Further, before obtaining the power allocation within the user pair according to the power allocation proportion coefficient of the user pair, further comprising:

[0014] The change range of the phase of each reflective unit in the reconfigurable intelligent surface and the change range of the power allocation proportion coefficient of the user pair are taken as the action space of the first agent;

[0015] All possible values of the decision factor corresponding to the user pairing matrix are taken as the action space of the second agent;

[0016] According to the effective data rates corresponding to all the users in the current time slot, rewards are set for the first agent and the second agent, respectively.

[0017] Further, the rewards are set for the first agent and the second agent, respectively, according to the effective data rates corresponding to all the users in the current time slot, including:

[0018] If the signal-to-interference-plus-noise ratio corresponding to the signal transmitted by the satellite received by all the users in the current time slot is greater than a preset threshold, rewards are set for the first agent and the second agent in the current training time step according to the effective data rates corresponding to all the users in the current time slot; wherein the signal received by each of the users is the signal transmitted by the satellite after transmission through the equivalent channel and reaching each of the users.

[0019] Further, the rewards are set for the first agent and the second agent, respectively, according to the effective data rates corresponding to all the users in the current time slot, including:

[0020] if all the users in the current time slot receive the signal transmitted by the satellite corresponding to a signal-to-interference-and-noise ratio less than the preset threshold, then according to the signal-to-interference-and-noise ratio corresponding to the signal received by each user, a penalty coefficient is set, and according to the penalty coefficient and the effective data rate corresponding to all the users in the current time slot, the first agent and the second agent in the current training time step are set with a reward; the penalty coefficient is set according to the difference between the preset threshold and the signal-to-interference-and-noise ratio below the preset threshold; wherein the signal transmitted by the satellite received by each user is the signal transmitted by the satellite after transmission through the equivalent channel and reaching each user.

[0021] Further, before the first agent obtains the user pairing matrix according to the decision factor output by the second agent, it further includes:

[0022] The second agent trains all possible user pairing results and outputs the decision factor; each decision factor corresponds to each user pairing result one by one.

[0023] Further, the power allocation within the user pair is obtained according to the power allocation proportion coefficient of the user pair, which includes:

[0024] determining the total power allocated by the satellite to each user pair;

[0025] obtaining the power corresponding to each user in each user pair, i.e. the power allocated by the satellite to each user in each user pair, according to the total power of each user pair and the power allocation proportion coefficient of each user pair.

[0026] Further, it further includes:

[0027] After the second neural network corresponding to the second agent in the current training time step calculates the reward corresponding to the action it executes, the second agent moves to the next state;

[0028] The second neural network corresponding to the second agent is trained and its network parameters are updated through the state corresponding to the second agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the second neural network is used to instruct the second agent to select the corresponding action and execute.

[0029] Further, it further includes:

[0030] After the first neural network corresponding to the first agent in the current training time step calculates the reward corresponding to the action it executes, the first agent moves to the next state;

[0031] The first neural network corresponding to the first agent is trained and its network parameters are updated using the state corresponding to the first agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the first neural network is used to instruct the first agent to select and execute a corresponding action.

[0032] Furthermore, maximizing the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix includes:

[0033] In the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

[0034] Furthermore, the method for obtaining the equivalent channel includes: obtaining the equivalent channel of the signal transmitted by the satellite to each of the users according to the reflection unit in the reconfigurable smart surface;

[0035] The step of obtaining the equivalent channel of the signal transmitted by the satellite to each user according to the reflective unit in the reconfigurable smart surface comprises:

[0036] Taking any one reflective unit in the reconfigurable smart surface as a reference, and obtaining a first component of a path loss of a link from the satellite to the reconfigurable smart surface according to a propagation direction of a signal transmitted from the satellite to the reconfigurable smart surface;

[0037] Obtaining a second component of the path loss of a link from the reconfigurable smart surface to the user according to a propagation direction of a signal from the reconfigurable smart surface to the user after the signal transmitted by the satellite is reflected by the reconfigurable smart surface;

[0038] The equivalent channel of the signal transmitted by the satellite to each of the users is obtained according to the first component, the distances between the first component and the satellite, the reconfigurable smart surface and the users.

[0039] Furthermore, obtaining the equivalent channel of the signal transmitted by the satellite to each of the users based on the first component, the first component, and the distances among the satellite, the reconfigurable smart surface, and the user includes:

[0040] Obtaining a channel from the satellite to the reconfigurable smart surface according to the first component and a distance between the satellite and the reconfigurable smart surface;

[0041] Obtaining a channel from the reconfigurable smart surface to the user according to the second component and a distance between the reconfigurable smart surface and the user;

[0042] Obtaining a channel from the satellite to the user according to a distance between the satellite and the user;

[0043] The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the channel from the satellite to the reconfigurable smart surface, the channel from the reconfigurable smart surface to the user, and the channel from the satellite to the user.

[0044] Furthermore, the method for obtaining the effective data rate includes: obtaining the effective data rate of the signal transmitted by the satellite to each of the users based on the equivalent channel of the signal transmitted by the satellite to each of the users;

[0045] The obtaining, based on the equivalent channel of the signal transmitted by the satellite to each user, the effective data rate at which the satellite transmits the signal to each user, comprises:

[0046] Obtaining a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user; the signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user;

[0047] An effective data rate at which the satellite transmits data to the user is obtained according to a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

[0048] Furthermore, all users within the coverage of the satellite are divided into multiple user pairs; the pairing results of the multiple user pairs constitute the user pairing matrix; each user pair includes multiple users; and users in the same user pair use the same time and frequency band.

[0049] The present application also discloses a satellite-to-ground downlink resource allocation system, which includes:

[0050] A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent;

[0051] A power allocation module, configured to obtain a power allocation within a user pair based on a power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair;

[0052] A user pairing module, configured for the first agent to obtain a user pairing matrix based on the decision factors output by the second agent; the user pairing matrix is ​​used to store all user pairing results;

[0053] A maximization module is configured to maximize the cumulative rewards corresponding to the first agent and the second agent to obtain an optimal phase of each reflective unit in the reconfigurable smart surface, power allocation of the satellite, and user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all the users.

[0054] The present application also discloses an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the resource allocation method for the satellite-to-ground downlink described in any one of the above items is implemented.

[0055] The present application also discloses a computer-readable storage medium, which includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes any one of the above-mentioned satellite-to-ground downlink resource allocation methods.

[0056] Due to the adoption of the above technical solution, the present application has the following advantages: the present application maximizes the average sum rate of the entire continuous service time of the downlink NOMA system by jointly optimizing the passive beamforming of the RIS, the power allocation of the satellite and the NOMA user pairing in each time slot, thereby improving the average sum rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0058] Figure 1 This is a schematic diagram of RIS-assisted downlink non-orthogonal multiple access in an urban scenario according to an embodiment of the present application;

[0059] Figure 2 This is a flowchart of a method for allocating resources for a satellite-to-ground downlink according to an embodiment of the present application;

[0060] Figure 3 This is a schematic diagram of the cumulative reward convergence of an embodiment of the present application;

[0061] Figure 4 Schematic diagram showing the effect of the number of RIS reflection units on the average sum rate according to an embodiment of the present application;

[0062] Figure 5This is a schematic diagram of the impact of system bandwidth on average sum rate in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The present application is further described with reference to the accompanying drawings and embodiments. The embodiments described are only a part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0064] A RIS-assisted downlink Non-Orthogonal Multiple Access (NOMA) system in an urban scenario is considered. Both satellite-to-user and satellite-to-RIS-to-user links are considered. The average downlink data rate over the entire service duration of the NOMA system is maximized by jointly optimizing the RIS passive beamforming, satellite power allocation, and NOMA user pairing in each time slot. This joint optimization problem involves both high-dimensional continuous and discrete optimization variables and is a long-term effective data rate maximization problem. This problem requires reconfiguring the phase of each reflector in the RIS, the satellite power allocation, and the NOMA user pairing in each time slot, making it difficult to solve using traditional optimization methods.

[0065] The embodiments of the present application provide a resource allocation method, allocation system, electronic device, and medium for a satellite-to-ground downlink, which integrates reconfigurable intelligent surface-assisted satellite downlink communication and non-orthogonal multiple access technology. By jointly optimizing the passive beamforming of the RIS in each time slot, the power allocation of the satellite, and the NOMA user pairing, the average sum rate of the downlink NOMA system over the entire continuous service time (the average value of the sum of the effective data rates of all users corresponding to all time slots in a single satellite data transmission cycle, i.e., the formula below) is maximized. , t is the time slot, T is the total number of time slots, K is the total number of user pairs, k is the sequence number of user pairs, and n is the user sequence number in each user pair. is the actual intra-pair power allocation).

[0066] Due to environmental factors and other factors, the channel quality of direct communication links between satellites and ground users can be poor. To enhance the user experience and increase the average and rate of the system during continuous service, a reconfigurable intelligent surface (RIS) can be used to assist in satellite-to-ground communication services. RIS consists of a large number of low-cost passive reflectors, each of which can be individually controlled. By dynamically adjusting the phase and amplitude of the incident signal, RIS achieves precise signal control and optimization, thereby improving link reliability and spectral efficiency.

[0067] As shown in FIG. 1, consider a RIS-aided downlink non-orthogonal multiple access system in an urban scenario. Exemplarily, in a three-dimensional Cartesian coordinate system, a satellite located at provides communication services for 2K single-antenna ground users, respectively, are coordinate values of the satellite on the x-axis, y-axis and z-axis. An RIS located at provides assisted indirect link communication services for users, respectively, are coordinate values of the RIS on the x-axis, y-axis and z-axis. The RIS has a total of reflective units, each column contains reflective units, and each row contains reflective units, and each unit has a horizontal and vertical interval of half a wavelength , and the set of reflective units is denoted as . The 2K ground users are paired to form K pairs of NOMA user pairs, and the set of user pairs is denoted as , and users in the same user pair use the same time / frequency band resources. The nth user in the kth user pair is denoted as , where . The coordinates of the user are denoted as .

[0068] The optimization problem has high-dimensional continuous optimization variables and discrete optimization variables, and is a long-term effective data rate maximization problem, that is, the phase of each reflective unit in the RIS, the power allocation of the satellite and the NOMA user pairing need to be reconfigured at each time slot, which is difficult to solve by traditional optimization methods. In order to stably provide communication services for ground users in the long term, the embodiments of the present application provide a resource allocation method for a satellite downlink, which considers jointly optimizing the passive beamforming of the RIS, the power allocation of the satellite and the NOMA user pairing in a certain transmission period (for example, the period of satellite transmission data) to maximize the average sum rate, that is, a deep reinforcement learning (Deep Reinforcement Learning, DRL) method that can optimize mixed actions is used to solve the optimization problem. Among the current popular DRL algorithms, the PPO algorithm can optimize continuous actions, and the DQN algorithm can optimize discrete actions. A hybrid DRL framework, namely the PPO-DQN algorithm, can be used, which constructs the MDP (Markov Decision Process) of the main environment and the mirror environment, thereby combining the two agents together to realize joint mixed action control.

[0069] The embodiments of the present application can be applied to the application scenarios shown in Figure 1 , and refer to Figure 2The satellite-to-ground downlink resource allocation method may include the following steps:

[0070] Step 201: Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent.

[0071] In one embodiment of the present application, the state used by the agent in the current training time step is defined as the equivalent channel of the signal transmitted by the satellite to each user in the previous time slot; the training time step corresponds to the time slot; and a single communication cycle of the satellite is evenly divided to obtain multiple time slots.

[0072] For example, the agent rationally plans the RIS reflection phase, satellite power allocation, and NOMA user pairing based on the channel conditions of the users in the network. Therefore, the equivalent channel from the satellite to all users is described as the agent state. The state space can be expressed as , There are 2K dimensions in total, and each dimension is continuous. is the equivalent channel of the 2Kth ground user. The state of the PPO agent is expressed as , the DQN agent state is represented as , and satisfies The state used in time slot t, i.e., time step t in the MDP Defined as The equivalent channel calculated by the time slot is thus expressed as , Indicates the equivalent channel of the 2Kth terrestrial user at time slot t-1. Directly set to .

[0073] Step 202: Obtain the power allocation within the user pair according to the power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair.

[0074] In one embodiment of the present application, obtaining power allocation within a user pair based on the power allocation ratio coefficient of the user pair includes: determining the total power allocated by the satellite to each user pair; obtaining the power corresponding to each user in each user pair based on the total power of each user pair and the power allocation ratio coefficient of each user pair, that is, the power allocated by the satellite to each user in each user pair;

[0075] Specifically, the goal of the embodiment of the present application is to optimize the phase of RIS by joint , Satellite power distribution and NOMA user pairing matrix To maximize the average sum rate of the system during the entire service duration. Since the equal power allocation method is used between user pairs, the total power of each NOMA user pair is determined. , is the total transmission power of the satellite, so when optimizing the internal power allocation, it is only necessary to confirm the proportional coefficient of the power allocation of the two users. , at this time the internal power distribution satisfies the following relationship:

[0076]

[0077] in, is a set of T time slots.

[0078] Step 203: The first agent obtains a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results.

[0079] Optionally, the PPO agent interacts with the primary environment (the communication environment where the user is located) and makes decisions based on the decision factors output by the second agent. Mapping the actual intra-pair power distribution , Mapping the actual NOMA user pairing matrix. Therefore, the PPO agent is trained at the time step The joint action performed on the main environment can be expressed as , Represents the reflection unit in time slot t phase.

[0080] Step 204: Maximize the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the reward is associated with the effective data rate corresponding to all users.

[0081] In one embodiment of the present application, in the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

[0082] In one embodiment of the present application, before step 202, the following steps are further included:

[0083] Step 21: The phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair are used as the action space of the first agent; all possible values ​​of the decision factor corresponding to the user pairing matrix are used as the action space of the second agent;

[0084] Step 22: If the signal-to-interference-and-noise ratios corresponding to the signals received by all users in the current time slot are all greater than the preset threshold, rewards are set for the first and second agents in the current training time step according to the effective data rates corresponding to all users in the current time slot; if the signal-to-interference-and-noise ratios corresponding to the signals received by all users in the current time slot are at least partially less than the preset threshold, a penalty coefficient is set according to the signal-to-interference-and-noise ratios corresponding to the signals received by each user, and rewards are set for the first and second agents in the current training time step according to the penalty coefficient and the effective data rates corresponding to all users in the current time slot; the penalty coefficient is set according to the difference between the preset threshold and the signal-to-interference-and-noise ratios less than the preset threshold; wherein the signal received by each user from the satellite is the signal transmitted by the satellite after transmission through the equivalent channel and reaching each user.

[0085] Optionally, step 21 can be implemented by the following method:

[0086] For the PPO agent, the action space is designed as the change range of the reflection phase of the RIS and the change range of the in-coverage power allocation proportion coefficient. Therefore, the continuous action space of the PPO agent is expressed as , is the change range of the reflection phase of the RIS, is the change range of the in-coverage power allocation proportion coefficient, has dimensions and changes continuously. The action output by the PPO agent at training time step t can be expressed as .

[0087] For the DQN agent, the action space is designed as all possible values of the NOMA pairing matrix decision factor. Therefore, the discrete action space is expressed as , is an array composed of non-negative integers. The action output by the DQN agent at training time step can be expressed as , is the pairing matrix decision factor at time step .

[0088] Optionally, step 22 can be implemented by the following method:

[0089] The DRL algorithm trains the network output decision with the purpose of maximizing the cumulative reward, which is consistent with maximizing the average sum rate of the entire duration of service, so the reward needs to include the sum rate. The reward function fed back by the main environment is consistent in form with the reward function fed back by the mirror environment , that is, The reward at training time step t is set to the weighted sum rate of all users at time slot t, which is specifically expressed as:

[0090]

[0091] in, is the normalized reward coefficient; This is a penalty coefficient set to ensure QoS and is defined as:

[0092]

[0093] To ensure QoS, Should not be lower than the pre-set SINR threshold If the SINR of the received signal of any user in time slot t is greater than the pre-set threshold, no penalty is required. Otherwise, the threshold and below The accumulated value of the difference between the actual SINRs is used as a coefficient, and a penalty coefficient in the form of a decreasing exponential function is further set. Satisfied in any case .

[0094] In one embodiment of the present application, before step 203, the following may also be included: the second agent trains all possible user pairing results and outputs the decision factor.

[0095] Optionally, each decision factor corresponds to each user pairing result. Make a decision. The DQN algorithm can process a one-dimensional discrete action space, and the output action is a non-negative integer. Therefore, before training, all pairing matrices will be exhausted. , a total of Since the channel environment changes dynamically over time, the output of the DQN agent training is Mapped to the corresponding user pairing matrix .

[0096] In one embodiment of the present application, after the second neural network corresponding to the second agent calculates the reward corresponding to its action in the current training time step, the second agent transitions to the next state. The second neural network corresponding to the second agent is trained and its network parameters are updated using the state, user pairing matrix, and reward corresponding to the second agent in the current training time step, as well as the state corresponding to the next training time step. The second neural network is used to instruct the second agent to select and execute a corresponding action. This embodiment can be implemented as described in Table 1 below.

[0097] In one embodiment of the present application, after the first neural network corresponding to the first agent calculates the reward corresponding to the action it performs in the current training time step, the first agent transitions to the next state. The first neural network corresponding to the first agent is trained and its network parameters are updated using the state, user pairing matrix, and reward corresponding to the first agent in the current training time step, as well as the state corresponding to the next training time step. The first neural network is used to instruct the first agent to select and execute a corresponding action. This embodiment can be implemented as described in Table 1 below.

[0098] Table 1 Average and rate maximization of NOMA network assisted by PPO-DQN

[0099]

[0100] In Table 1, is the mean of the action probability density function, is the variance of the action probability density function, X is the empirical batch size for training the network, and Y is the number of data reuses.

[0101] The specific simulation parameter settings are shown in Table 2.

[0102] Table 2 Simulation parameter settings

[0103]

[0104] In one embodiment of the present application, a method for acquiring an equivalent channel includes: obtaining an equivalent channel of a signal transmitted by a satellite to each user based on a reflection unit in a reconfigurable smart surface.

[0105] In one embodiment of the present application, obtaining an equivalent channel from a satellite-transmitted signal to each user based on a reflective unit in a reconfigurable smart surface includes:

[0106] Taking any reflection unit in the reconfigurable smart surface as a reference, the first component of the path loss of the link from the satellite to the reconfigurable smart surface is obtained according to the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface; the second component of the path loss of the link from the reconfigurable smart surface to the user is obtained according to the propagation direction of the signal after the reconfigurable smart surface reflects the signal transmitted by the satellite to the user; and the equivalent channel of the signal transmitted by the satellite to each user is obtained based on the first component, the second component, and the distances between the satellite, the reconfigurable smart surface, and the user.

[0107] In an embodiment of the present application, the equivalent channel of the signal transmitted by the satellite to each user is obtained according to the first component, the distance between the first component, the satellite, the reconfigurable intelligent surface and the user, comprising:

[0108] The channel from the satellite to the reconfigurable intelligent surface is obtained according to the first component and the distance between the satellite and the reconfigurable intelligent surface; the channel from the reconfigurable intelligent surface to the user is obtained according to the second component and the distance between the reconfigurable intelligent surface and the user; the channel from the satellite to the user is obtained according to the distance between the satellite and the user; and the equivalent channel of the signal transmitted by the satellite to each user is obtained according to the channel from the satellite to the reconfigurable intelligent surface, the channel from the reconfigurable intelligent surface to the user and the channel from the satellite to the user.

[0109] In an embodiment of the present application, the method for obtaining the equivalent channel can be realized by the following method:

[0110] Exemplarily, it is assumed that all channels follow a quasi-static block fading channel model, the channel coefficients remain approximately constant in each channel coherence block, and vary independently between blocks. Exemplarily, the transmission period is evenly divided into T time slots Therefore, in order to adapt to the fading channel block, it is necessary to reconfigure the reflection phase of the RIS, the power allocation of the satellite and the NOMA user pairing at the beginning of each time slot.

[0111] Let the reflection phase diagonal matrix of the RIS in the time slot be , denote the complex field with dimension , which can be expressed as:

[0112]

[0113] wherein denotes the phase of the reflection unit in the time slot t. In order to explore the performance limit of the RIS, it is assumed that can be continuously changed.

[0114] Let denote the channel from the satellite to the RIS in the time slot , let denote the channel from the RIS to the user in the time slot t, and let denote the channel from the satellite to the user in the time slot t. And and are both Rician fading channels, and Rayleigh fading channel, which can be respectively expressed as:

[0115]

[0116] in, Indicates reference distance Path loss at 、 and Respectively represent the distance from satellite to RIS, RIS to user The distance from the satellite to the user distance. 、 and represents the path loss exponent; and Both represent Rice factors. and Both represent LoS components; 、 and Both represent random NLoS components and obey complex Gaussian distribution .

[0117] Taking the first reflection unit in RIS as the reference, the coordinates are , then the first List The coordinates of each reflection unit can be expressed as:

[0118]

[0119] The propagation direction of the satellite signal to the RIS can be expressed as:

[0120]

[0121] No. List The phase of the first reflection unit relative to the first reflection unit can be expressed as:

[0122]

[0123] The LoS component of the satellite-to-RIS link can thus be Expressed as:

[0124]

[0125] Similarly, RIS to user The propagation direction of the signal can be expressed as:

[0126]

[0127] No. List The phase of the first reflection unit relative to the first reflection unit can be expressed as:

[0128]

[0129] therefore, It can be expressed as:

[0130]

[0131] Therefore, at time slot t, the satellite to user The equivalent channel can be expressed as:

[0132]

[0133] In one embodiment of the present application, the method for obtaining the effective data rate includes: obtaining the effective data rate of the signal sent by the satellite to each user based on the equivalent channel of the signal transmitted by the satellite to each user.

[0134] In one embodiment of the present application, obtaining the effective data rate at which the satellite transmits the signal to each user based on the equivalent channel of the signal transmitted by the satellite to each user includes:

[0135] A signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user is obtained based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user. The signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user. An effective data rate at which the satellite transmits data to the user is obtained based on the signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

[0136] In one embodiment of the present application, the method for obtaining the effective data rate can be implemented by the following method:

[0137] Based on the above equivalent channel acquisition process, the system bandwidth B is optionally evenly distributed to all NOMA user pairs, that is, there is no inter-NOMA pair interference in the system, only intra-pair interference. For the downlink NOMA scenario, assuming that in the user pair Satellites provide users with and The signals sent are and , satellites are allocated to users and The power is and , then the satellite is the user's The superimposed signal sent by all users in is:

[0138]

[0139] Therefore, after transmission through the equivalent channel integrated by RIS, the user The received signal is:

[0140]

[0141] in, Represents the signal the user expects to receive, is the signal of other users who are not expected, It's noise.

[0142] It's important to note that not all undesired signals cause interference, as Successive Interference Cancellation (SIC) allows users to partially cancel other users' signals. In downlink NOMA, after receiving the superimposed signals from the satellite, the receiving user uses SIC to decode the superimposed signals and cancel the interference. The core idea is to first detect the user signal with the highest power, then remove it from the superimposed signal, and then continue detecting the user signal with the second highest power.

[0143] In one possible implementation of this application, it is assumed that the user The first user has the worst channel condition. , the second user has good channel conditions ,Right now , and satisfies After SIC, the user The Signal to Interference plus Noise Ratio (SINR) of a user in the network can be calculated as:

[0144]

[0145]

[0146] in, and is the user's received noise power in the sub-band, is the transmitting antenna gain, is the antenna gain at the receiving end. In order to ensure the communication QoS (Quality of Service), SINR should not be lower than the preset threshold. In summary, the effective data rate can be expressed as:

[0147]

[0148] in, For bandwidth.

[0149] In one embodiment of the present application, the power allocation between NOMA user pairs adopts the equal power allocation method, so the embodiment of the present application will focus on the intra-pair power allocation, that is, the power allocation within the user pair. The NOMA user pairing matrix is ​​expressed as , Indicates the dimension The field of real numbers, whose elements are ,satisfy .like It means that user i and user j are paired successfully. This means that user i and user j are not in the same NOMA pair.

[0150] In one embodiment of the present application, the goal is to maximize the average sum rate of the downlink NOMA system over the entire service duration by jointly optimizing the passive beamforming of the RIS, the power allocation of the satellite, and the NOMA user pairing in each time slot. Therefore, the joint optimization problem (corresponding to maximizing the cumulative reward of the agent) can be expressed as:

[0151]

[0152] The first constraint is the phase variation range of the RIS reflection unit. represents the total transmission power of the satellite, and the second constraint indicates that the power distribution between NOMA user pairs is evenly distributed. The third and fourth constraints are constraints on SIC decoding. The fifth to eighth constraints are constraints on the NOMA pairing matrix, which means that in a certain time slot, all users cannot be paired with themselves and can only be paired with one other user. The ninth constraint represents the effective rate obtained to meet QoS. This joint optimization problem has both high-dimensional continuous optimization variables and discrete optimization variables, and is a long-term effective data rate maximization problem, that is, the phase of each reflection unit of RIS, the power distribution of the satellite, and the NOMA user pairing need to be reconfigured in each time slot, which is difficult to solve using traditional optimization methods.

[0153] The above embodiments of the present application fuse reconfigurable intelligent surface assisted satellite downlink communication and non-orthogonal multiple access technology, consider a reconfigurable intelligent surface assisted satellite downlink NOMA communication network, consider both satellite-user and satellite-RIS-user communication links, the RIS plays a role of assisting satellite and user communication in the scenario and enhances the signal, based on the considered satellite-user and satellite-RIS-user communication links, the equivalent channel from the satellite to the user is represented, and then the signal-to-interference-and-noise ratio of each user after SIC is represented, and then the effective data rate is expressed based on the Shannon formula; an optimization problem with the goal of maximizing the average sum rate of the entire downlink NOMA system is formed, and the constraint conditions include the constraints on the phase shift of the RIS reflecting unit, the power allocation of the satellite and the NOMA user pairing.

[0154] In order to verify the effectiveness of the scheme (Scheme One) in improving the effective data rate, the algorithm proposed in the present application is compared with other traditional schemes, which are as follows:

[0155] (1) Scheme One: the scheme proposed in the present application, which jointly optimizes the passive beamforming of the RIS, the power allocation of the satellite and the NOMA user pairing based on the PPO-DQN algorithm.

[0156] (2) Scheme Two: In this scheme, the passive beamforming of the RIS and the NOMA user pairing are optimized by using the proposed PPO-DQN algorithm. The power within the NOMA pair is randomly allocated under the condition of satisfying the constraint condition.

[0157] (3) Scheme Three: In this scheme, the phases of the RIS reflecting units are randomly configured. In addition, the NOMA user pairing and the power allocation within the user pair are optimized by using the proposed PPO-DQN algorithm.

[0158] (4) Scheme Four: In this scheme, the RIS technology is not used, and only the direct link between the BS and the user is considered. The NOMA pairing uses the traditional far-and-near pairing scheme, and the PPO algorithm is used to optimize the power allocation within the NOMA user pair.

[0159] Figure 3 The cumulative reward curves of the four schemes under the parameter settings of 50 RIS reflecting units, 1GHz system bandwidth and 60W total satellite transmit power are plotted. It is observed from Figure 3 that the cumulative rewards of all schemes gradually converge to a stable state with the increase of the training rounds. After training to a certain round, the cumulative rewards of the four schemes increase rapidly and gradually converge to a stable state because the agent has learned a better action. It is worth noting that the cumulative reward of the proposed Scheme One is significantly greater than that of the other traditional schemes, which indicates that the proposed Scheme One will bring a higher average sum rate and proves the effectiveness of the proposed algorithm.

[0160] like Figure 4 As shown in Figure 2, the simulation results demonstrate the impact of the number of RIS reflectors on the average sum rate. First, it can be observed that the average sum rates of the first three schemes increase with the increase in the number of RIS reflectors. This is because more RIS reflectors can reflect more signal paths and signal power, thereby increasing the data rate. More importantly, the curve of Scheme 1 is always higher than that of the other three schemes, indicating that the proposed PPO-DQN algorithm, which jointly optimizes RIS passive beamforming, satellite power allocation, and NOMA user pairing, can achieve the highest average sum rate compared to traditional schemes. Since Scheme 4 does not use RIS technology, the number of RIS reflectors has no impact on its performance, and the curve of this scheme is a straight line.

[0161] like Figure 5 As shown in Figure 2, simulation results demonstrate the impact of system bandwidth on the effective data rate. First, it can be seen that the performance of all schemes improves with increasing system bandwidth. This is because, according to Shannon's equation, bandwidth increases throughput. Second, Scheme 1 consistently outperforms the other schemes, demonstrating that the proposed PPO-DQN algorithm can effectively improve the average sum rate.

[0162] The embodiment of the present application further provides a satellite-to-ground downlink resource allocation system, which includes:

[0163] A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent;

[0164] A power allocation module is used to obtain the power allocation within the user pair based on the power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair;

[0165] A user pairing module is used for the first agent to obtain a user pairing matrix based on the decision factor output by the second agent; the user pairing matrix is ​​used to store all user pairing results;

[0166] The maximization module is configured to maximize the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all users.

[0167] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any of the above-mentioned satellite-to-ground downlink resource allocation methods is implemented.

[0168] An embodiment of the present application further provides a computer-readable storage medium, which includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes any of the above-mentioned satellite-to-ground downlink resource allocation methods.

[0169] Those skilled in the art should clearly understand that, for the convenience and brevity of description, the specific working processes of the resource allocation system, electronic device and computer-readable storage medium described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present application can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present application should be included in the scope of protection of the claims of the present application.

Claims

1. A method for allocating resources for a satellite-to-ground downlink, characterized in that: include: Determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; defining the state used by the agent in the current training time step as the equivalent channel of the signal transmitted by the satellite to each of the users in the previous time slot; Taking the phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair as the action space of the first intelligent agent; All possible values ​​of the decision factor corresponding to the user pairing matrix are used as the action space of the second agent; setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot; the effective data rate is determined according to the signal-to-interference-and-noise ratio and bandwidth of all users in each of the user pairs, and the signal-to-interference-and-noise ratio of each user is determined according to the power allocated by the satellite to each user in the user pair to which the user belongs and the equivalent channel of the signal transmitted by the satellite to each user in the user pair; Obtaining a power allocation within the user pair according to the power allocation ratio coefficient of the user pair; wherein the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; The second agent trains all possible user pairing results and outputs the decision factors; each decision factor corresponds to each user pairing result one by one; The first agent obtains a user pairing matrix based on the decision factor output by the second agent; The user pairing matrix is ​​used to store all user pairing results; Maximizing the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix; wherein the rewards are associated with the effective data rates corresponding to all the users; The method for obtaining the equivalent channel comprises: obtaining the equivalent channel of the signal transmitted by the satellite to each of the users according to the reflection unit in the reconfigurable smart surface; The step of obtaining the equivalent channel of the signal transmitted by the satellite to each user according to the reflective unit in the reconfigurable smart surface comprises: Taking any one of the reflective units in the reconfigurable smart surface as a reference, and according to the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface, obtaining a first component of the path loss of the link from the satellite to the reconfigurable smart surface; the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface is determined according to the relative position between each reflective unit in the reconfigurable smart surface and the reference, and the distance between the satellite and the reconfigurable smart surface; Obtaining a second component of the path loss of a link from the reconfigurable smart surface to the user based on a propagation direction of a signal transmitted by the satellite after being reflected by the reconfigurable smart surface to the user; the propagation direction of the signal transmitted by the satellite after being reflected by the reconfigurable smart surface to the user is determined based on a relative position between the user and the reference and a distance from the satellite to the user; The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the first component, the second component, the channel from the satellite to the user, and the reflection phase diagonal matrix of the reconfigurable smart surface.

2. The satellite-to-ground downlink resource allocation method according to claim 1, characterized in that: The training time step corresponds to the time slot; a single communication cycle of the satellite is evenly divided to obtain a plurality of the time slots.

3. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: The step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If the signal-to-interference-and-noise ratio (SINR) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is greater than a preset threshold, rewards are set for the first agent and the second agent in the current training time step according to the effective data rate corresponding to all the users in the current time slot; wherein the signal transmitted by the satellite received by each user is the signal transmitted by the satellite and transmitted through the equivalent channel to each user.

4. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: The step of setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot includes: If a signal-to-interference-plus-noise ratio (SIN) corresponding to the signal transmitted by the satellite received by all the users in the current time slot is at least partially less than a preset threshold, a penalty coefficient is set based on the SIN corresponding to the signal transmitted by the satellite received by each user, and a reward is set for the first agent and the second agent in the current training time step based on the penalty coefficient and the effective data rate corresponding to all the users in the current time slot; the penalty coefficient is set based on the difference between the preset threshold and the SIN below the preset threshold; wherein the signal transmitted by the satellite received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user.

5. The method for allocating satellite-to-ground downlink resources according to claim 1, wherein: The step of obtaining the power allocation within the user pair according to the power allocation ratio coefficient of the user pair includes: determining a total power allocated by the satellite to each of the user pairs; According to the total power of each user pair and the power allocation ratio coefficient of each user pair, the power corresponding to each user in each user pair is obtained, that is, the power allocated by the satellite to each user in each user pair.

6. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 5, characterized in that: Also includes: In the current training time step, after the second neural network corresponding to the second agent calculates the reward corresponding to the action it performs, the second agent transitions to the next state; Training the second neural network corresponding to the second agent and updating its network parameters using the state corresponding to the second agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; The second neural network is used to instruct the second agent to select and execute corresponding actions.

7. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 5, characterized in that: Also includes: In the current training time step, after the first neural network corresponding to the first agent calculates the reward corresponding to the action performed by the first agent, the first agent transitions to the next state; The first neural network corresponding to the first agent is trained and its network parameters are updated using the state corresponding to the first agent in the current training time step, the user pairing matrix and the reward, and the state corresponding to the next training time step; the first neural network is used to instruct the first agent to select and execute a corresponding action.

8. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 5, characterized in that: Maximizing the cumulative rewards corresponding to the first agent and the second agent to obtain the optimal phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix includes: In the current training time step, the cumulative rewards corresponding to the first agent and the second agent are maximized respectively to obtain the phase of each reflective unit in the reconfigurable smart surface, the power allocation of the satellite, and the user pairing matrix in the current time slot.

9. The satellite-to-ground downlink resource allocation method according to claim 1, wherein: Obtaining the equivalent channel from the signal transmitted by the satellite to each of the users according to the first component, the second component, the channel from the satellite to the user, and the reflection phase diagonal matrix of the reconfigurable smart surface includes: Obtaining a channel from the satellite to the reconfigurable smart surface according to the first component and a distance between the satellite and the reconfigurable smart surface; Obtaining a channel from the reconfigurable smart surface to the user according to the second component and a distance between the reconfigurable smart surface and the user; Obtaining a channel from the satellite to the user according to a distance between the satellite and the user; The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the channel from the satellite to the reconfigurable smart surface, the channel from the reconfigurable smart surface to the user, the channel from the satellite to the user, and the reflection phase diagonal matrix of the reconfigurable smart surface.

10. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 5, characterized in that: The method for obtaining the effective data rate includes: obtaining the effective data rate of the signal transmitted by the satellite to each of the users according to the equivalent channel of the signal transmitted by the satellite to each of the users; The obtaining, based on the equivalent channel of the signal transmitted by the satellite to each user, the effective data rate at which the satellite transmits the signal to each user, comprises: Obtaining a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user based on the power allocated by the satellite to each user and the equivalent channel corresponding to each user; the signal sent by the satellite and received by each user is a signal transmitted by the satellite and then transmitted through the equivalent channel to reach each user; An effective data rate at which the satellite transmits data to the user is obtained according to a signal-to-interference-and-noise ratio corresponding to the signal sent by the satellite and received by each user.

11. The satellite-to-ground downlink resource allocation method according to any one of claims 1 to 5, characterized in that: All users within the satellite coverage area are divided into multiple user pairs; the pairing results of the multiple user pairs constitute the user pairing matrix; each user pair includes two users; and users in the same user pair share the same time-frequency resource block.

12. A satellite-to-ground downlink resource allocation system, characterized in that: include: A state definition module is used to determine the state of an intelligent agent based on equivalent channels from the satellite to all users; the intelligent agent includes a first intelligent agent and a second intelligent agent; defining the state used by the agent in the current training time step as the equivalent channel of the signal transmitted by the satellite to each of the users in the previous time slot; Taking the phase variation range of each reflective unit in the reconfigurable smart surface and the power allocation ratio coefficient of the user pair as the action space of the first intelligent agent; All possible values ​​of the decision factor corresponding to the user pairing matrix are used as the action space of the second agent; setting rewards for the first agent and the second agent respectively according to the effective data rates corresponding to all the users in the current time slot; the effective data rate is determined according to the signal-to-interference-and-noise ratio and bandwidth of all users in each of the user pairs, and the signal-to-interference-and-noise ratio of each user is determined according to the power allocated by the satellite to each user in the user pair to which the user belongs and the equivalent channel of the signal transmitted by the satellite to each user in the user pair; A power allocation module, configured to obtain a power allocation within a user pair based on a power allocation ratio coefficient of the user pair; the power allocation within the user pair is used to indicate the power allocated by the satellite to each user in each user pair; The second agent trains all possible user pairing results and outputs the decision factors; each decision factor corresponds to each user pairing result one by one; A user pairing module, configured for the first agent to obtain a user pairing matrix based on the decision factor output by the second agent; The user pairing matrix is ​​used to store all user pairing results; a maximization module, configured to maximize the cumulative rewards corresponding to the first intelligent agent and the second intelligent agent to obtain an optimal phase of each reflective unit in the reconfigurable intelligent surface, power allocation of the satellite, and user pairing matrix; wherein the rewards are associated with effective data rates corresponding to all the users; The method for obtaining the equivalent channel comprises: obtaining the equivalent channel of the signal transmitted by the satellite to each of the users according to the reflection unit in the reconfigurable smart surface; The step of obtaining the equivalent channel of the signal transmitted by the satellite to each user according to the reflective unit in the reconfigurable smart surface comprises: Taking any one of the reflective units in the reconfigurable smart surface as a reference, and according to the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface, obtaining a first component of the path loss of the link from the satellite to the reconfigurable smart surface; the propagation direction of the signal transmitted by the satellite to the reconfigurable smart surface is determined according to the relative position between each reflective unit in the reconfigurable smart surface and the reference, and the distance between the satellite and the reconfigurable smart surface; Obtaining a second component of the path loss of a link from the reconfigurable smart surface to the user based on a propagation direction of a signal transmitted by the satellite after being reflected by the reconfigurable smart surface to the user; the propagation direction of the signal transmitted by the satellite after being reflected by the reconfigurable smart surface to the user is determined based on a relative position between the user and the reference and a distance from the satellite to the user; The equivalent channel from the signal transmitted by the satellite to each of the users is obtained according to the first component, the second component, the channel from the satellite to the user, and the reflection phase diagonal matrix of the reconfigurable smart surface.

13. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the satellite-to-ground downlink resource allocation method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer is enabled to execute the satellite-to-ground downlink resource allocation method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Resource allocation method in D2D-NOMA uplink communication system based on RIS

    CN118158802A

  • Method for optimizing intelligent Internet of Vehicles communication system model based on double-relay RIS

    CN119865267A