Satellite communication method and system based on rate division multiple access

By introducing rate segmentation multiple access technology and improved DQN algorithm in satellite communication systems, the rate allocation is dynamically adjusted, and the problems of low spectrum utilization and insufficient interference management in traditional satellite communication systems are solved, and the total reachable rate of the system is maximized and the robustness of network performance is improved.

CN120263277AActive Publication Date: 2025-07-04PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510734099.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Traditional satellite communication systems have limitations in terms of low spectrum utilization, insufficient interference management capabilities and inflexible resource allocation, and are difficult to adapt to scenarios where user needs are dynamic changes and business needs are highly explosive.

Method used

Rate segmented multiple access (RSMA) technology is used to build a satellite communication system model through stationary orbit satellites, high-altitude communication platforms and ground user terminals, and the improved DQN algorithm is used to optimize rate allocation, combining Markov decision-making process and multi-agent collaboration mechanism to dynamically adjust the rate allocation coefficients of public information flow and private information flow.

Benefits of technology

It improves the spectrum utilization rate and network performance of satellite communication systems, can effectively deal with channel quality fluctuations caused by user mobility and terrain occlusion, and maximizes the total reachable rate of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263277A_ABST
    Figure CN120263277A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of satellite communication, and particularly discloses a satellite communication method and system based on rate division multiple access, and the method comprises the steps: building a satellite communication system model through a stationary orbit satellite, a high-altitude communication platform and a ground user terminal; determining an optimization problem according to a satellite communication system model by taking maximization of the total reachable rate of the system as a target and taking power budget of a transmitter and the lowest rate of energy distribution as constraint conditions; modeling the optimization problem into a Markov decision process through a state space, an action space and a reward function; and solving a Markov decision process by using an improved DQN algorithm, and determining an optimized rate distribution coefficient based on the solution of the Markov decision process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite communication, and particularly relates to a satellite communication method and system based on rate-splitting multiple access. Background Art

[0002] The rate-splitting multiple access (RSMA) technology has been introduced into satellite communication systems mainly to address the limitations of traditional multiple access technologies in aspects such as spectral efficiency, dynamic resource allocation, and interference management. The following discussions are from two aspects: the disadvantages of traditional technologies and the advantages of RSMA.

[0003] (1) Disadvantages of traditional multiple access technologies With the access requirements of a large number of ground devices, the satellite communication system, as an infrastructure that can achieve seamless connection, has become the core of the construction of the next-generation wireless communication system. Early satellite communication systems mainly adopted multiple access technologies such as frequency-division multiple access (FDMA), time-division multiple access (TDMA), code-division multiple access (CDMA), and later orthogonal frequency-division multiple access (OFDMA) and non-orthogonal multiple access (NOMA) were also introduced. However, these technologies have the following problems in satellite scenarios: 1) The first problem is the low spectral utilization rate. For example, both FDMA and TDMA allocate spectrum resources by dividing fixed frequency bands or time slots, and users exclusively occupy the allocated spectrum resource units, resulting in low spectral utilization rate and being difficult to adapt to the characteristics of dynamically changing user demands in satellite communication. The CDMA access technology relies on spreading codes to distinguish users. When the number of users increases, the multiple access interference (MAI) increases significantly, and complex power control is required. The control is difficult in an environment with cross transmission delays such as satellite-ground.

[0004] 2) The second problem is the insufficient interference management ability. In a satellite communication system based on NOMA, although NOMA superimposes user signals in the power domain and uses successive interference cancellation (SIC) to improve the system capacity, in the satellite channel, the differences between user channels are small (for example, a geostationary satellite has a wide coverage and limited differences in user path losses), resulting in poor power multiplexing effects and prominent problems of SIC decoding error accumulation.

[0005] 3) Other problems include that traditional OFDMA technology is sensitive to frequency offset and phase noise and requires further development of complex equalization algorithms, which greatly increases the system overhead. Moreover, traditional technologies usually rely on static or semi-static resource allocation and are difficult to adapt to scenarios where user distribution is uneven and service demands are highly bursty (such as intermittent transmission of Internet of Things devices) in satellite communication, resulting in resource waste or congestion.

[0006] Therefore, based on the above multiple access technologies, how to use the Rate-Splitting Multiple Access (RSMA) technology to enhance the transmission performance of satellite communication systems is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0007] To achieve the object of the present invention, the present application provides a satellite communication method based on rate-splitting multiple access, including: Step S1: Construct a satellite communication system model by using a geostationary satellite, a high-altitude communication platform, and a ground user terminal; Step S2: Taking the maximization of the total achievable rate of the system as the objective and the power budget of the transmitter and the minimum rate of energy allocation as the constraint conditions, determine the optimization problem according to the satellite communication system model; Step S3: Model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Step S4: Use an improved DQN algorithm to solve the Markov decision process; Step S5: Determine the optimized rate allocation coefficient based on the solution of the Markov decision process.

[0008] In some specific embodiments, Step S1 includes: Step S11: The satellite emits a beam carrying data to a near-space node through a coherent laser link; Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident beam, and converts the modulated optical signal into a radio frequency signal; Step S13: The ground user terminal uses single-layer RSMA technology to receive the signal of the high-altitude communication platform through a direct transmission link.

[0009] In some specific embodiments, Step S2 includes: The optimization problem is determined according to the following formula: ; In the formula, represents the total achievable rate of the system, represents the received signal rate of the common information stream of the i th ground user, represents the rate splitting coefficient, w k represents the active transmission beamforming matrix of the high-altitude communication platform, represents the received signal rate of the private information stream of the i th ground user, c i represents the data rate of the i th ground user receiving the common information, and Indicates that the common information can be successfully decoded by each ground user, in denotes the achievable rate requirement of the th ground user, denotes the i th ground user's energy allocation requirement is greater than J i , denotes the maximum transmit power budget at the transmitter is less than , denotes the rate splitting coefficient limit.

[0010] In some specific embodiments, step S3 includes: Determining the state space according to the information set of all current ground user terminals, the action vectors selected by all current ground user terminals, and the immediate rewards in the local training phase of all current ground user terminals; Determining the action space according to the active transmit beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term in an unconstrained case and a penalty term under constraint satisfaction.

[0011] In some specific embodiments, step S4 includes: Step S41: Defining the direct link from each high-altitude communication platform to the ground user terminal as an autonomous decision-making unit, enabling each autonomous decision-making unit to have local state observation ability, independent action space, and personalized reward function; Step S42: Each autonomous decision-making unit shares a set of deep neural network parameters, learns a collaborative policy through a global experience pool, and broadcasts the deep neural network parameters of the trained policy to each agent; Step S43: According to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establishing a dual-mode parameter update mechanism, and updating the deep neural network parameters using mini-batch gradient descent; Step S44: Using an improved DQN algorithm to optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform; Step S45: Solving the Markov decision process according to the deep neural network parameters, the transmit beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

[0012] To achieve the same invention purpose, the present application also provides a satellite communication system based on rate splitting multiple access, including: Model construction module: used to construct a satellite communication system model using a geostationary satellite, a high-altitude communication platform, and a ground user terminal; Optimization problem determination module: used to determine an optimization problem according to the satellite communication system model with the goal of maximizing the total achievable rate of the system and subject to the power budget of the transmitter and the minimum rate of energy allocation; Decision process establishment module: used to model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Process solving module: used to solve the Markov decision process using an improved DQN algorithm; Solution verification module: used to determine an optimized rate allocation coefficient based on the solution of the Markov decision process.

[0013] In some specific embodiments, the model construction module is used to perform the following steps: Step S11: The satellite transmits a data-bearing light beam to the near-space node through a coherent laser link; Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident light beam, and converts the modulated optical signal into a radio frequency signal; Step S13: The ground user terminal receives the high-altitude communication platform signal through a direct transmission link using single-layer RSMA technology.

[0014] In some specific embodiments, in the model construction module, the optimization problem is determined according to the following formula: ; In the formula, represents the total achievable rate of the system, represents the received signal rate of the common information flow of the i th ground user, represents the rate splitting coefficient, w k represents the active transmission beamforming matrix of the high-altitude communication platform, represents the received signal rate of the private information flow of the i th ground user, c i represents the data rate of the i th ground user receiving the common information, and indicate that the common information can be successfully decoded by each ground user, in represents the achievable rate requirement of the th ground user, represents the energy allocation requirement of the i th ground user is greater than J i , represents the maximum transmit power budget at the transmitter is less than , represents the rate splitting coefficient limit.

[0015] In some specific embodiments, the decision process establishment module is configured to perform the following steps: Determine the state space according to the information set of all current ground user terminals, the action vectors selected by all current ground user terminals, and the immediate rewards in the local training phase of all current ground user terminals; Determine the action space according to the active transmission beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term under unconstrained conditions and a penalty term under satisfied constraint conditions.

[0016] In some specific embodiments, the process solving module is configured to perform the following steps: Step S41: Define the direct link from each high-altitude communication platform to a ground user terminal as an autonomous decision-making unit, so that each autonomous decision-making unit has the capabilities of local state observation, independent action space, and personalized reward function; Step S42: Each autonomous decision-making unit shares a set of deep neural network parameters, learns a collaborative policy through a global experience pool, and broadcasts the deep neural network parameters of the trained policy to each agent; Step S43: According to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establish a dual-mode parameter update mechanism, and use mini-batch gradient descent to update the deep neural network parameters; Step S44: Use the improved DQN algorithm to optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform; Step S45: Solve the Markov decision process according to the deep neural network parameters, the transmit beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

[0017] Advantages of the above technical solution: The method provided by this application breaks through the idealized assumption of the channel model in traditional beamforming algorithms, and establishes an end-to-end mapping relationship from the original channel state information to the beam weights through the DQN algorithm. Different from the fixed rate allocation mode in traditional systems, a rate splitting multiple access strategy multi-agent cooperation mechanism is introduced, and each agent corresponds to the rate splitting decision of a specific user or traffic flow. By defining a joint state space that includes user QoS requirements, service type priorities, and network load status, the DQN agent can learn how to dynamically adjust the rate allocation coefficient between the common information flow and the private information flow under different network conditions.

[0018] This application also proposes to model the communication process among satellites, high-altitude communication platforms, and ground terminals as a heterogeneous Markov decision process, where the beamforming decisions of high-altitude platforms and the power control strategies of ground terminals influence each other. By establishing a cross-air-layer joint reward function, the system can evaluate the overall network performance of the air-ground cooperation strategy, such as total throughput, energy efficiency, etc. In actual deployment, this cooperation framework can effectively cope with channel quality fluctuations caused by factors such as user mobility and terrain occlusion, and improve the robustness of network performance through the intelligent cooperation of air-ground nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0020] Figure 1 FIG. is a schematic flowchart of a satellite communication method based on rate-splitting multiple access according to an embodiment of the present invention; Figure 2 FIG. is a schematic structural diagram of a satellite communication system based on rate-splitting multiple access according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them.

[0022] Examples of the embodiments are shown in the drawings, where the same or similar reference signs represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0023] Embodiment 1 An embodiment of the present invention provides a satellite communication method based on rate-splitting multiple access. Referring to Figure 1 as shown, it includes: Step S1: Construct a satellite communication system model using a geostationary satellite, a high-altitude communication platform, and a ground user terminal; Step S2: With the goal of maximizing the total achievable rate of the system and subject to the power budget of the transmitter and the minimum rate of energy allocation as constraints, determine an optimization problem according to the satellite communication system model; Step S3: Model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Step S4: Solve the Markov decision process using an improved DQN algorithm.

[0024] In a specific embodiment of the present invention, Step S1 includes: Step S11: The satellite transmits a light beam carrying data to a near-space node through a coherent laser link; Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident light beam, and converts the modulated optical signal into a radio frequency signal; Step S13: The ground user terminal receives the signal of the high-altitude communication platform through a direct transmission link using single-layer RSMA technology.

[0025] The system model proposed by the present invention consists of a geostationary satellite, a high-altitude communication platform, and a ground user terminal. Its signal transmission mechanism is divided into two core stages according to the characteristics of the transmission medium: In the initial transmission stage, the satellite transmits an optical wave beam carrying data to a near-space node through a coherent laser link ; The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident light beam, and converts the modulated optical signal into a radio frequency signal through an optoelectronic conversion module. This conversion process follows the cascaded transformation principle of light-electricity-radio frequency, and its output signal can be characterized as: ; where, is the satellite transmission power, represents the conversion coefficient of the optical signal to the radio frequency signal, represents the environmental noise, represents the scalar channel fading coefficient from the satellite to the high-altitude communication platform. In addition, to solve the potential errors caused by atmospheric turbulence, an Equal Gain Combining (EGC) scheme is adopted at the high-altitude communication platform.

[0026] In the second transmission stage, the high-altitude communication platform applies single-layer RSMA technology, and the ground user receives the signal of the high-altitude communication platform through a direct transmission link.

[0027] Assume that the associated signal sent by the high-altitude communication platform to the th ground user is divided into two components, namely the common information part and the private information part . Then, all the common information parts are combined into , encoded into a common information stream . The The private part of each ground user terminal is separately encoded into a private information stream , satisfying . That is to say, the signal sent by the high-altitude communication platform is a common information stream and private information streams, totaling symbols, which reach the ground user equipment through the direct link. Then the signal sent by the high-altitude communication platform can be expressed as: ; In the formula, represents the active transmission beam vector of the th information stream . Therefore, the total transmission power at the high-altitude communication platform is expressed as: ; In the formula, and are the efficiency parameter of the power amplifier and the power consumption of each circuit module respectively. The time interval in the system is analyzed in a discrete form. In the th time slot, the signal received by the th ground user is expressed as: ; In the formula, represents the channel gain from the high-altitude communication platform to the th ground user.

[0028] In addition, the power splitting ratio of the th ground user is defined as . This ratio is crucial for how the received signal divides the energy allocation and signal transmission. It should be noted that the common information stream will serve as an effective energy allocation transmission carrier, be broadcast to all users, and will not interfere with the decoding of the private information stream. Specifically, part of the signal power is used for user identification, and the remaining part is used for energy allocation.

[0029] Therefore, in the RSMA-based satellite communication system, ground users adopt a layered decoding strategy to achieve information separation. Specifically, each user terminal first treats the private data streams of other users as additive interference terms and preferentially decodes the common information stream. This processing mechanism is based on the differences in dual information attributes: the common information stream, as the basic service data shared across the network, needs to ensure accessibility for all users; the private information stream, on the other hand, carries personalized services and only has payload value for the target user. After decoding the common information stream, the receiver reconstructs the common signal component through the successive interference cancellation (SIC) technique and projects and strips it in the received signal space. At this time, the channel observation model is transformed into an enhanced observation that only contains the private information of the target user, and then the second decoding operation is performed. Therefore, the received signal rates of the common information stream and the private information stream of the th ground user are respectively expressed as: ; and: ; where, represents the group channel coefficient between the high-altitude communication platform and the ground user terminal. To ensure that the common information stream can be successfully decoded by all ground users, the rate of the public information stream can be set to . Then, let be the data rate of receiving the common information at the th ground user, which needs to satisfy: ; Then, the total achievable rate of the considered system can be expressed as: ; In the formula, represents the system ergodic sum rate of the FSO link. Here, it is defaulted that the system ergodic sum rate of the FSO link is higher than that of the RF link. This means that the signal mainly experiences free space attenuation during transmission and is less blocked or reflected by other objects. Therefore, compared with the radio frequency transmission link, the signal loss of the FSO link is smaller, and it can more effectively maintain the signal strength and quality. Secondly, the radio frequency transmission link is limited by the physical characteristics of radio frequency signals. Radio frequency signals are affected by various factors during transmission, such as multipath effects, signal attenuation, interference, etc. These factors will reduce the modulation efficiency and spectrum utilization rate of the radio frequency transmission link.

[0030] In a specific embodiment of the present invention, the core optimization problem of the present invention is to maximize the total achievable rate of the system, aiming to ensure that the system can achieve the best energy utilization state while meeting various operation constraints. These constraints mainly include the power budget of the transmitter and the minimum rate requirements for energy allocation. To solve this optimization problem, a refined algorithm is planned to optimize multiple key parameters. These parameters include the beamforming matrix for active transmission of the high-altitude communication platform , and the rate splitting coefficient . The optimization problem is determined according to the following formula: ; In the formula, represents the total achievable rate of the system, represents the received signal rate of the common information flow of the i -th ground user, represents the rate splitting coefficient, w k represents the beamforming matrix for active transmission of the high-altitude communication platform, represents the received signal rate of the private information flow of the i -th ground user, c i represents the data rate of the common information received by the i -th ground user, and represent that the common information can be successfully decoded by each ground user, in represents the achievable rate requirement of the -th ground user, represents the energy allocation requirement of the i -th ground user is greater than J i , represents the maximum transmit power budget at the transmitter is less than , represents the rate splitting coefficient limit.

[0031] In a specific embodiment of the present invention, step S3 includes: Determining the state space according to the information set of all current ground user terminals, the action vectors selected by all current ground user terminals, and the immediate rewards in the local training phase of all current ground user terminals; Determining the action space according to the active transmission beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term in the unconstrained case and a penalty term under the satisfied constraint conditions.

[0032] Specifically, the above optimization problem can be modeled as a Markov decision process (MDP), and the specific content includes: 1) State space: According to the design principle of the state space, it should include as much environmental information related to the optimization problem as possible. In the considered system model, the state space should consist of three parts. The information set of all current ground user terminals , the selected action vector and the immediate reward in the local training phase . The information set of all current users is defined as: (4-24) In the formula, and respectively represent the received signal rates of the common information flow and the private information flow of the th ground user in the previous moment. The selected action vector represents the action selected using the improved DQN algorithm starting from the previous moment. The immediate reward is calculated from the current state and the current action, and this metric can directly reflect the ability of the Agent to solve the optimization problem in different situations. Therefore, the state space of the training node is represented as: ; 2) Action space: The action space in the local training phase mainly includes the active transmit beam vector , the power splitting coefficient and the achievable rate of the common information flow , etc. When using the improved DQN algorithm to solve the problem, the input information needs to be input into the Q network. Therefore, in the input or output of the neural network, the complex form needs to be split into the real part and the imaginary part and used as inputs respectively. Therefore, the active transmit beam vector at the high-altitude communication platform is split into: ; In the formula, and respectively represent the modulus of the corresponding transmit power and the unit length representing the beam direction. By such decomposition, the system can control the transmit signal strength and propagation direction separately. Here, the hyperbolic tangent function is used to smoothly vary the power output within a predetermined range. This is because the hyperbolic tangent function is a non-linear function, and its output value ranges from -1 to 1. At the same time, this function also has a smooth transition characteristic, that is, when the input value increases from negative infinity to positive infinity, the output value smoothly transitions from -1 to 1, avoiding sudden changes. Therefore, it is represented as: ; Pmax is the maximum transmit power at the transmitter. The main purpose of this formula is to limit the output of the transmit power within a specified range after meeting the corresponding constraints. And represents the output response of the activation function to the selected transmit power.

[0033] In a satellite communication system applying RSMA, the common information stream will be transmitted as the main energy carrier. Under this scheme, the decoding process of the common information stream will not interfere with the reception of the private information stream due to the successive interference cancellation technique. Therefore, the beamforming direction of the common information stream can adopt the maximum ratio combining (MRT) scheme. This method can maximize the power of the signal at the receiving end, thereby improving the reliability of transmission. For the private information stream, considering the need to reduce interference between users, a zero-forcing transmission scheme can be adopted. The zero-forcing transmission scheme precisely controls the direction of the transmitted signal to ensure that the signal interference is forced to zero in the direction of non-target users, thereby effectively reducing the interference between users and improving the signal quality and the overall performance of the system. Therefore, the active transmit power direction can be specifically expressed as: (4-28) In the formula, represents the combined channel coefficient from the high-altitude communication platform to the ground user terminal. Define , and represents the beam direction under the combined channel. At this time . On the other hand, the power splitting coefficient and the achievable rate of the common information stream can also be calculated using the hyperbolic tangent function, which is expressed as: ; and: ; In the formula, and are the achievable rate of the common information stream and the power splitting coefficient output under the activation function respectively. Therefore, the action space is expressed as: ; 3) Reward function: In the process of using the DQN algorithm to solve the optimization problem, the reward function is a key content to be considered in the design process. It not only needs to consider the optimization goal but also must take into account the constraints in the optimization problem P0. Usually, the reward function mainly consists of two parts: the first part is the immediate reward term reflecting the unconstrained situation, and the other part is the penalty term ensuring that the constraints are satisfied. To balance the relationship between the two, a demand-aware reward function needs to be carefully designed. In this section, the reward function can be expressed as: ; In the formula, , and respectively represent the constraint penalty terms in the corresponding optimization problem, and respectively represent the action penalty results under the conditions of not meeting the transmit power requirement, the common information decoding requirement, and the quality of service requirement. They can be respectively modeled as: ; ; .

[0034] In the constrained optimization problem, the design of the penalty function plays a key role in the convergence of the algorithm. By constructing a boundary barrier function and imposing dynamic weight penalties on the iteration points outside the feasible region, an effective constraint coupling mechanism can be established, enabling the optimization trajectory to naturally tend to the feasible solution space during the iteration process. This strategy significantly improves the global search ability of the algorithm under strict constraint conditions, ensuring that a physically feasible Pareto optimal solution can still be stably obtained near the complex constraint boundary. In the optimization framework of the satellite communication system based on reinforcement learning, the construction of the reward function embodies the concept of multi-objective collaborative design. Only when all quality of service (QoS) constraints and energy consumption limits are met, the system assigns a non-zero reward value. This threshold feedback mechanism strengthens the priority of constraint satisfaction. For the effective actions to improve the system achievable rate, the reward function adopts the principle of diminishing marginal benefit and gives differential positive incentives according to the magnitude of the rate increase, guiding the intelligent agent to form an adaptive balance between energy utilization efficiency and communication performance.

[0035] The design of this reward and punishment mechanism follows the dual-track regulation principle: the penalty term forms a repulsive force field through the non-linear mapping of the constraint violation degree to prevent the policy exploration from entering the infeasible region; the reward term constructs an attractive force field based on the improvement of the performance index to drive the intelligent agent to evolve towards the Pareto front. This two-way regulation mechanism realizes the organic unity of constraint satisfaction and performance optimization while ensuring the robustness of the system, providing a theoretically feasible optimization framework for the autonomous decision-making of complex communication systems.

[0036] In a specific embodiment of the present invention, step S4 includes: Step S41: Define the direct link from each high-altitude communication platform to the ground user terminal as an autonomous decision-making unit, enabling each autonomous decision-making unit to have local state observation ability, an independent action space, and a personalized reward function; Step S42: Each autonomous decision-making unit shares a set of deep neural network parameters, learns the collaborative strategy through the global experience pool, and broadcasts the deep neural network parameters of the trained strategy to each intelligent agent; Step S43: Establish a dual-mode parameter update mechanism according to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, and update the deep neural network parameters using mini-batch gradient descent; Step S44: Optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform using the improved DQN algorithm; Step S45: Solve the Markov decision process based on the deep neural network parameters, the transmit beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

[0037] Benefiting from the function approximation ability of the deep neural network in DQN, the proposed DQN framework can handle problems with huge state spaces and action spaces. However, since DQN only trains one neural network, there will be problems with unstable convergence speed. Therefore, to accelerate the convergence effect of the algorithm and the stability of the DQN algorithm, considering the future satellite communication scenario of large-scale ground user access, an improved DQN algorithm is proposed. Its core idea is as follows: (1) First, define the direct link of each high-altitude communication platform-ground user terminal as an autonomous decision-making unit. Each autonomous decision-making unit needs to have local state observation capabilities (including link quality, transmission status, etc.), an independent action space (power allocation decision), and a personalized reward function (based on local QoS satisfaction); (2) Then, all parameters adopt a hybrid paradigm of "centralized training - distributed execution". Specifically, in the training phase: all agents share the same set of deep neural network parameters , and learn collaborative strategies through the global experience pool. In particular, when the discount factor , the network degenerates into a first-order Markov decision process. At this time, only a single Q-training network needs to be maintained, significantly reducing the non-stationarity problem in the multi-agent system. In the execution phase: the trained policy parameters are broadcast to each agent to achieve fully distributed decision-making. Each node can complete real-time resource scheduling only relying on local observations; (3) Finally, in response to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establish a dual-mode parameter update mechanism, and use mini-batch gradient descent to finely adjust the deep neural network parameters , triggering the exploration mode to enhance the robustness of the policy.

[0038] The complete process of using the improved DQN algorithm to optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform can be divided into the following steps: ① Input stage: Initialize the experience pool, maximize the number of training rounds, and the maximum number of communication time slots ; ② Training stage:

[0039] ③ Output stage: After optimization 。

[0040] Specifically, for the proposed improved DQN optimization framework, its state space is composed of the joint representation of the refined channel state vector and environmental parameters, including multi-dimensional parameters such as instantaneous channel gain, interference power distribution, user movement pattern, and time-varying characteristics of network topology. The state representation is formed by dynamically adjusted tensor splicing. The network architecture adopts a deep residual design. The dimension of the input layer is adaptively adjusted according to the scale of state variables. The neurons in the hidden layer expand according to the exponential rule. The forward propagation complexity of a single time step is a product function of the square of state variables and the number of network layers. The global training complexity further considers the trajectory length and training cycles. The reward function design adopts a dual-factor mechanism. The main term is based on the logarithmic measure of the secrecy rate of the signal-to-interference ratio, and the constraint term punishes power overstep through an exponential function. The two achieve dynamic balance through the Lagrange multiplier method. When retraining, historical policy experience is retained through knowledge distillation. The progressive forgetting mechanism makes the influence of historical experience decay over time. At the same time, a dual-mode mechanism is used to pre-train hyperparameters to achieve fast policy adaptation in the new environment. Simulation experiments show that this mechanism can approach the performance level of traditional retraining with a small number of parameter updates.

[0041] Step S5: Determine the optimized rate allocation coefficient based on the solution of the Markov decision process.

[0042] In the aspect of dynamic beam optimization driven by reinforcement learning, this application breaks through the idealized assumption of the channel model in traditional beamforming algorithms and establishes an end-to-end mapping relationship from the original channel state information to the beam weights through the DQN algorithm. The core of the algorithm lies in designing a hybrid network architecture that includes spatial feature extraction and decision fusion. Among them, the convolutional neural network is responsible for processing the spatial correlation of the multi-dimensional channel matrix, while the fully connected layer generates the beamforming vector based on the extracted features. This data-driven optimization method not only avoids the computational overhead brought by complex matrix operations, but also can adaptively capture non-ideal factors in the actual communication environment, such as atmospheric attenuation, multipath effect, etc., so as to realize the online evolution of beam patterns in the high-altitude communication scenario with dynamic topology.

[0043] Different from the fixed rate allocation mode in traditional systems, this application introduces a rate splitting multiple access strategy multi-agent cooperation mechanism, and each agent corresponds to the rate splitting decision of a specific user or traffic flow. By defining a joint state space that includes user QoS requirements, service type priorities, and network load status, the DQN agent can learn how to dynamically adjust the rate allocation coefficient between the common information flow and the private information flow under different network conditions.

[0044] The construction of the collaborative optimization framework for satellite communication systems marks the leap of this patent from single - technology optimization to system - level solutions. This application proposes to model the communication processes among satellites, high - altitude communication platforms, and ground terminals as a heterogeneous Markov decision process, where the beamforming decisions of high - altitude platforms and the power control strategies of ground terminals influence each other. By establishing a cross - altitude joint reward function, the system can evaluate the overall network performance of the air - ground collaborative strategy, such as total throughput, energy efficiency, etc. In actual deployment, this collaborative framework can effectively cope with channel quality fluctuations caused by factors such as user mobility and terrain occlusion, and through the intelligent cooperation of air - ground nodes, achieve a robust improvement in network performance.

[0045] For the unique strong - interference environment of high - altitude communication, this application innovatively proposes an anti - interference robustness enhancement design scheme. On the one hand, by leveraging the generalization ability of DQN, the interference avoidance mechanism is incorporated into the beamforming optimization process. By adjusting parameters such as beam direction and null - depth, while ensuring the communication quality of target users, the impact of interference signals on system performance is minimized. In addition, for the signal fading problem under multipath channel conditions, this application proposes an exploration - exploitation strategy based on dual - parameter update to intelligently search for the optimal parameter variation in the beamforming parameter space, significantly enhancing the anti - interference robustness of the system.

[0046] Embodiment 2 An embodiment of the present invention provides a satellite communication system based on rate - splitting multiple access. Referring to Figure 2 as shown, it includes: Model construction module 10: used to construct a satellite communication system model by using geostationary satellites, high - altitude communication platforms, and ground user terminals; Optimization problem determination module 20: used to determine the optimization problem according to the satellite communication system model with the goal of maximizing the total achievable rate of the system and with the power budget of the transmitter and the minimum rate of energy allocation as constraint conditions; Decision - making process establishment module 30: used to model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Process solving module 40: used to solve the Markov decision process using an improved DQN algorithm; Solution verification module 50: used to determine the optimized rate allocation coefficient based on the solution of the Markov decision process.

[0047] In a specific embodiment of the present invention, the model construction module 10 is used to perform the following steps: Step S11: The satellite transmits a data - bearing light beam to a near - space node through a coherent laser link; Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident light beam, and converts the modulated optical signal into a radio frequency signal; Step S13: The ground user terminal uses the single-layer RSMA technology to receive the high-altitude communication platform signal through the direct transmission link.

[0048] In a specific embodiment of the present invention, in the model construction module 20, the optimization problem is determined according to the following formula: ; In the formula, represents the total achievable rate of the system, represents the received signal rate of the common information flow of the i th ground user, represents the rate splitting coefficient, w k represents the beamforming matrix actively transmitted by the high-altitude communication platform, represents the received signal rate of the private information flow of the i th ground user, c i represents the data rate of the i th ground user receiving the common information, and represent that the common information can be successfully decoded by each ground user, in represents the achievable rate requirement of the th ground user, represents the i th ground user's energy allocation requirement is greater than J i , represents the maximum transmit power budget at the transmitter is less than , represents the rate splitting coefficient limit.

[0049] In a specific embodiment of the present invention, the decision process establishment module 30 is used to perform the following steps: Determine the state space according to the information set of all current ground user terminals, the action vectors selected by all current ground user terminals, and the immediate rewards in the local training phase of all current ground user terminals; Determine the action space according to the active transmit beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term under unconstrained conditions and a penalty term under constraint satisfaction conditions.

[0050] In a specific embodiment of the present invention, the process solving module 40 is used to perform the following steps: Step S41: Define the direct link from each high-altitude communication platform to the ground user terminal as an autonomous decision-making unit, enabling each autonomous decision-making unit to have the capabilities of local state observation, independent action space, and personalized reward function; Step S42: Each autonomous decision-making unit shares a set of deep neural network parameters, learns a collaborative policy through a global experience pool, and broadcasts the deep neural network parameters of the trained policy to each agent; Step S43: According to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establish a dual-mode parameter update mechanism, and use mini-batch gradient descent to update the deep neural network parameters; Step S44: Use the improved DQN algorithm to optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform; Step S45: Solve the Markov decision process according to the deep neural network parameters, the transmit beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

[0051] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0052] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.

[0053] The method and device provided by the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0054] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A satellite communication method based on rate-splitting multiple access, characterized in that, including: Step S1: Construct a satellite communication system model using a geostationary satellite, a high-altitude communication platform, and a ground user terminal; Step S2: With the goal of maximizing the total achievable rate of the system and subject to the power budget of the transmitter and the minimum rate of energy allocation as constraints, determine the optimization problem according to the satellite communication system model; Step S3: Model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Step S4: Solve the Markov decision process using an improved DQN algorithm; Step S5: Determine the optimized rate allocation coefficient based on the solution of the Markov decision process.

2. The satellite communication method based on rate-splitting multiple access according to claim 1, wherein, Step S1 includes: Step S11: The satellite transmits a data-bearing light beam to a near-space node through a coherent laser link; Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident light beam, and converts the modulated optical signal into a radio frequency signal; Step S13: The ground user terminal uses single-layer RSMA technology to receive the high-altitude communication platform signal through a direct transmission link.

3. The satellite communication method based on rate-splitting multiple access according to claim 1, characterized in that Step S2 includes: The optimization problem is determined according to the following formula: ; wherein, represents the total achievable rate of the system, represents the received signal rate of the common information flow of the i -th ground user, represents the rate splitting coefficient, w k represents the beamforming matrix actively transmitted by the high-altitude communication platform, represents the received signal rate of the private information flow of the i -th ground user, c i represents the data rate of the i -th ground user for receiving common information, and represents that the common information can be successfully decoded by each ground user, in represents the achievable rate requirement of the -th ground user, represents the energy allocation requirement of the i -th ground user is greater than J i , represents the maximum transmit power budget at the transmitter is less than , represents the rate splitting coefficient limit.

4. The satellite communication method based on rate-splitting multiple access according to claim 1, characterized in that, Step S3 includes: Determine the state space based on the information set of all current ground user terminals, the action vector selected by all current ground user terminals, and the immediate reward in the local training phase of all current ground user terminals; Determine the action space based on the active transmission beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term under unconstrained conditions and a penalty term under conditions that satisfy the constraints.

5. The satellite communication method based on rate-splitting multiple access according to claim 3, wherein Step S4 includes: Step S41: Define the direct link from each high-altitude communication platform to a ground user terminal as an autonomous decision-making unit, so that each autonomous decision-making unit has the ability of local state observation, an independent action space, and a personalized reward function; Step S42: Each autonomous decision-making unit shares a set of deep neural network parameters, learns a collaborative policy through a global experience pool, and broadcasts the trained deep neural network parameters to each agent; Step S43: According to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establish a dual-mode parameter update mechanism, and use mini-batch gradient descent to update the deep neural network parameters; Step S44: Use an improved DQN algorithm to optimize the transmission beamforming matrix and rate splitting coefficient of the high-altitude communication platform; Step S45: Solve the Markov decision process according to the deep neural network parameters, the transmission beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

6. A satellite communication system based on rate-splitting multiple access, characterized in that, including: Model construction module: used to construct a satellite communication system model using a geostationary satellite, a high-altitude communication platform, and a ground user terminal; Optimization problem determination module: used to determine the optimization problem according to the satellite communication system model with the goal of maximizing the total achievable rate of the system and subject to the power budget of the transmitter and the minimum rate of energy allocation as constraints; Decision process establishment module: used to model the optimization problem as a Markov decision process through a state space, an action space, and a reward function; Process solution module: used to solve the Markov decision process using an improved DQN algorithm; Solution verification module: used to determine the optimized rate allocation coefficient based on the solution of the Markov decision process.

7. The satellite communication system based on rate-splitting multiple access according to claim 6, characterized in that, The model construction module is used to perform the following steps: Step S11: The satellite emits a beam carrying data to the near-space node through a coherent laser link. Step S12: The high-altitude communication platform uses an adaptive optical receiving array to perform wavefront correction and energy coupling on the incident beam, and converts the modulated optical signal into a radio frequency signal. Step S13: The ground user terminal receives the signal of the high-altitude communication platform through the direct transmission link using the single-layer RSMA technology.

8. The satellite communication system based on rate-splitting multiple access according to claim 6, characterized in that, In the model construction module, the optimization problem is determined according to the following formula: ; wherein, represents the total achievable rate of the system, represents the received signal rate of the common information flow of the i -th ground user, represents the rate splitting coefficient, w k represents the beamforming matrix actively transmitted by the high-altitude communication platform, represents the received signal rate of the private information flow of the i -th ground user, c i represents the data rate of the i -th ground user for receiving common information, and represents that the common information can be successfully decoded by each ground user, in represents the achievable rate requirement of the -th ground user, represents the energy allocation requirement of the i -th ground user is greater than J i , represents the maximum transmit power budget at the transmitter is less than , represents the rate splitting coefficient limit.

9. The satellite communication system based on rate-splitting multiple access according to claim 6, wherein, The decision process establishment module is used to perform the following steps: Determine the state space according to the information set of all current ground user terminals, the action vector selected by all current ground user terminals, and the immediate reward in the local training phase of all current ground user terminals; Determine the action space according to the active transmit beam vector, the power splitting coefficient, and the achievable rate of the common information flow; The reward function includes: an immediate reward term under unconstrained conditions and a penalty term under satisfied constraint conditions.

10. The satellite communication system based on rate-splitting multiple access according to claim 8, characterized in that, The process solving module is used to perform the following steps: Step S41: Define the direct link from each high-altitude communication platform to the ground user terminal as an autonomous decision unit, so that each autonomous decision unit has the ability of local state observation, an independent action space, and a personalized reward function; Step S42: Each autonomous decision unit shares a set of deep neural network parameters, learns a collaborative strategy through a global experience pool, and broadcasts the deep neural network parameters of the trained strategy to each agent; Step S43: According to the time-varying characteristics of the link between the ground user terminal and the high-altitude communication platform, establish a dual-mode parameter update mechanism, and use mini-batch gradient descent to update the deep neural network parameters; Step S44: Use the improved DQN algorithm to optimize the transmit beamforming matrix and rate splitting coefficient of the high-altitude communication platform; Step S45: Solve the Markov decision process according to the deep neural network parameters, the transmit beamforming matrix of the high-altitude communication platform, and the rate splitting coefficient.

Citation Information

Patent Citations

  • Dynamic node scheduling method based on DQN algorithm in wireless body area network

    CN111465031A

  • Intelligent resource allocation method and device in low-orbit satellite communication

    CN115913317A

  • Satellite-ground convergence network multi-task unloading method and device, medium and equipment

    CN119483722A

  • Resource allocation strategy generation method, device and equipment based on satellite edge calculation

    CN119946720A

  • Coordinated multiple access method for multi-cell ground-to-air data transmission

    US12219374B1