Satellite-ground joint nonlinear linear distortion correction method and device based on reinforcement learning
By establishing a joint distortion model and action space based on reinforcement learning, and training the correction model using deep reinforcement learning algorithms, the problems of power amplifier nonlinear distortion and channel linear distortion in satellite-to-ground communication links were solved, thereby improving signal quality and transmission efficiency, and enhancing the system's dynamic adaptability and resource utilization efficiency.
Patent Information
- Application Number
- CN202511203493.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing technologies cannot effectively combine power amplifier nonlinearity distortion and channel linearity distortion in satellite-to-ground communication links, resulting in limited transmission rates and data capacity of the communication links. Furthermore, fixed parameter compensation methods cannot track the time-varying characteristics of satellite channels in real time.
A reinforcement learning-based approach is adopted to establish a joint distortion model, state space, action space, and joint reward function. The correction model is trained using a deep reinforcement learning algorithm. Through the joint adjustment of the predistorter, clipper, and equalizer, the nonlinear distortion of the power amplifier and the linear distortion of the channel are corrected.
The signal quality, power amplifier efficiency, and bit error rate have been optimized, improving the transmission efficiency and reliability of the communication link, simplifying on-board processing, reducing system resource consumption, and enhancing dynamic adaptability and global optimization capabilities.
Smart Images

Figure CN120729398B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of satellite communication technology, and in particular to a method and apparatus for joint satellite-ground nonlinear linear distortion correction based on reinforcement learning. Background Technology
[0002] With the development of satellite communication technology, satellite-to-ground communication links face numerous challenges when transmitting massive amounts of data. Specifically, these links suffer from power amplifier nonlinearity distortion and channel linearity distortion. Power amplifier nonlinearity distortion is mainly caused by the increased peak-to-average power ratio (PAPR) of higher-order modulation signals, leading the power amplifier to operate in the nonlinear region and affecting signal quality. Channel linearity distortion, on the other hand, is caused by factors such as channel unevenness, out-of-band rejection, and group delay fluctuations. These distortion problems severely affect the performance of higher-order modulation schemes, limiting the transmission rate and data capacity of satellite-to-ground communication links.
[0003] Traditional methods typically handle nonlinear and linear distortions independently, resulting in suboptimal compensation effects and failing to achieve global optimization. Furthermore, methods using fixed parameters cannot track the time-varying characteristics of satellite channels in real time. Summary of the Invention
[0004] This invention provides a reinforcement learning-based method and apparatus for joint nonlinear linear distortion correction between satellite and ground communication links, which solves the technical problem that existing technologies cannot jointly correct power amplifier nonlinear distortion and channel linear distortion in satellite-ground communication links.
[0005] On the one hand, this invention provides a reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method, comprising:
[0006] A joint distortion model is established; wherein the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link;
[0007] Establish a state space; wherein the state space includes signal characteristics, channel environment, and power amplifier status in the satellite-to-ground communication link;
[0008] Establish an action space; wherein, the action space includes adjustments to the predistorter at the transmitter end, the clipping threshold at the transmitter end, and the equalizer at the receiver end in the satellite-to-ground communication link;
[0009] Establish a joint reward function; wherein the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link;
[0010] A correction model for correcting the joint distortion model is trained using a deep reinforcement learning algorithm based on the state space, the action space, and the joint reward function.
[0011] The joint distortion model is corrected using the correction model.
[0012] On the other hand, the present invention also provides a reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device, comprising:
[0013] A joint distortion module is used to establish a joint distortion model; wherein, the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link;
[0014] A state space module is used to establish a state space, wherein the state space includes signal characteristics, channel environment, and power amplifier status in the satellite-to-ground communication link.
[0015] An action space module is used to establish an action space; wherein, the action space includes adjustments to the predistorter at the transmitter end, the clipping threshold at the transmitter end, and the equalizer at the receiver end in the satellite-to-ground communication link;
[0016] The reward function module is used to establish a joint reward function; wherein, the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link;
[0017] The reinforcement learning module is used to train a correction model for correcting the joint distortion model by employing a deep reinforcement learning algorithm based on the state space, the action space, and the joint reward function.
[0018] A correction module is used to correct the joint distortion model using the correction model.
[0019] The present invention provides a satellite-to-ground joint nonlinear linear distortion correction method and apparatus based on reinforcement learning. By establishing a joint distortion model, state space, action space and joint reward function, and training the correction model using a deep reinforcement learning algorithm, the joint correction of power amplifier nonlinear distortion and channel linear distortion in satellite-to-ground communication links is achieved. This optimizes signal quality, power amplifier energy efficiency and bit error rate, improves the transmission efficiency and reliability of communication links, simplifies on-board processing, reduces system resource consumption, and enhances the system's dynamic adaptability and global optimization capabilities. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1This is a flowchart illustrating the satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning provided in this embodiment of the invention.
[0022] Figure 2 This is a schematic diagram of the architecture of the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method provided in an embodiment of the present invention;
[0023] Figure 3 This is a schematic diagram of the power amplifier model provided in an embodiment of the present invention;
[0024] Figure 4 This is a schematic diagram of simulated channel group delay fluctuations provided in an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of the in-band amplitude characteristics of the simulated received signal provided in an embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of the 128QAM and 256QAM distortion-free constellations provided in the embodiments of the present invention;
[0027] Figure 7 These are the constellation diagram and scatter plot of the simulated EsNo of 19.63dB under the 128QAM modulation method provided in the embodiments of the present invention;
[0028] Figure 8 The constellation diagram and scatter plot of the simulated EsNo of 34.63dB under the 128QAM modulation mode provided in the embodiment of the present invention are shown.
[0029] Figure 9 The constellation diagram and scatter plot of the simulated EsNo of 49.63dB under the 128QAM modulation method provided in the embodiment of the present invention are shown.
[0030] Figure 10 This is a constellation diagram of 128QAM modulation under different conditions provided in the embodiments of the present invention;
[0031] Figure 11 This is a constellation diagram of 256QAM modulation under different conditions provided in the embodiments of the present invention;
[0032] Figure 12 This is a schematic diagram of the structure of the satellite-ground joint nonlinear linear distortion correction device based on reinforcement learning provided in an embodiment of the present invention;
[0033] Figure 13 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0035] Figure 1 This is a flowchart illustrating the satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning provided in this embodiment of the invention.
[0036] See Figure 1 The reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method may include the following steps 101 to 106.
[0037] Step 101: Establish a joint distortion model; wherein, the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link.
[0038] Specifically, the joint distortion model may include the following formula (1):
[0039] (1);
[0040] in, The received signal after nonlinear distortion and linear distortion; It is an ideal modulated signal; Here, K represents the power amplifier's nonlinearity coefficient; K is the nonlinearity order. This is the channel impulse response; is noise; k is an index variable, belonging to 1 to K.
[0041] This step, by unifying the modeling of power amplifier nonlinear distortion and channel linear distortion, can comprehensively consider various distortion factors in the communication link, avoiding the limitations of handling nonlinear and linear distortion separately.
[0042] Step 102: Establish the state space; wherein, the state space includes the signal characteristics, channel environment and power amplifier status in the satellite-to-ground communication link.
[0043] Specifically, the signal characteristics include:
[0044] Modulation order, instantaneous peak-to-average power ratio of the output signal, demodulated signal, and demodulation bit error rate;
[0045] Channel environment includes: received signal-to-noise ratio, Doppler frequency shift, in-band amplitude fluctuation, and group delay fluctuation;
[0046] Amplifier status includes:
[0047] Output power back-off value, temperature.
[0048] The state space in this step encompasses signal characteristics, channel environment, and power amplifier status, providing a comprehensive understanding of the current state of the communication link and enabling more accurate decision-making. Parameters in the state space (such as received signal-to-noise ratio, Doppler shift, and power amplifier temperature) reflect the dynamic changes in the communication link in real time, allowing the correction model to adaptively adjust its strategy to cope with time-varying channel environments.
[0049] Step 103: Establish the action space; wherein, the action space includes the adjustment of the predistorter at the transmitter end, the adjustment of the clipping threshold at the transmitter end, and the adjustment of the equalizer at the receiver end in the satellite-to-ground communication link.
[0050] Specifically, the adjustment of the transmitter predistorter includes:
[0051] Adjustment of memory depth and memory coefficient of the transmitter predistorter;
[0052] Adjustments to the clipping threshold at the transmitting end include:
[0053] The clipping threshold at the transmitter is dynamically adjusted using a deep reinforcement learning algorithm.
[0054] The adjustment of the receiver equalizer includes:
[0055] Adjustment of tap coefficients and update step size of the receiver equalizer.
[0056] In this step, the transmitting end can refer to the signal transmission section, especially the predistorter associated with the power amplifier. The transmitting end can refer to the signal processing stage, particularly the signal modulation and preprocessing stages. Alternatively, the transmitting end and the transmitting end can be considered as a single transmission section. By jointly adjusting the parameters of the transmitting end, the transmitting end, and the receiving end, global optimization of the communication link can be achieved, avoiding local optimization and thus improving overall performance. Adjustment of the transmitting end predistorter can also include the predistortion order and predistortion coefficients.
[0057] Step 104: Establish a joint reward function; wherein, the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link.
[0058] Specifically, the joint reward function includes the following formula (2):
[0059] (2);
[0060] Where R is the joint reward function; α These are the weighting coefficients for signal reception quality; is a measure of signal reception quality; e is the base of the natural logarithm; EVM is the error vector magnitude; β This is the weighting factor for energy efficiency; As a measure of energy efficiency; This is the actual output power of the power amplifier; This is the maximum output power of the power amplifier; This is the weighting factor for the bit error rate; is a measure of bit error rate; BER is the bit error rate.
[0061] This step provides a clear optimization direction for reinforcement learning algorithms by comprehensively considering multiple objectives such as signal quality, power amplifier efficiency, and bit error rate, thereby achieving a dynamic balance among multiple objectives.
[0062] Step 105: Using a deep reinforcement learning algorithm, a correction model is trained based on the state space, action space, and joint reward function to correct the joint distortion model.
[0063] Specifically, the formula (3) for the deep reinforcement learning algorithm is as follows:
[0064] (3);
[0065] in, L ( θ ) is the objective function of the deep reinforcement learning algorithm PPO; θ These are the parameters of the policy network; Indicates the expected value for time step t; This represents the probability of choosing action a according to policy π in state s; This represents the probability that the old policy will choose action a in state s; This represents the dominance function at time step t; The clipping parameter is typically set to 0.2. `clip` represents the clipping function. `min` indicates taking the minimum value to form a pessimistic estimate, avoiding over-optimization due to excessively high policy probabilities.
[0066] This step leverages the self-learning and adaptive capabilities of deep reinforcement learning algorithms to achieve global optimization of the communication link and dynamically adjust the correction strategy to adapt to complex time-varying environments.
[0067] The aforementioned advantage function can be calculated using generalized advantage estimation (GAE), as shown in formulas (4) and (5) below:
[0068] (4);
[0069] Where T is the total duration of the round, that is, the total number of time steps in the complete decision sequence; This is the offset of the time step; This is the discount factor (usually 0.99). This is the GAE smoothing factor (usually taken as 0.95).
[0070] (5);
[0071] in, Temporal Difference Error (TDEF) represents the difference between the state value at time step t (after adding the discount) and the current state value. The immediate reward obtained at time step t; The current state Value function estimation; The next state Value function estimation; Indicates at time step The reward at that time plus the difference between the discounted state value and the current state value.
[0072] If the advantage function A value greater than 0 indicates that the current action is better than average, encouraging an increase in its probability; if the advantage function... A value less than 0 indicates that the current action is worse than average, and its probability is suppressed.
[0073] Step 106: Use the calibration model to calibrate the joint distortion model.
[0074] This step significantly improves the signal quality and transmission efficiency of the communication link by real-time correction of power amplifier nonlinear distortion and channel linear distortion, while simplifying on-board processing and reducing resource consumption.
[0075] In this embodiment, by establishing a joint distortion model, state space, action space, and joint reward function, and using a deep reinforcement learning algorithm to train the correction model, joint correction of power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link is achieved. This optimizes signal quality, power amplifier energy efficiency, and bit error rate, improves the transmission efficiency and reliability of the communication link, simplifies on-board processing, reduces system resource consumption, and enhances the system's dynamic adaptability and global optimization capabilities.
[0076] In one embodiment of this specification, the calibration model includes:
[0077] The predistorter, using a memory polynomial model, has its output expressed by the following formula (6):
[0078] (6);
[0079] in, This is the output signal of the predistorter; K is the nonlinear order; k is the index variable, belonging to 1 to K;M For memory depth; The coefficients of the predistorter; These are past samples of the input signal. n Indicates the current time step. m Indicates the delay relative to the current time step;
[0080] A clipper, used to limit the peak amplitude of a signal, has its output expressed by the following formula (7):
[0081] (7);
[0082] The signal is after clipping. The original input signal represents the signal value at time t; The clipping threshold; The phase angle of the original input signal; For signal The absolute value of , where j represents the imaginary unit.
[0083] In this embodiment, the predistorter employs a memory polynomial model, enabling it to dynamically adjust the output signal based on past samples of the input signal, effectively compensating for the nonlinear distortion of the power amplifier. The clipper limits the peak amplitude of the signal, reducing the peak-to-average power ratio (PAPR), thereby decreasing the probability of the power amplifier entering the nonlinear region and further improving signal quality. This combined design, through the synergistic effect of predistortion and clipping, effectively improves the nonlinear distortion problem in satellite-to-ground communication links, enhancing the performance and reliability of the communication system.
[0084] In one embodiment of this specification, the calibration model further includes:
[0085] An equalizer is used to compensate for linear distortion and residual nonlinear distortion at the receiver.
[0086] In this embodiment, by introducing an equalizer at the receiving end, the system can further optimize the signal during the receiving stage, making up for the distortion that the predistorter and clipper at the transmitting end may not have fully corrected. This allows the entire correction model to not only handle nonlinear distortion but also effectively cope with linear distortion problems, further improving the overall performance of the communication link and enhancing the robustness and adaptability of the system.
[0087] In one embodiment of this specification, the equalizer is represented by the following formula (8):
[0088] (8);
[0089] in, In time step n The received signal output; For the firstk Volterra core, Indicates the input signal at time step n Previous The sample values at each time step; ∏ represents the multiplication symbol.
[0090] In this embodiment, the equalizer design based on the Volterra core makes the distortion compensation at the receiver more accurate and efficient, effectively handling the distortion problem caused by high-order modulated signals in complex channel environments, and significantly improving the signal quality of the system at the receiver.
[0091] Figure 2 This is a schematic diagram of the architecture of the reinforcement learning-based joint satellite-ground nonlinear linear distortion correction method provided in this embodiment of the invention. In this architecture, the system communicates with two different ground stations (e.g., ground station #1 and ground station #2) at two different time points (e.g., time T1 and time T2). Telemetry and control station #1 is responsible for establishing a telemetry and remote control link with the satellite for communication control and monitoring. Microwave data transmission station #1 is responsible for establishing a downlink high-speed data transmission link with the satellite for data reception and preliminary processing. Telemetry and control station #2 is similar to telemetry and control station #1, responsible for establishing a telemetry and remote control link with the satellite for communication control and monitoring. It sends control commands to the satellite and receives satellite status information and telemetry data. Microwave data transmission station #2 is similar to microwave data transmission station #1, responsible for establishing a downlink high-speed data transmission link with the satellite for data reception and preliminary processing. It receives high-speed data sent by the satellite and transmits it to ground station #2 for further demodulation and equalization processing.
[0092] The data center not only stores different satellite models but also handles the loading and aggregation of shared models from ground stations and satellite orbit insertion models. Using a federated learning framework, each ground station trains its model locally, uploads its gradients to the data center for aggregation, generates a global model, and then broadcasts the aggregated model to all ground station nodes. A downlink high-speed data transmission link is used to transmit data that has undergone pre-distortion and equalization processing, while a telemetry and remote control link is used to transmit power amplifier parameters and receive control commands from ground stations.
[0093] The workflow in the above architecture can be referenced as follows.
[0094] Onboard predistortion and parameter acquisition:
[0095] At time T1, the satellite communicates with ground station #1. The satellite completes data predistortion processing, and the predistortion parameters are remotely uploaded by the ground station. At the same time, the satellite collects power amplifier parameters, including power amplifier temperature, design output power, and actual output power, and transmits them to ground station #1 via telemetry.
[0096] Ground equilibration and model training:
[0097] After receiving telemetry data from the satellite, ground station #1 completes downlink data demodulation, equalization, and other processing. The ground station uses historical channel data and real-time information (such as signal quality indicators EVM, BER, environmental state parameters, etc.) to train a correction model and output equalization coefficients (used for local demodulation equalization) and predistortion coefficients (uploaded to the satellite periodically, or when the predistortion coefficients differ significantly).
[0098] Federated Learning Network:
[0099] Ground stations #1 and #2 share local experience data through the data center to improve the model's generalization ability. The data center is responsible for storing different satellite models, sharing models between ground stations, and loading and aggregating satellite orbit insertion models.
[0100] Communication at time T2:
[0101] At time T2, the satellite communicates with ground station #2. Ground station #2 completes equalization and model training, and uploads the predistortion coefficients to the satellite.
[0102] In some other embodiments of this specification, after correcting the joint distortion model using a correction model, the method may further include:
[0103] Updating federated learning parameters mainly includes the following steps:
[0104] Step 1: Ground Station Upload local model parameters at time t To the data center.
[0105] Step 2: The data center generates a global model by weighted aggregation of data from various ground station nodes, as shown in the following formula (9):
[0106] (9);
[0107] in, For ground station The local model parameters at time t, For the global model at time t+1, As weight, The calculation is shown in the following formula (10):
[0108] (10);
[0109] in, For the first The amount of empirical data per ground station node; This represents the total number of ground station nodes. Let be the amount of empirical data for the j-th ground station node.
[0110] Step 3: The data center broadcasts the aggregated global model to all ground station nodes. The broadcast method can be either a full broadcast or a differential update. Differential updates can reduce bandwidth usage because they only transmit parameter changes, thus reducing bandwidth consumption.
[0111] Parameter change It can be shown by the following formula (11):
[0112] (11);
[0113] Ground station nodes can use the model stored in the ontology or merge it with the globally updated model, as shown in the following formula (12):
[0114] (12);
[0115] in, For trust weights, take This means prioritizing the adoption of global knowledge.
[0116] If there is no local model, then .
[0117] Step 4: Calculate the deviation of the contribution of ground station nodes, as shown in the following formula (13):
[0118] (13);
[0119] in, Let be the deviation of the i-th ground station node. The norm of the global model parameters. This represents the distance between the model parameters of ground station i and the global model parameters.
[0120] If the deviation is greater than the preset threshold ( If the threshold is set (as a preset threshold), it is determined to be an abnormal node, and the data center will force an update of the ground station node model.
[0121] Step 5: Dynamically adjust the weights of ground station nodes based on channel quality;
[0122] Specifically, the weight of nodes with high SNR is increased to enhance the influence of reliable data. The adjusted weight can be represented by the following formula (14):
[0123] (14);
[0124] in, Let be the signal-to-noise ratio of the received signal at the i-th ground station node. The average signal-to-noise ratio of the received signal across all ground stations. The weight is the adjusted weight for the i-th ground station node.
[0125] In this embodiment, federated learning is a distributed machine learning method that allows multiple devices or servers to collaboratively train a model while maintaining data privacy and locality. The role of federated learning parameter updates is to achieve knowledge sharing and model fusion among multiple ground stations through the federated learning framework, thereby improving the generalization ability, adaptability, and performance of the entire satellite-ground joint nonlinear-linear distortion correction system.
[0126] The specific processing flow of the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method can be seen as follows:
[0127] Initialization phase:
[0128] Generally, simulation data can be generated using the Hammerstein-Wiener power amplifier model and the linear distortion channel model to initially train deep reinforcement learning (DRL).
[0129] Online calibration process:
[0130] Step 1: State Observation:
[0131] The satellite acquires signal feature matrices (including the instantaneous peak-to-average power ratio (PAPR), power amplifier output power backoff (OPBO), and power amplifier temperature (T)) and transmits them to the ground via telemetry; the ground acquires signal feature matrices in real time (including demodulation EVM, demodulation bit error rate, received signal-to-noise ratio (SNR), and Doppler shift). (Intra-band amplitude fluctuations, group delay fluctuations); combine satellite and ground signal characteristics to extract joint frequency / time domain feature vectors.
[0132] Step 2: Action Decision:
[0133] DRL takes a state vector as input and outputs the optimal action combination (predistortion parameters + equalization configuration parameters).
[0134] Step 3: Implementation of joint compensation:
[0135] The predistortion parameters are remotely fed to the satellite, where the transmitter applies digital predistortion (DPD) to correct the power amplifier's nonlinearity; the receiver initiates Volterra equalization to complete residual nonlinear equalization and linear equalization.
[0136] Step 4: Effectiveness Evaluation
[0137] Calculate the corrected EVM, BER, and power amplifier power consumption to generate an instant reward value.
[0138] Step 5: Experience Storage and Sharing:
[0139] Store the (state, action, reward, new state) tuple in the local experience pool and upload it to the data center as needed.
[0140] In some other embodiments, a deep reinforcement learning algorithm is used to train a correction model for correcting the joint distortion model based on the state space, action space, and joint reward function. This model may include:
[0141] Step 1: Network Structure Design
[0142] Using an Actor network, the output action distribution based on the state space, including predistortion parameters and equalizer parameters, can be expressed by formula (15):
[0143] (15);
[0144] in, Let be the mean of the actions output by the Actor network at time t; Let be the standard deviation of the actions output by the Actor network at time t;
[0145] The state value is estimated using a Critic network, which can be expressed by formula (16):
[0146] (16);
[0147] in, The state value estimated by the Critic network at time t. For Critic network.
[0148] Step Two, Training Process:
[0149] Collect experience points; these may include states, actions, rewards, and new states; states, actions, rewards, and new states can be represented as follows: ;
[0150] Calculate the dominance value at each time step using GAE ;
[0151] Perform multiple small-batch network updates for each batch of data, as shown in the following formula (17):
[0152] (17);
[0153] in, The parameters before the update. For the updated parameters, Indicates to Find the gradient, i.e. ; The learning rate can be dynamically adjusted according to channel conditions, such as the signal-to-noise ratio, as shown in formula (18):
[0154] (18);
[0155] The Critic network is updated to minimize the value error; where the value error is given by the following formula (19):
[0156] (19);
[0157] in, Let the target state value function be... This represents the average of the structures calculated at different times. This is the value error.
[0158] Step 3: Dynamic adjustment of weighting coefficients:
[0159] It can dynamically adjust according to real-time channel conditions, power amplifier operating parameters, and service requirements, for example:
[0160] Design a rule base to adjust weights based on current state observations. For example, increase weights when the signal-to-noise ratio is poor (low SNR). , , When the power amplifier temperature is too high, increase the weight. .
[0161] The following section provides a specific example for performance simulation and comparison. The simulation parameters are shown below:
[0162] Modulation methods: 128QAM, 256QAM;
[0163] Symbol rate: 750 Msps;
[0164] Residual frequency offset: 10 kHz;
[0165] Power amplifier nonlinear model: Saleh model;
[0166] Power amplifier linearity: 18dBm (output 24dBm), as shown in the figure below;
[0167] Amplifier P-1dB (input power at 1dB compression point): 23dBm (output 28dBm);
[0168] Simulated average input power: 23dBm (output 28dBm).
[0169] Figure 3 This is a schematic diagram of the power amplifier model provided in the embodiment of the present invention, showing the linear and nonlinear operating regions of the power amplifier and the 1dB compression point; Figure 4 This is a schematic diagram of simulated channel group delay fluctuation provided by an embodiment of the present invention, showing the simulation results of 5ns group delay fluctuation at a frequency of 450MHz; Figure 5 This is a schematic diagram of the amplitude characteristics of the simulated received signal within the band provided in an embodiment of the present invention, illustrating the power spectral density (PSD) characteristics of the received signal when the simulated channel unevenness is 1 dB.
[0170] Figure 6 This is a schematic diagram of a distortion-free constellation of 128QAM and 256QAM provided in an embodiment of the present invention, wherein, Figure 6 a is a diagram of a 128QAM distortion-free constellation. Figure 6 b represents a 256QAM distortion-free constellation. Figure 7 These are the constellation diagram and scatter plot of the simulated EsNo of 19.63dB under the 128QAM modulation method provided in this embodiment of the invention; wherein, Figure 7 'a' represents a constellation diagram. Figure 7 b is a scatter plot. Figure 8 These are the constellation diagram and scatter plot of the simulated EsNo of 34.63dB under the 128QAM modulation method provided in this embodiment of the invention; wherein, Figure 8 'a' represents a constellation diagram. Figure 8 b is a scatter plot. Figure 9 These are the constellation diagram and scatter plot of the simulated EsNo of 49.63dB under the 128QAM modulation method provided in this embodiment of the invention; wherein, Figure 9 'a' represents a constellation diagram. Figure 9 b is a scatter plot. The above figures show the signal constellation diagram and scatter plot under different signal-to-noise ratios (EsNo) and 128QAM modulation schemes, used to analyze and verify the performance of the satellite-ground joint nonlinear-linear distortion correction system proposed in this invention.
[0171] Figure 10 This is a constellation diagram of 128QAM modulation under different conditions provided in an embodiment of the present invention; wherein, Figure 10 a is the receiver constellation diagram when the in-band amplitude ripple is 1dB and EbNo=25dB; Figure 10 b represents the receiver constellation diagram when the in-band amplitude ripple is 1dB and EbNo=40dB; Figure 10 c represents the receiver constellation diagram with an in-band group delay ripple of 5ns and EbNo=40dB. Figure 11 This is a constellation diagram of 256QAM modulation under different conditions provided in the embodiments of the present invention; wherein, Figure 11 a represents the receiver constellation diagram when the in-band amplitude ripple is 1 dB and EbNo = 25 dB. Figure 11 b represents the receiver constellation diagram when the in-band amplitude ripple is 1 dB and EbNo = 40 dB. Figure 11c represents the receiver constellation diagram with an in-band group delay ripple of 5 ns and EbNo = 40 dB. The above figures illustrate the receiver constellation diagrams under different modulation schemes, different signal-to-noise ratios (EbNo), and different impairment conditions, used to analyze and verify the performance of the satellite-ground joint nonlinear-linear distortion correction system proposed in this invention.
[0172] The above Figures 7 to 11 In the diagram, red dots represent ideal constellation points, and blue dots represent corrected constellation points. Figures 7 to 11 In the diagram, the horizontal axis usually represents the in-phase component of the signal, also known as the real part; the vertical axis represents the quadrature component of the signal, also known as the imaginary part.
[0173] As can be seen from the above figures, the present invention optimizes signal quality, power amplifier efficiency, and bit error rate, thereby improving the transmission efficiency and reliability of the communication link.
[0174] Based on the same general inventive concept, this invention also protects a reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device, such as... Figure 12 As shown, Figure 12 This is a schematic diagram of the structure of the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device provided in an embodiment of the present invention. The reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device provided by the present invention will be described below. The reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device described below can be referred to in correspondence with the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method described above.
[0175] The reinforcement learning-based satellite-ground joint nonlinear linear distortion correction device includes a joint distortion module 1201, a state space module 1202, an action space module 1203, a reward function module 1204, a reinforcement learning module 1205, and a correction module 1206.
[0176] The joint distortion module 1201 is used to establish a joint distortion model; wherein, the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link;
[0177] The state space module 1202 is used to establish the state space; wherein, the state space includes the signal characteristics, channel environment and power amplifier status in the satellite-to-ground communication link;
[0178] The action space module 1203 is used to establish the action space; wherein, the action space includes the adjustment of the predistorter at the transmitter end, the adjustment of the clipping threshold at the transmitter end, and the adjustment of the equalizer at the receiver end in the satellite-to-ground communication link;
[0179] The reward function module 1204 is used to establish a joint reward function; wherein, the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link;
[0180] The reinforcement learning module 1205 is used to train a correction model for correcting the joint distortion model by using a deep reinforcement learning algorithm based on the state space, action space and joint reward function.
[0181] The correction module 1206 is used to correct the joint distortion model using the correction model.
[0182] Figure 13 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention.
[0183] like Figure 13 As shown, the electronic device may include a processor 1310, a communications interface 1320, a memory 1330, and a communication bus 1340. The processor 1310, communications interface 1320, and memory 1330 communicate with each other via the communication bus 1340. The processor 1310 can call logic instructions from the memory 1330 to execute a reinforcement learning-based joint satellite-ground nonlinear linear distortion correction method.
[0184] Furthermore, the logical instructions in the aforementioned memory 1330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0185] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method provided by the above methods.
[0186] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the reinforcement learning-based satellite-ground joint nonlinear linear distortion correction method provided by the above methods.
[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning, characterized in that, include: A joint distortion model is established; wherein the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link; Establish a state space; wherein the state space includes signal characteristics, channel environment, and power amplifier status in the satellite-to-ground communication link; Establish an action space; wherein, the action space includes adjustments to the predistorter at the transmitter end, the clipping threshold at the transmitter end, and the equalizer at the receiver end in the satellite-to-ground communication link; Establish a joint reward function; wherein the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link; A correction model for correcting the joint distortion model is trained using a deep reinforcement learning algorithm based on the state space, the action space, and the joint reward function. The joint distortion model is corrected using the correction model.
2. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The joint distortion model includes the following formula: ; in, The received signal after nonlinear distortion and linear distortion; It is an ideal modulated signal; Here, K represents the power amplifier's nonlinearity coefficient; K is the nonlinearity order. This is the channel impulse response; is noise; k is an index variable, belonging to 1 to K.
3. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The signal characteristics include: Modulation order, instantaneous peak-to-average power ratio of the output signal, demodulated signal, and demodulation bit error rate; The channel environment includes: received signal-to-noise ratio, Doppler frequency shift, in-band amplitude fluctuation, and group delay fluctuation. The power amplifier status includes: Output power back-off value, temperature.
4. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The adjustment of the transmitter predistorter includes: Adjustment of memory depth and memory coefficient of the transmitter predistorter; The adjustment of the clipping threshold at the transmitting end includes: The clipping threshold at the transmitter is dynamically adjusted using a deep reinforcement learning algorithm. The adjustment of the receiver equalizer includes: Adjustment of tap coefficients and update step size of the receiver equalizer.
5. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The joint reward function includes the following formula: ; Where R is the joint reward function; α These are the weighting coefficients for signal reception quality; is a measure of signal reception quality; e is the base of the natural logarithm; EVM is the error vector magnitude; β This is the weighting factor for energy efficiency; As a measure of energy efficiency; This is the actual output power of the power amplifier; This is the maximum output power of the power amplifier; This is the weighting factor for the bit error rate; is a measure of bit error rate; BER is the bit error rate.
6. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The formula for the deep reinforcement learning algorithm is shown below: ; in, L ( θ ) is the objective function of the deep reinforcement learning algorithm PPO; θ These are the parameters of the policy network; Indicates the expected value for time step t; This represents the probability of choosing action a according to policy π in state s; This represents the probability that the old policy will choose action a in state s; This represents the dominance function at time step t; For shearing parameters.
7. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 1, characterized in that, The correction model includes: The predistorter, employing a memory polynomial model, has its output expressed by the following formula: ; in, This is the output signal of the predistorter; K is the nonlinear order; k is the index variable, belonging to 1 to K; M For memory depth; These are the coefficients of the predistorter; These are past samples of the input signal. n Indicates the current time step. m This indicates a delay relative to the current time step; A clipper is used to limit the peak amplitude of a signal, and its output is expressed by the following formula: ; The signal is after clipping. The original input signal represents the signal value at time t; The clipping threshold; The phase angle of the original input signal; For signal The absolute value of.
8. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 7, characterized in that, The correction model also includes: An equalizer is used to compensate for linear distortion and residual nonlinear distortion at the receiver.
9. The satellite-ground joint nonlinear linear distortion correction method based on reinforcement learning according to claim 8, characterized in that, The equalizer is represented by the following formula: ; in, y ( n ) for time step n The received signal output; For the first k Volterra core, Indicates the input signal at time step n Previous The sample values at each time step; ∏ represents the multiplication symbol.
10. A satellite-ground joint nonlinear linear distortion correction device based on reinforcement learning, characterized in that, include: A joint distortion module is used to establish a joint distortion model; wherein, the joint distortion model includes power amplifier nonlinear distortion and channel linear distortion in the satellite-to-ground communication link; A state space module is used to establish a state space, wherein the state space includes signal characteristics, channel environment, and power amplifier status in the satellite-to-ground communication link. An action space module is used to establish an action space; wherein, the action space includes adjustments to the predistorter at the transmitter end, the clipping threshold at the transmitter end, and the equalizer at the receiver end in the satellite-to-ground communication link; The reward function module is used to establish a joint reward function; wherein, the joint reward function includes the error vector magnitude, power amplifier output power, and bit error rate in the satellite-to-ground communication link; The reinforcement learning module is used to train a correction model for correcting the joint distortion model by employing a deep reinforcement learning algorithm based on the state space, the action space, and the joint reward function. A correction module is used to correct the joint distortion model using the correction model.
Citation Information
Patent Citations
Power amplifier linearization thermal compensation method based on reinforcement learning
CN120067610A
OFDM (Orthogonal Frequency Division Multiplexing) signal time domain nonlinear distortion recovery method and device for satellite communication
CN120075001A