A decision-making method for anti-aging strategies based on deep reinforcement learning and meteorological data
By combining deep reinforcement learning with meteorological data, a multi-dimensional state space and intelligent decision engine are constructed, which solves the problems of timeliness of rain attenuation and decision isolation in satellite communication, realizes proactive prediction and multi-dimensional optimization, and improves link availability and system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies for dealing with rain attenuation in satellite communications suffer from problems such as insufficient timeliness, isolated decision-making, lack of predictive intelligence, and insufficient adaptability, making it difficult to achieve collaborative optimization decision-making based on multi-dimensional information.
By combining deep reinforcement learning with meteorological data, a comprehensive state space is constructed that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data. This space is then trained using deep reinforcement learning algorithms to form an intelligent decision engine capable of forward-looking and multi-dimensional joint optimization, thereby enabling a shift from passive response to proactive prediction.
It significantly improves link availability and overall system performance, achieves collaborative optimization of multi-dimensional anti-fading measures, can predict future fading events, avoids instantaneous link interruptions caused by response delays, and improves link reliability and resource utilization efficiency.
Smart Images

Figure CN121333397B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication and relates to a decision-making method for anti-fading strategy based on deep reinforcement learning and meteorological data. Background Technology
[0002] To meet the gigabits per second capacity requirements of communication satellite systems, modern internet satellites generally use high-frequency bands such as QV, Ka, and Ku as feed links. However, these high-frequency electromagnetic waves are highly susceptible to absorption and scattering by atmospheric factors, especially rainfall, clouds, and condensates, resulting in severe signal attenuation, known as "rain attenuation" (see [Alozie E, Abdulkarim A, Abdullahi I, et al. A review on rain signal attenuationmodeling, analysis and validation techniques: Advances, challenges and futuredirection[J]. Sustainability, 2022, 14(18): 11744.]). Rain attenuation can easily lead to signal loss exceeding 20 dB, far exceeding the range of conventional power control, and is the most significant factor restricting the reliability and availability of high-throughput satellite systems. Therefore, researching anti-rain attenuation strategies for low-Earth orbit satellites is of great importance.
[0003] Currently, decision-making methods for addressing rain attenuation challenges primarily rely on Uplink Power Control (UPC) and Adaptive Coding and Modulation (ACM) as the two core technologies, as seen in the reference [Jaiswal A, Jain VK, Kar S. Adaptive coding and modulation (ACM) technique for performance enhancement of FSO Link[C] / / 2016 IEEE First International Conference on Control, Measurement and Instrumentation (CMI). IEEE, 2016: 53-57.]. UPC infers and compensates for uplink attenuation by monitoring downlink beacon signal attenuation; ACM dynamically adjusts the modulation order and coding rate based on real-time channel conditions, switching to a more robust but less efficient modulation and coding scheme to maintain link connectivity when the channel deteriorates. However, these methods mainly rely on measurements of fading that has already occurred, resulting in control loop delays. Furthermore, the switching threshold for ACM is usually set based on long-term statistical experience, lacking adaptability to short-term meteorological dynamics.
[0004] To improve the predictive ability of rainfall attenuation events, research on meteorological data has been incorporated into the decision-making process. Shahrban M et al. proposed a rainfall attenuation prediction model based on numerical weather prediction (NWP), as shown in the reference [Shahrban M, Walker JP, Wang QJ, et al. An evaluation of numerical weather prediction based rainfall forecasts[J]. Hydrological Sciences Journal, 2016, 61(15):2704-2717.]. This model can predict the rainfall area and intensity in advance, providing early warning for link scheduling. However, this method relies on the accuracy of NWP, which itself suffers from computational complexity and long update cycles, making it difficult to meet the needs of rapid dynamic regulation at the minute or even second level. Another approach is station diversity, which uses multiple geographically dispersed gateway stations to avoid localized rainfall. Emiliani LD investigated the gains achievable through a multi-site diversity scheme and found that clouds still have a significant impact (see [Emiliani LD, Luini L. A Combined Rain and Cloud Attenuation FieldSimulator and its Application to Gateway Diversity Analysis at Ka, Q and Vbands[C] / / 34th AIAA International Communications Satellite SystemsConference. 2016: 5742.]). This decision-making approach is singular in dimension, lacks joint optimization with power, coding and modulation, and fails to fully utilize real-time radar extrapolation and other short-term forecast information.
[0005] In summary, existing technologies have the following shortcomings: First, they lack timeliness. Traditional UPC and ACM only take action after fading occurs, and the control loop delay leads to poor performance in dealing with rapidly changing deep fading. Second, decision-making is isolated. Anti-fading measures such as power control, ACM switching, and site diversity are usually managed by independent controllers, lacking collaborative decision-making and making it difficult to achieve global optimization. Third, they lack predictive intelligence, failing to effectively integrate and utilize high-precision, short-term weather forecast big data to predict future fading events. Fourth, they lack adaptability: thresholds or strategies based on fixed rules are difficult to adapt to changes in climate characteristics across different regions and seasons. Therefore, there is an urgent need in this field for an anti-fading method that can integrate multi-source information, has predictive capabilities, and can perform joint optimization decision-making. Summary of the Invention
[0006] To address the problems in the background technology, this invention proposes a decision-making method for anti-fading strategies based on deep reinforcement learning and meteorological data. This method constructs a comprehensive state space that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data, and trains it using advanced deep reinforcement learning algorithms. Ultimately, this forms an intelligent decision engine capable of forward-looking, multi-dimensional joint optimization, achieving a shift from passive response to proactive prediction.
[0007] The technical solution adopted in this invention is as follows:
[0008] A decision-making method for anti-attenuation strategies based on deep reinforcement learning and meteorological data, the specific steps of which are as follows:
[0009] Step 1: Build the STK+ns-3 co-simulation platform and construct the state space of satellite multi-source heterogeneous data fusion, including link state, network service state, meteorological state and system resource state; among them, the link state vector includes the real-time quality and historical change trend of the communication link, the network service state vector includes data load and service quality requirements, the meteorological state vector includes multi-source meteorological data, and the system resource state vector includes the status of available resources.
[0010] Step 2: Construct a DRL model based on near-end policy optimization, including a policy network and a value network. The DRL model outputs the corresponding action space based on the current state space and real-time weather conditions. The action space includes power control actions, adaptive coding and modulation selection actions, and site diversity handover actions. The power control actions define the strategy for adjusting the transmit power. The adaptive coding and modulation actions include the modulation scheme, coding rate, spectral efficiency, and minimum required SINR. The site handover actions control the rerouting decision of the service.
[0011] Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality.
[0012] Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
[0013] Furthermore, the state space in step 1 is:
[0014]
[0015] in, The specific definitions and modeling methods for each component at the decision moment are as follows:
[0016] (1) Link state vector for:
[0017]
[0018] In the formula, Indicates time The received beacon power measurement value, of which , Set value; Indicates time The signal-to-interference-plus-noise ratio (SIR) is calculated using the following formula: , Let be the received beacon power measurement value at time t. For interference power, The noise power spectral density; For a moment The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio.
[0019] (2) Network service state vector for:
[0020]
[0021] In the formula, Indicates the length of the send buffer queue; This represents the average latency of high-priority service packets; This indicates the proportion of high-priority services in the total service volume;
[0022] (3) Meteorological state vector for:
[0023]
[0024] In the formula, This indicates the future obtained by extrapolating from real-time weather radar data. The predicted path attenuation value at time t is obtained using an optical flow extrapolation model:
[0025]
[0026] in, Radar reflectivity factor For the echo moving vector field, The signal propagation path length; extrapolated time step. For set values, 1≤ ≤ ; This represents the average rainfall rate over a future set time period provided by numerical weather forecasts; This indicates the station's climate statistics attenuation value;
[0027] (4) System resource state vector for:
[0028]
[0029] In the formula, This indicates the current load rate of the backup gateway station; This indicates the current available power margin of the main gateway station.
[0030] Furthermore, the action space in step 2 is:
[0031]
[0032] In the formula, Indicates at time The action combination performed by the intelligent agent specifically includes:
[0033] (1) Power control action for:
[0034]
[0035] Action meaning: at the nominal transmission power Power adjustment based on Actual execution power: ;
[0036] (2) Adaptive coding and modulation selection action for:
[0037]
[0038] Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR;
[0039] (3) Site diversity switching action for:
[0040]
[0041] in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway station; This indicates a complete switch to the backup gateway station;
[0042] Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows,
[0043]
[0044] Each action combination represents a complete anti-aging strategy;
[0045] The action constraints are as follows:
[0046] power constraint: ;in, and Determined by the characteristics of the power amplifier hardware;
[0047] Switching delay constraint: ;
[0048] Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
[0049] Furthermore, the reward function expression in step 3 is:
[0050]
[0051] In the formula, These are the weighting coefficients for each reward item;
[0052] The specific items are as follows:
[0053] (1) Link availability reward To ensure basic connectivity of the communication link, a piecewise function is designed:
[0054]
[0055] The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link outage; Reward magnitude: positive rewards encourage maintaining high-quality connections, while negative penalties prevent link degradation and outages.
[0056] (2) System throughput bonus Positively correlated with spectral efficiency:
[0057]
[0058] in: This represents the spectral efficiency of the current modulation and coding scheme (MCS). The packet error rate is represented by BER (Bag Error Rate). , This refers to the length of the package. These are the normalization coefficients;
[0059] (3) Power efficiency penalty :
[0060]
[0061] in, This is the power adjustment amount; The theoretically optimal power adjustment is calculated based on the channel capacity formula. This is the maximum allowable power adjustment amount; The penalty factor is used to balance power efficiency and system performance.
[0062] (4) Switch the stability penalty item :
[0063]
[0064] in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load; Based on the switching penalty coefficient;
[0065] (5) Service quality awards :
[0066]
[0067] The time-delay reward function is:
[0068]
[0069] The throughput guarantee reward function is:
[0070]
[0071] in, This represents the average latency of high-priority service packets; Indicates the target latency for VoIP services; This represents the actual throughput; Indicates the minimum throughput required by the business; and Indicates the weighting coefficient;
[0072] Normalization of the reward function:
[0073]
[0074] in, and The moving average and standard deviation for rewards are dynamically updated during training.
[0075] Compared with the prior art, the present invention has the following advantages:
[0076] (1) The anti-fading strategy decision method based on deep reinforcement learning and meteorological data in this invention realizes the transformation from passive response to active prediction. By fusing meteorological radar extrapolation and NWP data, the agent can predict future fading events, thereby achieving a smooth transition, effectively avoiding instantaneous link interruption caused by response delay, and significantly improving the availability of the link.
[0077] (2) The present invention realizes the collaborative optimization of multi-dimensional anti-attenuation measures based on deep reinforcement learning and meteorological data anti-attenuation strategy decision-making method. Through a unified DRL agent, it simultaneously outputs power, MCS and site selection instructions. During the training process, the agent automatically learns the inherent correlation and trade-off between these actions to achieve the optimal overall system performance (throughput, energy consumption, and service continuity). Attached Figure Description
[0078] Figure 1 This is a flowchart of the overall process of the anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data in this invention.
[0079] Figure 2 This is a schematic diagram of the state space structure for multi-source heterogeneous data fusion according to the present invention.
[0080] Figure 3 This is a schematic diagram of the calculation mechanism of the multi-objective trade-off reward function of the present invention.
[0081] Figure 4 This is a diagram of the algorithm architecture for fusing deep reinforcement learning with meteorological data in this invention. Detailed Implementation
[0082] The present invention will now be described in further detail with reference to the accompanying drawings and examples;
[0083] This invention presents a decision-making method for anti-fading strategies based on deep reinforcement learning and meteorological data. By constructing a comprehensive state space that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data, and training it using advanced deep reinforcement learning algorithms, an intelligent decision engine capable of forward-looking and multi-dimensional joint optimization is finally formed, realizing the transformation from passive response to proactive prediction.
[0084] like Figure 1 As shown, the present invention includes the following steps:
[0085] Step 1: Build the STK+ns-3 co-simulation platform and construct the state space for satellite multi-source heterogeneous data fusion. :
[0086] This step aims to construct a comprehensive and accurate representation of the environment's state, providing a basis for decision-making in deep reinforcement learning agents, such as... Figure 2 As shown. The state space is designed as a comprehensive vector containing four main categories of information, and its mathematical representation is:
[0087]
[0088] in, The specific definitions and modeling methods for each component at the decision moment are as follows:
[0089] (1) Link state vector :
[0090] Link state vectors capture the real-time quality and historical trends of communication links, and are defined as follows:
[0091]
[0092] in, Indicates time The received beacon power measurement (unit: dBm), of which This sequence reflects short-term historical changes in signal strength. The value is usually 10, corresponding to the historical window of the past 10 seconds. Indicates time The signal-to-interference-plus-noise ratio (SIR / NOT) (unit: dB) is calculated using the following formula: ,in For interference power, This represents the noise power spectral density. Indicates time The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio.
[0093] (2) Network service state vector :
[0094] Network service state vectors represent data load and quality of service requirements, and are defined as follows:
[0095]
[0096] in, This indicates the length of the send buffer queue (in bytes), reflecting the backlog of data to be transmitted. This represents the average latency (in milliseconds) of high-priority service packets (such as VoIP). This indicates the proportion of high-priority services in the total service volume.
[0097] (3) Meteorological state vector :
[0098] Meteorological state vectors are key to achieving forecasting capabilities, integrating multi-source meteorological data:
[0099]
[0100] in, This indicates the future obtained by extrapolating from real-time weather radar data. Predicted path attenuation at time (in dB). Extrapolation model using optical flow:
[0101]
[0102] in, The radar reflectivity component at various velocities is calculated using the classic Horn-Schunck algorithm. Radar reflectivity factor For the echo moving vector field, This represents the signal propagation path length. The extrapolated time step is typically set to... minutes, that is . : Average rainfall rate for the next hour provided by numerical weather prediction (unit: mm / h). This represents the climatological attenuation value of a site calculated based on the model in ITU-RP.618 Recommendation, which is the attenuation value for an average of more than 1% of the time per year at that location, serving as a long-term reference benchmark.
[0103] (4) System resource state vector :
[0104] The system resource state vector describes the status of available resources:
[0105]
[0106] in, This represents the current load factor of the standby gateway (a decimal between 0 and 1). This indicates the current available power margin of the main gateway station (unit: dB).
[0107] Step 2: Construct a Deep Reinforcement Learning (DRL) model based on proximal policy optimization, including a policy network and a value network; the DRL model outputs the corresponding action space based on the current state space and real-time weather conditions.
[0108] This invention employs the Proximal Policy Optimization (PPO) algorithm as a training framework for deep reinforcement learning due to its stability and efficiency in continuous control problems. PPO avoids drastic fluctuations during training by limiting the step size of policy updates.
[0109] Policy Network (Actor) Architecture: Input Layer: 24-dimensional state vector Hidden layer 1: Fully connected layer, 128 neurons, ReLU activation function; Hidden layer 2: LSTM layer, 128 hidden units, handling temporal dependencies; Hidden layer 3: Fully connected layer, 64 neurons, ReLU activation function; Output layer: 90-dimensional action probability distribution.
[0110] Value Network (Critic) Architecture: Input Layer: 24-dimensional state vector Hidden layer: shares the first three layers with the policy network; Output layer: 1-dimensional state-value function, linear activation.
[0111] This invention designs a multi-dimensional joint optimization action space, enabling an agent to collaboratively control multiple anti-degradation mechanisms. The action space is defined as follows:
[0112]
[0113] in, Indicates at time The action combination executed by the intelligent agent specifically includes control instructions in the following three dimensions:
[0114] (1) Power control action :
[0115] The power control action defines the strategy for adjusting the transmit power:
[0116] (Unit: dB)
[0117] Action meaning: at the nominal transmission power Power adjustments based on the following parameters: Maximum gain +9dB: This corresponds to a compensation capability of approximately 20dB attenuation (considering the system's inherent margin); Minimum gain -3dB: This appropriately reduces power consumption under good channel conditions; 3dB step: This meets the control accuracy requirements of a practical power amplifier. Actual power output: This section lists a discrete selection region pre-defined for the reinforcement learning action space, which is adjusted based on the reinforcement learning process using the reward function and learning results.
[0118] (2) Adaptive coding and modulation selection action :
[0119] MCS actions are discrete selection variables, chosen from a predefined modulation and coding scheme library:
[0120]
[0121] Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR; here is a discrete selection region pre-defined for the reinforcement learning action space, which is adjusted by the reinforcement learning itself through the reward function and learning results.
[0122] The specific MCS parameter configurations are shown in Table 1 below.
[0123] Table 1
[0124]
[0125] (3) Site diversity switching action :
[0126] Rerouting decisions for site switching action control services:
[0127]
[0128] in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway (gradual switching). This indicates a complete switch to the backup gateway station (emergency switch).
[0129] Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows,
[0130]
[0131] Each action combination represents a complete anti-degradation strategy, for example:
[0132] Conservative strategy: ;
[0133] Aggressive strategy: ;
[0134] Emergency Response Strategy: .
[0135] The action constraints are as follows:
[0136] power constraint: ;in, and Determined by the characteristics of the power amplifier hardware;
[0137] Switching delay constraint: ;
[0138] Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
[0139] Action execution cycle: Basic execution cycle is 1 second; Emergency mode execution cycle: 100 milliseconds (when prediction decay exceeds 15dB); Action effective time: millisecond.
[0140] Through this multi-dimensional action space, the intelligent agent can autonomously select the optimal combination of anti-fading strategies based on real-time channel conditions, service requirements, and weather forecasts, significantly improving the system's reliability and resource utilization efficiency.
[0141] Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality.
[0142] The design of the reward function is crucial to the success of deep reinforcement learning; it needs to accurately reflect the system's multi-objective optimization requirements. The reward function of this invention comprehensively considers five dimensions: link reliability, spectral efficiency, power efficiency, network stability, and quality of service. Figure 3 As shown, its mathematical expression is as follows:
[0143]
[0144] in, The optimal values for the weighting coefficients of each reward item were determined through expert experience and grid search: , , , , .
[0145] (1) Link availability reward :
[0146] This reward ensures basic connectivity of the communication link and is designed as a piecewise function:
[0147]
[0148] The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link interruption; reward magnitude: positive rewards encourage maintaining a high-quality connection, while negative penalties prevent link degradation and interruption.
[0149] (2) System throughput bonus :
[0150] This award encourages efficient data transmission and is positively correlated with spectrum efficiency:
[0151]
[0152] in: The spectral efficiency (bps / Hz) of the current modulation and coding scheme (MCS) ranges from [0.67 to 5.40]. Packet error rate, estimated using BER: , This refers to the length of the package. This is a normalization factor to ensure that this reward matches the magnitude of other items.
[0153] (3) Power efficiency penalty :
[0154] This penalty controls power consumption to prevent over-firing.
[0155]
[0156] in, For power control action; The theoretically optimal power adjustment is calculated based on the channel capacity formula. dB is the maximum allowable power adjustment; This is a penalty factor used to balance power efficiency and system performance.
[0157] (4) Switch the stability penalty item :
[0158] This penalty reduces unnecessary site switching:
[0159]
[0160] in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load; The penalty coefficient is switched based on the base.
[0161] (5) Service quality awards :
[0162] This reward ensures the performance of high-priority services:
[0163]
[0164] The time-delay reward function is:
[0165]
[0166] The throughput guarantee reward function is:
[0167]
[0168] in, This represents the average latency of high-priority service packets (such as VoIP). ms represents the target latency for VoIP services; This represents the actual throughput; Indicates the minimum throughput required by the business; and These represent the weighting coefficients, set to 3.0 and 2.0 respectively.
[0169] Normalization of the reward function:
[0170] To ensure training stability, the final reward is normalized:
[0171]
[0172] in, and The moving average and standard deviation of the reward are dynamically updated during training. Through this carefully designed reward function, the agent can learn the optimal strategy to balance multiple competing objectives under complex constraints, ensuring both the basic reliability of the link and the efficient utilization of system resources.
[0173] Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
[0174] The core idea of the PPO algorithm is to limit the magnitude of policy updates through a pruning mechanism. Its objective function is defined as:
[0175]
[0176] in, This represents the probability ratio between the old and new strategies; This is the estimated value of the dominance function; The pruning parameters are typically set to 0.1-0.3; the advantage function is calculated using Generalized Advantage Estimation (GAE). The network is trained using the Adam optimizer, which stands for Adaptive Moment Estimation. It combines the advantages of momentum and RMSProp, providing a personalized learning rate for each parameter by adaptively estimating the first moment (mean) and second moment (uncentered variance) of the gradient. It is the core engine driving the efficient and stable training of deep neural networks (Actor network and Critic network).
[0177] To ensure that the trained DRL agent has strong generalization ability, its core environmental parameters and range are shown in Table 2.
[0178] Table 2
[0179]
[0180] The deep reinforcement learning training process based on proximal policy optimization (PPO) unfolds in a systematic manner to ensure that the algorithm can effectively learn the optimal anti-depletion policy. The entire training process is as follows: Figure 4 As shown.
[0181] Simulation verification and comparative analysis:
[0182] Simulation verification was performed using the STK+ns-3 co-simulation platform, and the channel model parameters are as follows:
[0183] Rainfall attenuation model: ITU-RP.618-13 model is adopted.
[0184]
[0185] in , (Ka-band horizontal polarization); Indicates the rainfall rate; The effective path length is expressed in km. In addition, multipath fading is considered and the Loo multipath fading model is adopted, which is suitable for satellite mobile channels. The shadow fading model adopts a log-normal distribution with a standard deviation of 3dB.
[0186] The service model parameters are set as follows: FTP service: file size 100MB, arrival rate Poisson distribution (λ=0.1); VoIP service: packet size 40 bytes, arrival rate fixed 50ms interval; video stream: bitrate 2Mbps, packet size 1500 bytes.
[0187] In 100 independent test scenarios, the method of this invention was compared with two existing methods. The first comparison method was the traditional reactive adaptive coding modulation and power control (reactive ACM), which makes decisions based solely on the current instantaneous channel measurements and is the benchmark scheme used in most current satellite communication systems. The second comparison method was a predictive rule-based method, which incorporates meteorological forecast information but uses simple rules based on fixed thresholds for decision-making.
[0188] The final simulation results are compared in Table 3.
[0189] Table 3
[0190]
[0191] Simulation results show that in terms of link availability, this invention achieves an excellent level of 99.2%, significantly improving upon traditional methods (94.3%) and predictive methods (97.5%). Particularly in heavy rainfall scenarios, it controls link quality fluctuations within 3dB through coordinated control 5 minutes in advance, completely avoiding the multiple outages encountered in the comparative methods. Regarding system efficiency, the average spectral efficiency of this invention reaches 3.45bps / Hz, a 47.4% improvement over traditional methods; its power efficiency reaches 12.8kbps / dBm, 33.3% higher than predictive methods. Furthermore, its site handover success rate is as high as 98.5%, and unnecessary handovers are reduced by 62%, significantly lowering network overhead.
[0192] In summary, while traditional reactive ACMs are simple to implement, their passive response mechanisms cannot cope with deep fading. Predictive rule-based methods, while introducing foresight, rely on fixed rules and struggle to adapt to complex and ever-changing real-world environments. In contrast, the core advantage of this invention lies in its intelligent decision-making capabilities and multi-method collaborative optimization mechanism: it automatically discovers the optimal strategy through end-to-end learning, achieving the best trade-off among multiple competing objectives such as link reliability, spectral efficiency, and power consumption. Although the training phase requires significant resources, its inference latency is low, fully meeting real-time control requirements, and providing a solid technical foundation for building next-generation intelligent, efficient, and reliable satellite communication systems.
Claims
1. A method for anti-attenuation strategy decision-making based on deep reinforcement learning and meteorological data, characterized in that, The specific steps are as follows: Step 1: Build the STK+ns-3 co-simulation platform and construct the state space of satellite multi-source heterogeneous data fusion, including link state, network service state, meteorological state and system resource state; among them, the link state vector includes the real-time quality and historical change trend of the communication link, the network service state vector includes data load and service quality requirements, the meteorological state vector includes multi-source meteorological data, and the system resource state vector includes the status of available resources. Step 2: Construct a DRL model based on near-end policy optimization, including a policy network and a value network. The DRL model outputs the corresponding action space based on the current state space and real-time weather conditions. The action space includes power control actions, adaptive coding and modulation selection actions, and site diversity handover actions. The power control actions define the strategy for adjusting the transmit power. The adaptive coding and modulation actions include the modulation scheme, coding rate, spectral efficiency, and minimum required SINR. The site diversity handover actions control the rerouting decision of the service. Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality. Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
2. The anti-aging strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The state space in step 1 is: ; in, The specific definitions and modeling methods for each component at the decision moment are as follows: (1) Link state vector for: ; In the formula, Indicates time The received beacon power measurement value, of which , Set value; Indicates time The signal-to-interference-plus-noise ratio (SIR) is calculated using the following formula: , Let be the received beacon power measurement value at time t. For interference power, The noise power spectral density; For a moment The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio. (2) Network service state vector for: ; In the formula, Indicates the length of the send buffer queue; This represents the average latency of high-priority service packets; This indicates the proportion of high-priority services in the total service volume; (3) Meteorological state vector for: ; In the formula, This indicates the future obtained by extrapolating from real-time weather radar data. The predicted path attenuation value at time t is obtained using an optical flow extrapolation model: ; in, Radar reflectivity factor For the echo moving vector field, The signal propagation path length; extrapolated time step. For set values, 1≤ ≤ ; This represents the average rainfall rate over a future set time period provided by numerical weather forecasts; This indicates the station's climate statistics attenuation value; (4) System resource state vector for: ; In the formula, This indicates the current load rate of the backup gateway station; This indicates the current available power margin of the main gateway station.
3. The anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The motion space in step 2 is: ; In the formula, Indicates at time The action combination performed by the intelligent agent specifically includes: (1) Power control action for: ; Action meaning: At the nominal transmission power Power adjustment based on Actual execution power: ; (2) Adaptive coding and modulation selection action for: ; Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR; (3) Site diversity switching action for: ; in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway station; This indicates a complete switch to the backup gateway station; Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows, ; Each action combination represents a complete anti-aging strategy; The action constraints are as follows: power constraint: ;in, and Determined by the characteristics of the power amplifier hardware; Switching delay constraint: , =500ms; Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
4. The anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The reward function expression in step 3 is: ; In the formula, These are the weighting coefficients for each reward item; The specific items are as follows: (1) Link availability reward To ensure basic connectivity of the communication link, a piecewise function is designed: ; The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link outage; Reward magnitude: positive rewards encourage maintaining high-quality connections, while negative penalties prevent link degradation and outages. (2) System throughput bonus Positively correlated with spectral efficiency: ; in: This indicates the current spectral efficiency of the MCS; The packet error rate is represented by BER (Bag Error Rate). , This refers to the length of the package. These are the normalization coefficients; (3) Power efficiency penalty : ; in, This is the power adjustment amount; The theoretically optimal power adjustment is calculated based on the channel capacity formula. This is the maximum allowable power adjustment amount; The penalty factor is used to balance power efficiency and system performance. (4) Switch the stability penalty item : ; in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load; Based on the switching penalty coefficient; (5) Service quality awards : ; The time-delay reward function is: ; The throughput guarantee reward function is: ; in, This represents the average latency of high-priority service packets; Indicates the target latency for VoIP services; This represents the actual throughput; Indicates the minimum throughput required by the business; and Indicates the weighting coefficient; Normalization of the reward function: ; in, and The moving average and standard deviation for rewards are dynamically updated during training.
Citation Information
Patent Citations
Channel information self-adaption-oriented intelligent cooperative transmission method between low-orbit satellites
CN114050855A
Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning
CN119962403A