Anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data
By combining deep reinforcement learning with meteorological data, a comprehensive state space was constructed and trained, which solved the timeliness and adaptability problems of satellite communication systems under rain attenuation, realized proactive prediction and multi-dimensional optimization, and improved link availability and system efficiency.
Patent Information
- Application Number
- CN202511881392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-12-15
AI Technical Summary
Existing satellite communication systems suffer from insufficient timeliness, isolated decision-making, lack of predictive intelligence, and insufficient adaptability when facing rain attenuation, resulting in poor performance when the signal decays rapidly.
By combining deep reinforcement learning with meteorological data, a comprehensive state space is constructed that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data. This space is then trained using deep reinforcement learning algorithms to form an intelligent decision engine capable of forward-looking and multi-dimensional joint optimization, thereby enabling a shift from passive response to proactive prediction.
It significantly improves link availability and overall system performance. By predicting future attenuation events, it avoids instantaneous link interruptions caused by response delays and achieves collaborative optimization of multi-dimensional anti-attenuation measures.
Smart Images

Figure CN121333397A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of satellite communication and relates to an anti-fading strategy decision method based on deep reinforcement learning and meteorological data. BACKGROUND
[0002] To meet the capacity demand of gigabit per second level for communication satellite systems, modern Internet satellites generally use high-frequency bands such as QV, Ka and Ku as feeder links. However, these high-frequency electromagnetic waves are easily absorbed and scattered by atmospheric environment, especially rain, cloud, water condensate and other factors, resulting in serious signal attenuation, i.e. rain attenuation, as described in the literature [Alozie E, Abdulkarim A, Abdullahi I, et al. A review on rain signal attenuation modeling, analysis and validation techniques: Advances, challenges and future direction [J]. Sustainability, 2022, 14(18): 11744.]. Rain attenuation can easily cause more than 20 dB of signal loss, far exceeding the range of conventional power control, and is the main factor restricting the reliability and availability of high-throughput satellite systems, so it is of great significance to study the anti-rain attenuation strategy method of low-orbit satellites.
[0003] Currently, the decision methods for rain attenuation challenges are mainly based on two core technologies of uplink power control (UPC) and adaptive coding and modulation (ACM), as described in the literature [Jaiswal A, Jain V K, Kar S. Adaptive coding and modulation (ACM) technique for performance enhancement of FSO Link [C] / / 2016 IEEE First International Conference on Control, Measurement and Instrumentation (CMI). IEEE, 2016: 53-57.]. UPC estimates and compensates the attenuation of the uplink by monitoring the attenuation of the downlink beacon signal; ACM dynamically adjusts the modulation order and coding rate according to the real-time channel conditions, and switches to a more robust but less efficient modulation and coding scheme when the channel deteriorates to maintain link connectivity. However, these methods mainly rely on the measurement of the already occurred fading, and there is a control loop delay. In addition, the switching threshold of ACM is usually set based on long-term statistical experience, lacking adaptability to short-term meteorological dynamics.
[0004] To improve the prediction ability of rain fade events, research on meteorological data is introduced into the decision-making process. Shahrban M et al. proposed a rain fade prediction model based on numerical weather prediction (NWP), see [Shahrban M, Walker J P, Wang Q J, et al. An evaluation of numerical weather prediction based rainfall forecasts [J]. Hydrological Sciences Journal, 2016, 61 (15): 2704-2717.]. This model can predict the rainfall area and intensity in advance, providing early warning for link scheduling. However, this method relies on the accuracy of NWP, which itself has the problems of complex calculation and long update cycle, making it difficult to meet the needs of rapid dynamic regulation and control at the minute or even second level. Another approach is site diversity, which avoids local rainfall by using multiple gateway stations geographically dispersed. Emiliani LD investigated the gain achieved by the multi-site diversity scheme and found that clouds still have a significant impact, see [Emiliani L D, Luini L. A Combined Rain and Cloud Attenuation Field Simulator and its Application to Gateway Diversity Analysis at Ka, Q and V bands [C] / / 34th AIAA International Communications Satellite Systems Conference. 2016: 5742.]. This decision-making dimension is single and not optimized jointly with power, coding and modulation, and does not fully utilize real-time radar extrapolation and other short-term forecast information.
[0005] In summary, the existing technology has the following shortcomings: first, the timeliness is insufficient, traditional UPC and ACM take action only after the fade occurs, and the control loop delay leads to poor performance in dealing with rapid changes in deep fading. Second, the decision-making is isolated, and the power control, ACM switching, and site diversity are usually managed by independent controllers, lacking collaborative decision-making and making it difficult to achieve global optimization. Third, there is a lack of prediction intelligence, and high-precision, short-term meteorological forecast data has not been effectively integrated and utilized to predict future fade events. Fourth, the adaptability is insufficient: threshold or strategy based on fixed rules is difficult to adapt to changes in climate characteristics in different regions and different seasons. Therefore, there is an urgent need in the field for an anti-fade method that can integrate multi-source information, have prediction ability, and make joint optimization decisions. SUMMARY
[0006] To address the problems in the background technology, this invention proposes a decision-making method for anti-fading strategies based on deep reinforcement learning and meteorological data. This method constructs a comprehensive state space that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data, and trains it using advanced deep reinforcement learning algorithms. Ultimately, this forms an intelligent decision engine capable of forward-looking, multi-dimensional joint optimization, achieving a shift from passive response to proactive prediction.
[0007] The technical solution adopted in this invention is as follows:
[0008] A decision-making method for anti-attenuation strategies based on deep reinforcement learning and meteorological data, the specific steps of which are as follows:
[0009] Step 1: Build the STK+ns-3 co-simulation platform and construct the state space of satellite multi-source heterogeneous data fusion, including link state, network service state, meteorological state and system resource state; among them, the link state vector includes the real-time quality and historical change trend of the communication link, the network service state vector includes data load and service quality requirements, the meteorological state vector includes multi-source meteorological data, and the system resource state vector includes the status of available resources.
[0010] Step 2: Construct a DRL model based on near-end policy optimization, including a policy network and a value network. The DRL model outputs the corresponding action space based on the current state space and real-time weather conditions. The action space includes power control actions, adaptive coding and modulation selection actions, and site diversity handover actions. The power control actions define the strategy for adjusting the transmit power. The adaptive coding and modulation actions include the modulation scheme, coding rate, spectral efficiency, and minimum required SINR. The site handover actions control the rerouting decision of the service.
[0011] Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality.
[0012] Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
[0013] Furthermore, the state space in step 1 is:
[0014]
[0015] in, The specific definitions and modeling methods for each component at the decision moment are as follows:
[0016] (1) Link state vector for:
[0017]
[0018] In the formula, Indicates time The received beacon power measurement value, of which , Set value; Indicates time The signal-to-interference-plus-noise ratio (SIR) is calculated using the following formula: , Let be the received beacon power measurement value at time t. For interference power, The noise power spectral density; For a moment The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio.
[0019] (2) Network service state vector for:
[0020]
[0021] In the formula, Indicates the length of the send buffer queue; This represents the average latency of high-priority service packets; This indicates the proportion of high-priority services in the total service volume;
[0022] (3) Meteorological state vector for:
[0023]
[0024] In the formula, This indicates the future obtained by extrapolating from real-time weather radar data. The predicted path attenuation value at time t is obtained using an optical flow extrapolation model:
[0025]
[0026] in, Radar reflectivity factor For the echo moving vector field, The signal propagation path length; extrapolated time step. For set values, 1≤ ≤ ; This represents the average rainfall rate over a future set time period provided by numerical weather forecasts; This indicates the station's climate statistics attenuation value;
[0027] (4) System resource state vector for:
[0028]
[0029] In the formula, This indicates the current load rate of the backup gateway station; This indicates the current available power margin of the main signaling station.
[0030] Furthermore, the action space in step 2 is:
[0031]
[0032] In the formula, Indicates at time The action combination performed by the intelligent agent specifically includes:
[0033] (1) Power control action for:
[0034]
[0035] Action meaning: At the nominal transmission power Power adjustment based on Actual execution power: ;
[0036] (2) Adaptive coding and modulation selection action for:
[0037]
[0038] Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR;
[0039] (3) Site diversity switching action for:
[0040]
[0041] in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway station; This indicates a complete switch to the backup gateway station;
[0042] Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows,
[0043]
[0044] Each action combination represents a complete anti-aging strategy;
[0045] The action constraints are as follows:
[0046] power constraint: ;in, and Determined by the characteristics of the power amplifier hardware;
[0047] Switching delay constraint: ;
[0048] Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
[0049] Furthermore, the reward function expression in step 3 is:
[0050]
[0051] In the formula, These are the weighting coefficients for each reward item;
[0052] The specific items are as follows:
[0053] (1) Link availability reward To ensure basic connectivity of the communication link, a piecewise function is designed:
[0054]
[0055] The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link outage; Reward magnitude: positive rewards encourage maintaining high-quality connections, while negative penalties prevent link degradation and outages.
[0056] (2) System throughput bonus Positively correlated with spectral efficiency:
[0057]
[0058] in: This represents the spectral efficiency of the current modulation and coding scheme (MCS). The packet error rate is estimated using BER: , This refers to the length of the package. These are the normalization coefficients;
[0059] (3) Power efficiency penalty :
[0060]
[0061] in, This is the power adjustment amount; The theoretically optimal power adjustment is calculated based on the channel capacity formula. This is the maximum allowable power adjustment amount; The penalty factor is used to balance power efficiency and system performance.
[0062] (4) Switch the stability penalty item :
[0063]
[0064] in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load. Based on the switching penalty coefficient;
[0065] (5) Service quality awards :
[0066]
[0067] The time-delay reward function is:
[0068]
[0069] The throughput guarantee reward function is:
[0070]
[0071] in, This represents the average latency of high-priority service packets; Indicates the target latency for VoIP services; This represents the actual throughput; Indicates the minimum throughput required by the business; and Indicates the weighting coefficient;
[0072] Normalization of the reward function:
[0073]
[0074] in, and The moving average and standard deviation for rewards are dynamically updated during training.
[0075] Compared with the prior art, the present invention has the following advantages:
[0076] (1) The anti-fading strategy decision method based on deep reinforcement learning and meteorological data in this invention realizes the transformation from passive response to active prediction. By fusing meteorological radar extrapolation and NWP data, the agent can predict future fading events, thereby achieving a smooth transition, effectively avoiding instantaneous link interruption caused by response delay, and significantly improving the availability of the link.
[0077] (2) The present invention realizes the collaborative optimization of multi-dimensional anti-attenuation measures based on deep reinforcement learning and meteorological data anti-attenuation strategy decision-making method. Through a unified DRL agent, it simultaneously outputs power, MCS and site selection instructions. During the training process, the agent automatically learns the inherent correlation and trade-off between these actions to achieve the optimal overall system performance (throughput, energy consumption, and service continuity). Attached Figure Description
[0078] Figure 1 This is a flowchart of the overall process of the anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data in this invention.
[0079] Figure 2 This is a schematic diagram of the state space structure for multi-source heterogeneous data fusion according to the present invention.
[0080] Figure 3 This is a schematic diagram of the calculation mechanism of the multi-objective trade-off reward function of the present invention.
[0081] Figure 4 This is a diagram of the algorithm architecture for fusing deep reinforcement learning with meteorological data in this invention. Detailed Implementation
[0082] The present invention will now be described in further detail with reference to the accompanying drawings and examples;
[0083] This invention presents a decision-making method for anti-fading strategies based on deep reinforcement learning and meteorological data. By constructing a comprehensive state space that integrates real-time channel state, network load, and multi-dimensional meteorological forecast data, and training it using advanced deep reinforcement learning algorithms, an intelligent decision engine capable of forward-looking and multi-dimensional joint optimization is finally formed, realizing the transformation from passive response to proactive prediction.
[0084] like Figure 1 As shown, the present invention includes the following steps:
[0085] Step 1: Build the STK+ns-3 co-simulation platform and construct the state space for satellite multi-source heterogeneous data fusion. :
[0086] This step aims to construct a comprehensive and accurate representation of the environment's state, providing a basis for decision-making in deep reinforcement learning agents, such as... Figure 2 As shown. The state space is designed as a comprehensive vector containing four main categories of information, and its mathematical representation is:
[0087]
[0088] in, The specific definitions and modeling methods for each component at the decision moment are as follows:
[0089] (1) Link state vector :
[0090] Link state vectors capture the real-time quality and historical trends of communication links, and are defined as follows:
[0091]
[0092] in, Indicates time The received beacon power measurement (unit: dBm), of which This sequence reflects short-term historical changes in signal strength. The value is usually 10, corresponding to the historical window of the past 10 seconds. Indicates time The signal-to-interference-plus-noise ratio (SIR / NOT) (unit: dB) is calculated using the following formula: ,in For interference power, This represents the noise power spectral density. Indicates time The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio.
[0093] (2) Network service state vector :
[0094] Network service state vectors represent data load and quality of service requirements, and are defined as follows:
[0095]
[0096] in, This indicates the length of the send buffer queue (in bytes), reflecting the backlog of data to be transmitted. This represents the average latency (in milliseconds) of high-priority service packets (such as VoIP). This indicates the proportion of high-priority services in the total service volume.
[0097] (3) Meteorological state vector :
[0098] Meteorological state vectors are key to achieving forecasting capabilities, integrating multi-source meteorological data:
[0099]
[0100] in, This indicates the future obtained by extrapolating from real-time weather radar data. Predicted path attenuation at time (in dB). Extrapolation model using optical flow:
[0101]
[0102] in, The radar reflectivity component at various velocities is calculated using the classic Horn-Schunck algorithm. Radar reflectivity factor For the echo moving vector field, This represents the signal propagation path length. The extrapolated time step is typically set to... minutes, that is . : Average rainfall rate for the next hour provided by numerical weather prediction (unit: mm / h). This represents the climatological attenuation value of a site calculated based on the model in ITU-RP.618 Recommendation, which is the attenuation value for an average of more than 1% of the time per year at that location, serving as a long-term reference benchmark.
[0103] (4) System resource state vector :
[0104] The system resource state vector describes the status of available resources:
[0105]
[0106] in, This represents the current load factor of the standby gateway (a decimal between 0 and 1). This indicates the current available power margin of the main gateway station (unit: dB).
[0107] Step 2: Construct a Deep Reinforcement Learning (DRL) model based on proximal policy optimization, including a policy network and a value network; the DRL model outputs the corresponding action space based on the current state space and real-time weather conditions.
[0108] This invention employs the Proximal Policy Optimization (PPO) algorithm as a training framework for deep reinforcement learning due to its stability and efficiency in continuous control problems. PPO avoids drastic fluctuations during training by limiting the step size of policy updates.
[0109] Policy Network (Actor) Architecture: Input Layer: 24-dimensional state vector Hidden layer 1: Fully connected layer, 128 neurons, ReLU activation function; Hidden layer 2: LSTM layer, 128 hidden units, handling temporal dependencies; Hidden layer 3: Fully connected layer, 64 neurons, ReLU activation function; Output layer: 90-dimensional action probability distribution.
[0110] Value Network (Critic) Architecture: Input Layer: 24-dimensional state vector Hidden layer: shares the first three layers with the policy network; Output layer: 1-dimensional state-value function, linear activation.
[0111] This invention designs a multi-dimensional joint optimization action space, enabling an agent to collaboratively control multiple anti-degradation mechanisms. The action space is defined as follows:
[0112]
[0113] in, Indicates at time The action combination executed by the intelligent agent specifically includes control instructions in the following three dimensions:
[0114] (1) Power control action :
[0115] The power control action defines the strategy for adjusting the transmit power:
[0116] (Unit: dB)
[0117] Action meaning: At the nominal transmission power Power adjustments based on the following parameters: Maximum gain +9dB: This corresponds to a compensation capability of approximately 20dB attenuation (considering the system's inherent margin); Minimum gain -3dB: This appropriately reduces power consumption under good channel conditions; 3dB step: This meets the control accuracy requirements of a practical power amplifier. Actual power output: This section lists a discrete selection region pre-defined for the reinforcement learning action space, which is adjusted based on the reinforcement learning process using the reward function and learning results.
[0118] (2) Adaptive coding and modulation selection action :
[0119] MCS actions are discrete selection variables, chosen from a predefined modulation and coding scheme library:
[0120]
[0121] Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR; here is a discrete selection region pre-defined for the reinforcement learning action space, which is adjusted by the reinforcement learning itself through the reward function and learning results.
[0122] The specific MCS parameter configurations are shown in Table 1 below.
[0123] Table 1
[0124] (3) Site diversity switching action :
[0125] Rerouting decisions for site switching action control services:
[0126]
[0127] in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway (gradual switching). This indicates a complete switch to the backup gateway station (emergency switch).
[0128] Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows,
[0129]
[0130] Each action combination represents a complete anti-degradation strategy, for example:
[0131] Conservative strategy: ;
[0132] Aggressive strategy: ;
[0133] Emergency Response Strategy: .
[0134] The action constraints are as follows:
[0135] power constraint: ;in, and Determined by the characteristics of the power amplifier hardware;
[0136] Switching delay constraint: ;
[0137] Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
[0138] Action execution cycle: Basic execution cycle is 1 second; Emergency mode execution cycle: 100 milliseconds (when prediction decay exceeds 15dB); Action effective time: millisecond.
[0139] Through this multi-dimensional action space, the intelligent agent can autonomously select the optimal combination of anti-fading strategies based on real-time channel conditions, service requirements, and weather forecasts, significantly improving the system's reliability and resource utilization efficiency.
[0140] Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality.
[0141] The design of the reward function is crucial to the success of deep reinforcement learning; it needs to accurately reflect the system's multi-objective optimization requirements. The reward function of this invention comprehensively considers five dimensions: link reliability, spectral efficiency, power efficiency, network stability, and quality of service. Figure 3 As shown, its mathematical expression is as follows:
[0142]
[0143] in, The optimal values for the weighting coefficients of each reward item were determined through expert experience and grid search: , , , , .
[0144] (1) Link availability reward :
[0145] This reward ensures basic connectivity of the communication link and is designed as a piecewise function:
[0146]
[0147] The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link interruption; reward magnitude: positive rewards encourage maintaining a high-quality connection, while negative penalties prevent link degradation and interruption.
[0148] (2) System throughput bonus :
[0149] This award encourages efficient data transmission and is positively correlated with spectrum efficiency:
[0150]
[0151] in: The spectral efficiency (bps / Hz) of the current modulation and coding scheme (MCS) ranges from [0.67 to 5.40]. Packet error rate, estimated using BER: , This refers to the length of the package. This is a normalization factor to ensure that this reward matches the magnitude of other items.
[0152] (3) Power efficiency penalty :
[0153] This penalty controls power consumption to prevent over-firing.
[0154]
[0155] in, For power control action; The theoretically optimal power adjustment is calculated based on the channel capacity formula. dB is the maximum allowable power adjustment; This is a penalty factor used to balance power efficiency and system performance.
[0156] (4) Switch the stability penalty item :
[0157] This penalty reduces unnecessary site switching:
[0158]
[0159] in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load. The penalty coefficient is switched based on the base.
[0160] (5) Service quality awards :
[0161] This reward ensures the performance of high-priority services:
[0162]
[0163] The time-delay reward function is:
[0164]
[0165] The throughput guarantee reward function is:
[0166]
[0167] in, This represents the average latency of high-priority service packets (such as VoIP). ms represents the target latency for VoIP services; This represents the actual throughput; Indicates the minimum throughput required by the business; and These represent the weighting coefficients, set to 3.0 and 2.0 respectively.
[0168] Normalization of the reward function:
[0169] To ensure training stability, the final reward is normalized:
[0170]
[0171] in, and The moving average and standard deviation of the reward are dynamically updated during training. Through this carefully designed reward function, the agent can learn the optimal strategy to balance multiple competing objectives under complex constraints, ensuring both the basic reliability of the link and the efficient utilization of system resources.
[0172] Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
[0173] The core idea of the PPO algorithm is to limit the magnitude of policy updates through a pruning mechanism. Its objective function is defined as:
[0174]
[0175] in, This represents the probability ratio between the old and new strategies; This is the estimated value of the dominance function; The pruning parameters are typically set to 0.1-0.3; the advantage function is calculated using Generalized Advantage Estimation (GAE). The network is trained using the Adam optimizer, which stands for Adaptive Moment Estimation. It combines the advantages of momentum and RMSProp, providing a personalized learning rate for each parameter by adaptively estimating the first moment (mean) and second moment (uncentered variance) of the gradient. It is the core engine driving the efficient and stable training of deep neural networks (Actor network and Critic network).
[0176] To ensure that the trained DRL agent has strong generalization ability, its core environmental parameters and range are shown in Table 2.
[0177] Table 2
[0178] The deep reinforcement learning training process based on proximal policy optimization (PPO) unfolds in a systematic manner to ensure that the algorithm can effectively learn the optimal anti-depletion policy. The entire training process is as follows: Figure 4 As shown.
[0179] Simulation verification and comparative analysis:
[0180] Simulation verification was performed using the STK+ns-3 co-simulation platform, and the channel model parameters are as follows:
[0181] Rainfall attenuation model: ITU-RP.618-13 model is adopted.
[0182]
[0183] in , (Ka-band horizontal polarization); Indicates the rainfall rate; The effective path length is expressed in km. In addition, multipath fading is considered and the Loo multipath fading model is adopted, which is suitable for satellite mobile channels. The shadow fading model adopts a log-normal distribution with a standard deviation of 3dB.
[0184] The service model parameters are set as follows: FTP service: file size 100MB, arrival rate Poisson distribution (λ=0.1); VoIP service: packet size 40 bytes, arrival rate fixed 50ms interval; video stream: bitrate 2Mbps, packet size 1500 bytes.
[0185] In 100 independent test scenarios, the method of this invention was compared with two existing methods. The first comparison method was the traditional reactive adaptive coding modulation and power control (reactive ACM), which makes decisions based solely on the current instantaneous channel measurements and is the benchmark scheme used in most current satellite communication systems. The second comparison method was a predictive rule-based method, which incorporates meteorological forecast information but uses simple rules based on fixed thresholds for decision-making.
[0186] The final simulation results are compared in Table 3.
[0187] Table 3
[0188] Simulation results show that in terms of link availability, this invention achieves an excellent level of 99.2%, significantly improving upon traditional methods (94.3%) and predictive methods (97.5%). Particularly in heavy rainfall scenarios, it controls link quality fluctuations within 3dB through coordinated control 5 minutes in advance, completely avoiding the multiple outages encountered in the comparative methods. Regarding system efficiency, the average spectral efficiency of this invention reaches 3.45bps / Hz, a 47.4% improvement over traditional methods; its power efficiency reaches 12.8kbps / dBm, 33.3% higher than predictive methods. Furthermore, its site handover success rate is as high as 98.5%, and unnecessary handovers are reduced by 62%, significantly lowering network overhead.
[0189] In summary, while traditional reactive ACMs are simple to implement, their passive response mechanisms cannot cope with deep fading. Predictive rule-based methods, while introducing foresight, rely on fixed rules and struggle to adapt to complex and ever-changing real-world environments. In contrast, the core advantage of this invention lies in its intelligent decision-making capabilities and multi-method collaborative optimization mechanism: it automatically discovers the optimal strategy through end-to-end learning, achieving the best trade-off among multiple competing objectives such as link reliability, spectral efficiency, and power consumption. Although the training phase requires significant resources, its inference latency is low, fully meeting real-time control requirements, and providing a solid technical foundation for building next-generation intelligent, efficient, and reliable satellite communication systems.
Claims
1. A method for anti-attenuation strategy decision-making based on deep reinforcement learning and meteorological data, characterized in that, The specific steps are as follows: Step 1: Build the STK+ns-3 co-simulation platform and construct the state space of satellite multi-source heterogeneous data fusion, including link state, network service state, meteorological state and system resource state; among them, the link state vector includes the real-time quality and historical change trend of the communication link, the network service state vector includes data load and service quality requirements, the meteorological state vector includes multi-source meteorological data, and the system resource state vector includes the status of available resources. Step 2: Construct a DRL model based on near-end policy optimization, including a policy network and a value network. The DRL model outputs the corresponding action space based on the current state space and real-time weather conditions. The action space includes power control actions, adaptive coding and modulation selection actions, and site diversity handover actions. The power control actions define the strategy for adjusting the transmit power. The adaptive coding and modulation actions include the modulation scheme, coding rate, spectral efficiency, and minimum required SINR. The site handover actions control the rerouting decision of the service. Step 3: The STK+ns-3 co-simulation platform performs simulation based on the corresponding action space. After simulation, the state space of the next moment is obtained, and the reward function is used to calculate the reward values of the five dimensions of fused link reliability, spectral efficiency, power efficiency, network stability and service quality. Step 4: The training engine updates the DRL model based on the current state space, action space, reward value, and the state space of the next time step, using the state space of the next time step as the state space of the current time step, and returns to step 2.
2. The anti-aging strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The state space in step 1 is: ; in, The specific definitions and modeling methods for each component at the decision moment are as follows: (1) Link state vector for: ; In the formula, Indicates time The received beacon power measurement value, of which , Set value; Indicates time The signal-to-interference-plus-noise ratio (SIR) is calculated using the following formula: , Let be the received beacon power measurement value at time t. For interference power, The noise power spectral density; For a moment The bit error rate is estimated through real-time demodulation statistics or based on a mapping table of signal-to-interference-plus-noise ratio. (2) Network service state vector for: ; In the formula, Indicates the length of the send buffer queue; This represents the average latency of high-priority service packets; This indicates the proportion of high-priority services in the total service volume; (3) Meteorological state vector for: ; In the formula, This indicates the future obtained by extrapolating from real-time weather radar data. The predicted path attenuation value at time t is obtained using an optical flow extrapolation model: ; in, Radar reflectivity factor For the echo moving vector field, The signal propagation path length; extrapolated time step. For set values, 1≤ ≤ ; This represents the average rainfall rate over a future set time period provided by numerical weather forecasts; This indicates the station's climate statistics attenuation value; (4) System resource state vector for: ; In the formula, This indicates the current load rate of the backup gateway station; This indicates the current available power margin of the main gateway station.
3. The anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The motion space in step 2 is: ; In the formula, Indicates at time The action combination performed by the intelligent agent specifically includes: (1) Power control action for: ; Action meaning: At the nominal transmission power Power adjustment based on Actual execution power: ; (2) Adaptive coding and modulation selection action for: ; Each adaptive coding and modulation action includes the modulation scheme, coding rate, spectral efficiency, and minimum required SINR; (3) Site diversity switching action for: ; in, This indicates that the current main gateway connection will be maintained; This indicates that 50% of the service traffic will be switched to the backup gateway station; This indicates a complete switch to the backup gateway station; Joint optimization characteristics of the action space: The total number of combinations in the action space is as follows, ; Each action combination represents a complete anti-aging strategy; The action constraints are as follows: power constraint: ;in, and Determined by the characteristics of the power amplifier hardware; Switching delay constraint: ; Service continuity constraint: The QoS requirements of high-priority services must be guaranteed during the handover process.
4. The anti-attenuation strategy decision-making method based on deep reinforcement learning and meteorological data according to claim 1, characterized in that, The reward function expression in step 3 is: ; In the formula, These are the weighting coefficients for each reward item; The specific items are as follows: (1) Link availability reward To ensure basic connectivity of the communication link, a piecewise function is designed: ; The threshold setting for the bit error rate (BER) is based on satellite communication industry standards. In response to the requirements of high-quality business, The threshold for link outage; Reward magnitude: positive rewards encourage maintaining high-quality connections, while negative penalties prevent link degradation and outages. (2) System throughput bonus Positively correlated with spectral efficiency: ; in: This indicates the current spectral efficiency of the MCS; The packet error rate is represented by BER (Bag Error Rate). , This refers to the length of the package. These are the normalization coefficients; (3) Power efficiency penalty : ; in, This is the power adjustment amount; The theoretically optimal power adjustment is calculated based on the channel capacity formula. This is the maximum allowable power adjustment amount; The penalty factor is used to balance power efficiency and system performance. (4) Switch the stability penalty item : ; in, This is an indicator function; it is 1 when a switch occurs, and 0 otherwise. For backup site load rates, the penalty increases with increasing load. Based on the switching penalty coefficient; (5) Service quality awards : ; The time-delay reward function is: ; The throughput guarantee reward function is: ; in, This represents the average latency of high-priority service packets; Indicates the target latency for VoIP services; This represents the actual throughput. Indicates the minimum throughput required by the business; and Indicates the weighting coefficient; Normalization of the reward function: ; in, and The moving average and standard deviation for rewards are dynamically updated during training.
Citation Information
Patent Citations
Channel information self-adaption-oriented intelligent cooperative transmission method between low-orbit satellites
CN114050855A
Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning
CN119962403A
Microwave communication data transmission method
CN120150762A
Unmanned aerial vehicle anti-interference communication link control method based on heterogeneous network convergence
CN120769281A
Deep reinforcement learning intelligent decision-making platform based on unified artificial intelligence framework
US20240338570A1