Self-adaptive coding modulation method and system for relieving sun-borne interference in ultra-dense low-orbit satellite network
Through deep reinforcement learning and channel prediction models, the coding and modulation scheme is dynamically adjusted to solve the communication instability caused by solar eclipse interference in ultra-dense low-orbit satellite networks, achieve real-time response and efficient resistance to solar eclipse interference, and improve system performance and resource utilization efficiency.
Patent Information
- Application Number
- CN202510798524.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have failed to effectively deal with solar eclipse interference in ultra-dense low-orbit satellite networks, resulting in frequent switching of communication links and difficulty in ensuring communication continuity and stability. In addition, channel prediction and MCS selection methods cannot respond to rapidly changing channel conditions in a timely manner, and the computational complexity and strategy convergence performance are insufficient.
A deep reinforcement learning method is combined with a long short-term memory network and a proximal strategy optimization algorithm of knowledge distillation. Through the ground station's estimation and prediction of the satellite-to-ground channel state, the coding and modulation scheme is dynamically selected. Combined with the hard target and soft target training models, real-time response and efficient defense against solar eclipse interference are achieved.
It significantly improves the system's spectrum utilization, accurately predicts the channel status during solar eclipse interference, reduces the bit error rate, improves the system's anti-interference capability and resource utilization efficiency in complex space environments, and promotes the development of satellite communication networks towards high robustness and intelligence.
Smart Images

Figure CN120658350A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of satellite communication technology and relates to the design of an adaptive coding modulation scheme for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network. Background Art
[0002] Ultra-dense low-Earth Orbit (LEO) satellite communication networks (ULSNs), as a complement to terrestrial networks, offer wide coverage, flexible deployment, and low latency. They can effectively expand the communication range of terrestrial networks, such as those in oceanic and mountainous areas. Currently, with the rapid growth in the number of satellites and the increasing demand for user access, effectively improving the efficiency and throughput of inter-satellite and satellite-to-ground communications in ULSNs has become a key challenge that needs to be addressed. In this context, accurate modeling and performance evaluation of highly dynamic inter-satellite and satellite-to-ground links have become core tasks in ULSN research. Operating LEO satellites must navigate a complex space interference environment, particularly solar transits. High-frequency, broadband inter-satellite laser links and satellite-to-ground microwave links are extremely sensitive to solar transit interference, easily impacting them and leading to frequent link handoffs, which can impact system performance. When strong solar electromagnetic interference coincides with the direction of satellite downlink signals, the sun becomes a significant noise source for the link's receiving antenna. The additional solar noise power can cause demodulator performance to drop below the threshold, or even disrupt the communication link in severe cases.
[0003] Therefore, accurately calculating the interference impact of solar eclipse on the link, predicting its occurrence time, and proposing an effective adaptive coding modulation (ACM) strategy to alleviate solar eclipse interference are key links in optimizing the overall performance of satellite communication systems.
[0004] A review of existing literature reveals that solar transit interference has been extensively studied in the satellite-to-ground communication links of geosynchronous earth orbit (GEO) satellites and in inter-satellite communications within global navigation satellite systems. Because the sun is approximately spherical and can be observed as a disk of equal area on Earth's surface, the occurrence of a solar transit can be determined by analyzing the relative positions of the sun, satellites, and ground stations. Furthermore, the noise interference power caused by the transit can be estimated by calculating the solar surface temperature. This information allows for early preparation for potential severe interference or signal outages. Hughes Aircraft previously proposed a dual-receiver system that circumvented solar radiation interference by establishing cross-link communications between the affected ground station and two satellites during a solar transit. However, this system required up to 28 minutes of redundant transmission time, resulting in reduced communication efficiency. In a 2024 article titled “A local pre-rerouting algorithm to combat sun outage for inter-satellite links in low Earth orbit satellite networks,” published in Applied Sciences, J. Jin et al. proposed a local pre-routing algorithm in which a central node collects link information and service data on the link to be interrupted before the link is interrupted, and performs rerouting to reduce the probability of service interruption. Furthermore, in a 2010 article titled “Prediction and compensation of sun transit in LEO satellite systems,” published in the IEEE International Conference on Communications and Mobile Computing, W. Jiang et al. introduced an SNR-based adaptive modulation mechanism to offset the interference caused by solar transit and reduce the resource waste caused by frequent link switching. However, none of the above methods fully consider the network characteristics of ULSN, such as its complex structure and frequent link dynamics, and are difficult to adapt to larger-scale and rapidly changing constellation deployment environments.
[0005] A search also found that the ACM mechanism is widely used in terrestrial communication networks and satellite communication systems to deal with parameter selection problems in dynamic channel environments. In an article titled "Link-level performance analysis of DVB standards in ultra-dense LEO satellite-terrestrial networks" published by X. Zhang et al. at the IEEE Vehicular Technology Conference in 2024, they introduced an LSTM network to predict the channel state and selected the appropriate MCS through an offline calculation table to effectively balance the BER and system complexity. However, this method only takes BER as the optimization goal and does not comprehensively consider other key performance indicators of the system. In addition, deep reinforcement learning (DRL) is widely used in various decision-making tasks in the field of wireless communications because it combines deep learning with policy optimization, can obtain reward signals through environmental interaction, and realize adaptive learning of complex tasks. In their 2021 article titled “Deepreinforcement learning based adaptive modulation with outdated CSI” published in IEEE Communications Letters, S. Mashhad et al. used the DRL algorithm to replace the traditional offline calculation method and maximized the system throughput under the average BER constraint by setting a specific reward function. In their 2020 article titled “DQN-based adaptive modulation scheme over wireless communication channels” published in IEEE Communications Letters, D. Lee et al. further proposed a method combining a differentiable neural dictionary with neural scenario control (NEC) to improve the search efficiency and convergence performance of DRL.
[0006] In summary, the problems of the existing technologies are as follows: (1) the existing strategies for coping with solar eclipse interference fail to fully consider the complex characteristics of frequent switching and dynamic topology in ULSN, making it difficult to ensure the continuity and stability of communications in large-scale networks; (2) the existing channel prediction and MCS selection methods mostly rely on static or offline table calculations, which cannot respond to rapidly changing channel conditions in a timely manner and are difficult to meet the needs of real-time modulation and adjustment; (3) although some studies have introduced reinforcement learning algorithms, there are still deficiencies in computing complexity control and strategy convergence performance in dealing with large-scale network structures. Summary of the Invention
[0007] Purpose of the invention: In response to the problems existing in the prior art, the purpose of the present invention is to provide an adaptive coding modulation method and system for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network, thereby achieving real-time response and efficient resistance to solar eclipse interference and improving system performance.
[0008] Technical solution: To achieve the above-mentioned purpose, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides an adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network, comprising the following steps:
[0010] Modeling of satellite-to-ground and inter-satellite communication links for ultra-dense low-orbit satellite networks, and modeling of the timing of solar eclipses and their interference to low-orbit satellite networks;
[0011] The ground station estimates and predicts the satellite-to-ground channel state and returns the predicted signal-to-noise ratio (SNR) to the satellite.
[0012] Based on the SNR returned by the ground station and the SNR generated by the solar eclipse, a deep reinforcement learning method is used to select the coding modulation scheme (MCS) for downlink data transmission, and to decide whether to select a lower-level satellite for data offload before transmitting it to the ground station. The input state of the deep reinforcement learning includes the sum of the SNR returned by the ground station and the SNR generated by the solar eclipse, the satellite number used for communication with the ground station, and the satellite number used to establish the inter-satellite link. The output action is the MCS, and the reward takes into account the system's bit error rate and spectrum utilization.
[0013] Furthermore, considering that a certain layer of low-orbit satellites continuously transmits data downlink to a fixed ground station, and that solar eclipse interference mainly affects the receiving end in the downlink of the satellite communication system, the received signal-to-noise ratio during the solar eclipse interference period is predicted based on the downlink channel model, so as to dynamically adjust the MCS to mitigate the interference impact.
[0014] Furthermore, the ground station uses the least squares (LS) method to estimate the SNR based on the pilot data in the received data, uses the long short-term memory network (LSTM) to perform time series prediction based on the historical SNR data, and returns the predicted SNR to the satellite.
[0015] Furthermore, the deep reinforcement learning method uses the proximal policy optimization (PPO) algorithm based on knowledge distillation (KD) to guide the student model to select the optimal MCS faster in new environments through an offline generated teacher model.
[0016] Furthermore, hard targets and soft targets are used to jointly train the student model. The hard target is the original loss function of the student model, ensuring that the student model can extract useful information from the environment; the soft target is the cross-entropy loss between the output of the student model and the output of the teacher model, enabling the student model to learn the strategy of the teacher model.
[0017] Furthermore, the policy loss function and value loss function of the student model are as follows:
[0018]
[0019] Among them, ρ is the weight parameter, Temp is the distillation temperature, θ and They are the policy network and value network parameters, L Pf (θ) and are hard targets, namely the true policy loss function and value loss function of the student model; and They are the soft targets corresponding to the policy network and the value network, respectively, and are defined as the cross entropy loss between the output of the student model and the output of the teacher model; y i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the student model, z i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the teacher model.
[0020] Furthermore, the satellite downlinks data to the ground station and performs the following steps:
[0021] When the satellite in the N-1 layer transmits data to the ground station, if the interference of the solar eclipse on the ground station is less than the set switching threshold, the data is directly downloaded to the ground station; otherwise, the satellite in the N layer is searched for data unloading, and the satellite in the N layer then downloads the data to the ground station; when selecting the satellite in the N layer, the satellite that is not interfered by the solar eclipse and has the smallest SNR for communication with the ground station is selected.
[0022] In a second aspect, the present invention provides an adaptive coding and modulation system for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network, comprising:
[0023] A simulation module is used to model the satellite-to-ground and inter-satellite communication links of an ultra-dense low-orbit satellite network, as well as to model and calculate the timing of solar eclipses and their interference to the low-orbit satellite network;
[0024] The channel SNR prediction module is used by the ground station to estimate and predict the satellite-to-ground channel status and return the predicted SNR to the satellite;
[0025] The adaptive coding and modulation module is used to select the coding and modulation scheme (MCS) for downlink data transmission based on the SNR returned by the ground station and the SNR generated by the solar transit using a deep reinforcement learning method, and decide whether to select a lower-level satellite for data offload and then transmit it to the ground station. The input state of the deep reinforcement learning includes the sum of the SNR returned by the ground station and the SNR generated by the solar transit, the satellite number used for communication with the ground station, and the satellite number used to establish the inter-satellite link. The output action is the MCS, and the reward takes into account the system's bit error rate and spectrum utilization.
[0026] In a third aspect, the present invention provides a computer system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor. When the computer program / instruction is executed by the processor, the steps of the adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network are implemented.
[0027] In a fourth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the adaptive coding modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network.
[0028] Beneficial effects: Compared with the existing technology, the advantages and positive effects of the present invention are as follows: First, the ACM method for alleviating solar eclipse interference proposed in the present invention can dynamically adjust the MCS according to the actual channel state, effectively reducing the bit error rate while significantly improving the spectrum utilization rate; second, the present invention uses the LSTM time series prediction model, which can accurately predict the channel state during solar eclipse interference, realize interference estimation and modulation scheme pre-configuration in advance, and improve the system's adaptability to high-latency satellite-to-ground communication environments; third, the introduction of the KD-PPO algorithm for MCS selection not only effectively reduces the algorithm complexity and satellite-side computing load, but also takes into account the multi-objective optimization requirements of system performance indicators, reflecting the practicality and advancement of the present invention in actual engineering deployment. In summary, based on the characteristics of solar eclipse interference and the actual needs of ultra-dense satellite networks, the present invention constructs an intelligent modulation mechanism that integrates channel prediction, real-time adaptive modulation selection and computational efficiency control, which can effectively improve the anti-interference ability and resource utilization efficiency of the communication system in a complex space environment, and promote the development of a new generation of satellite communication networks towards higher robustness and intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a diagram of a ULSN solar eclipse scenario according to an embodiment of the present invention.
[0030] Figure 2 1 is a diagram of the satellite-to-ground and inter-satellite link layer mechanisms of an embodiment of the present invention; (a) is the satellite-to-ground communication link, and (b) is the inter-satellite communication link.
[0031] Figure 3 4 is a KD-PPO algorithm framework diagram of an embodiment of the present invention.
[0032] Figure 4 This is a schematic diagram of total SNR attenuation when the first-layer satellite communicates with the GS after considering solar eclipse interference, provided by an embodiment of the present invention.
[0033] Figure 5 This is a schematic diagram of the second-layer satellite being interfered with by inter-satellite solar transit provided by an embodiment of the present invention.
[0034] Figure 6 This is a schematic diagram comparing the average bit error rates of networks using different algorithms provided by an example of the present invention.
[0035] Figure 7 This is a schematic diagram comparing the average spectrum utilization of networks using different algorithms provided by an example of the present invention.
[0036] Figure 8 It is a schematic diagram of the total SNR change and corresponding MCS selection provided by an example of the present invention. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following detailed description of an embodiment of the present invention is given in conjunction with the accompanying drawings. This embodiment is implemented based on the technical solutions of the present invention, and provides a detailed implementation method and specific operation process. It should be understood that the specific examples described herein are only used to illustrate the present invention, and the scope of protection of the present invention is not limited to the following embodiments.
[0038] An embodiment of the present invention discloses an adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network, which mainly includes:
[0039] Step 1: Model the satellite-to-ground and inter-satellite communication links of the ultra-dense low-orbit satellite network, and calculate the time when the solar eclipse occurs and the interference to the low-orbit satellite network.
[0040] In this step, based on the physical layer standards of Digital Video Broadcasting-Satellite Second Generation Extensions (DVB-S2X) and Consultative Committee for Space Data Systems (CCSDS), satellite-to-ground and inter-satellite communication links for an ultra-dense low-orbit satellite network are established. Based on the International Telecommunication Union (ITU) standard ITU-R BO.1506, the occurrence time of solar eclipses is modeled to quantify the interference impact of solar eclipses on the downlink receiver.
[0041] Specifically, in some embodiments, Figure 2 As shown in the figure, the DVB-S2X standard is used to model the satellite-to-ground communication link in a ULSN. During signal transmission, BCH codes are first introduced as the outer error correction mechanism, and LDPC codes as the inner error correction mechanism. The bit sequence is then interleaved, followed by symbol mapping and modulation. The receiving end then performs demodulation, symbol demapping, bit deinterleaving, LDPC decoding, and BCH decoding, completing the complete data recovery process. The CCSDS standard is used to model the inter-satellite communication link in a ULSN. The transmitting end uses Reed-Solomon coding as the forward error correction technology, performs a bit interleaving, and then performs non-return-to-zero BPSK modulation. The receiving end performs demodulation, bit deinterleaving, and Reed-Solomon decoding.
[0042] Using ITU-R BO.1506 to model solar eclipses, the angle between the satellite, the sun, and the ground station is calculated based on their relative positions. This allows prediction of when solar eclipse interference will occur in satellite-to-ground and intersatellite links within ULSN. Furthermore, by analyzing the increase in noise temperature caused by the sun on the receiving antenna, the intensity of solar eclipse interference in the downlink—the increase in signal-to-noise ratio (SNR) caused by the solar eclipse—is quantified, providing a basis for subsequent compensation strategies.
[0043] Step 2: The ground station estimates and predicts the satellite-ground channel state and returns the predicted signal-to-noise ratio (SNR) to the satellite.
[0044] In some embodiments, the ground station can use the least squares (LS) method to estimate the SNR based on the pilot data in the received data, use the long short-term memory network (LSTM) to perform time series prediction based on the historical SNR data, and return the predicted SNR to the satellite.
[0045] Step 3: Based on the SNR returned by the ground station and the SNR generated by the solar transit, a deep reinforcement learning method is used to select the coding modulation scheme (MCS) for downlink data transmission, and to decide whether to select a lower-level satellite for data offload before transmitting it to the ground station. The input state of the deep reinforcement learning includes the sum of the SNR returned by the ground station and the SNR generated by the solar transit, the satellite number used for communication with the ground station, and the satellite number used for establishing the intersatellite link. The output action is the MCS, and the reward takes into account the system's bit error rate and spectrum utilization.
[0046] In this step, a certain layer of satellites needs to continuously download data to a fixed ground station. During the download process, an adaptive coding modulation (ACM) method based on deep reinforcement learning is used to alleviate the interference caused by solar eclipse.
[0047] In some embodiments, the ACM method based on deep reinforcement learning can perform the following steps: the transmitter transmits data with an inserted pilot signal, and the receiver extracts the pilot data using the least squares (LS) method to estimate the signal-to-noise ratio (SNR). The receiver inputs historical SNR data into an LSTM network for time series prediction, predicts the SNR at the next moment, and returns the SNR information to the transmitter. The transmitter uses the knowledge distillation-based proximal policy optimization (KD-PPO) algorithm to select the optimal coding and modulation scheme (MCS) for downlink data transmission based on the received SNR information and the results obtained by the on-board solar transit calculation program. A switching threshold is set. When the total SNR of the transit and the channel exceeds this threshold, a satellite that is not affected by intersatellite transit interference is selected for data offloading. This satellite then executes the ACM method to download the data to the ground station. When the total SNR of the transit and the channel does not exceed this threshold, the ACM method is used to directly download the data to the ground station.
[0048] The following combination Figure 1 The solar eclipse scenario in ULSN shown in the figure is used to describe in detail an adaptive coding and modulation method based on deep reinforcement learning for mitigating solar eclipse interference disclosed in an embodiment of the present invention. First, consider the scenario of a fixed ground station (GS). The GS faces many overhead satellites, which need to transmit the data they carry to the GS. When the GS receives a signal from a satellite S ij Download data if the solar outage interference is below the threshold L Thresh , then GS directly uses the ACM strategy to communicate with it. Otherwise, it is necessary to first unload the data to other satellites and then download it to the ground station. Due to the dense distribution of satellites in ULSN, the transmission between satellites in the same layer is still very likely to be affected by the interruption of sunlight between the satellite and the ground. Therefore, in order to avoid direct sunlight on the satellite-borne laser signal receiver, the satellite S in the lower layer that is not affected by solar eclipse interference and has the lowest SNR in communication with the ground station is selected. qk Unload data, then S qk An ACM strategy is used for satellite-to-ground communications. The prediction and interference calculation of solar eclipse events refer to Recommendation ITU-R BO.1506, the physical layer mechanism of satellite-to-ground link communications refers to the DVB-S2X standard, and the physical layer standard of inter-satellite link communications refers to the CCSDS standard.
[0049] The basic goal of this embodiment is to implement an efficient adaptive coding modulation method through deep reinforcement learning, specifically including channel state estimation, channel state prediction, and MCS selection strategy. In the channel state estimation part, the LS estimator is used to determine a set of optimal model parameters by minimizing the sum of the squares of the errors between the model prediction value and the actual observation value, thereby achieving the best fit effect. Assuming that the pilot data sent by the transmitter is X, the pilot data received by the receiver is Y, the channel matrix is H, and the channel noise is n, by minimizing the square error between the received signal and the reconstructed signal, the estimated value of the channel matrix can be obtained:
[0050]
[0051] The SNR obtained by the LS estimator can be calculated as follows:
[0052]
[0053] In the channel state prediction part, this embodiment uses a long short-term memory neural network (LSTM) for time series prediction. LSTM is a special recurrent neural network (RNN). Compared with traditional RNN, LSTM can effectively alleviate the long-term dependency problem, retain long-term memory information, and solve stability problems such as gradient disappearance or explosion during training. t-1 ) Input into LSTM network to obtain the SNR prediction value at time t
[0054] like Figure 3 As shown in the figure, in the MCS selection strategy part, the proximal policy optimization (PPO) algorithm based on knowledge distillation (KD) is adopted, so that the student PPO model in the real scene can refer to the more complex teacher PPO model that has been trained offline in advance during training. PPO is a DRL algorithm based on policy gradient. It selects actions through the interaction between the agent and the environment according to the current policy, with the goal of maximizing the cumulative reward. In this embodiment, the PPO agent has the following parts: S = {(S1, S2, S2)} is a three-dimensional state space, where S1 = {s 11 ,s 12 ,s 13 ,...} represents the continuous SNR value, and its domain is [0,35]. 21 ,s 22 ,s 23 ,...} represents the high-level satellite number used to communicate with GS, i.e. ij. S3={s 31 ,s 32 ,s 33,...} represents the satellite number used to establish the intersatellite link, i.e. qk. If the high-level satellite directly communicates with the ground, S3=0; A={a1,a2,...,a 42} represents the action taken by the agent. The agent should be deployed at the cluster head of a satellite cluster to inform the satellite communicating with the GS which MCS to adopt. It consists of 42 MCSs, including '8APSK 100 / 180', '16APSK 18 / 30', etc.; R = {r(s, a)} is the reward function that the agent can immediately obtain by interacting with the environment. In order to balance the system's bit error rate and spectrum utilization, this embodiment assumes that the agent is in state s at time t. t Take action a t The rewards obtained are defined as follows:
[0055]
[0056] Among them, BER t is the bit error rate of network communication at time t, R t Indicates the modulation order, C t Indicates the bit rate. PPO consists of two deep learning networks, one is the policy network and the other is the value network, using θ and Represent the network parameters of the policy network and the value network respectively. Its policy network generates the current policy π θ (a t |s t ), that is, in state s t Next take action a t The probability distribution of the value network outputs the estimated state value function under this strategy The loss function of the value network is calculated by the mean squared error between the current network output and the actual return:
[0057]
[0058] Where ε is the discount factor, which is used to control the degree of attenuation of future returns, r t is the immediate reward at time t. The policy network directly adjusts the policy parameters through gradient ascent or descent to maximize the expected cumulative reward. First, the policy function is parameterized, whose goal is to maximize the cumulative expected reward obtained by the agent from the initial state s0 and interacting with the environment. The expected reward function is expressed as follows:
[0059]
[0060] At the same time, define G represents the expected return from taking action a in state s under strategy π. t represents the cumulative rewards obtained from time t. Their mathematical expressions are as follows:
[0061]
[0062] Among them, S t and A t Represents the state at time t and the action taken. Then use the policy gradient (PG) theorem to calculate the gradient and update the policy parameters:
[0063]
[0064] Where κ is the learning rate. The goal of this algorithm is to find Based on this goal, we need to find a new parameter θ′ so that J(θ′) ≥ J(θ), so the optimization objective can be achieved in the new strategy π θ′ The following is expressed as:
[0065]
[0066] The difference in the objective function between the new and old strategies can be expressed as:
[0067]
[0068] Define the time series difference residual as the advantage function AD:
[0069]
[0070] The algorithm ignores the change in state access distribution between the new and old strategies, uses the state distribution of the old strategy instead, and uses KL divergence to measure the distance between the new and old strategies. The optimization objective can be updated as follows:
[0071]
[0072] Subsequently, KL divergence can be used to measure the distance between strategies, and the overall optimization formula is as follows:
[0073]
[0074] in, Indicates the strategy obtained in the kth update Based on the actual state distribution, we make an expectation for a certain quantity. δ represents the KL distance threshold, which is used to limit the difference between the new and old policies to not exceed a preset upper limit, thereby ensuring that each policy update is sufficiently stable. PPO uses the clip method to introduce constraints into the objective function to ensure that the change between the new and old policy parameters is not too large. The loss function of the policy network can be expressed as:
[0075] L Pf (θ)=-L θ (θ′)
[0076] KD is a model compression technology in deep learning. Its core idea is to allow the student model to learn knowledge from the teacher model, thereby achieving a balance between performance and model complexity. In this embodiment, a large-scale, high-performance teacher model consisting of 1114 neurons is trained using training data with a simulation step length of 30 seconds. The teacher model generates a "softened" output probability distribution as the training target. The outputs of the policy network and the value network can be softened using the following formula:
[0077]
[0078] Where Temp is the distillation temperature, z i is the original output value of the teacher model. A higher Temp value will produce a smoother probability distribution. This embodiment uses hard targets and soft targets to jointly train the student model. The hard target is the original loss function of the student model, which ensures that the student model can extract useful information from the current environment; the soft target is the cross entropy loss between the output of the student model and the output of the teacher model, which enables the student model to learn the strategy of the teacher model. The student model inputs training data with a simulation step length of 1 minute and consists of 662 neurons. The updated strategy loss function and value loss function are as follows:
[0079]
[0080]
[0081] Among them, ρ is a weight parameter used to adjust the relative importance of hard targets and soft targets. Let y i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the student model, z i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the teacher model, then the soft target can be expressed as:
[0082]
[0083] We will simulate and verify the effectiveness of our proposed adaptive coding and modulation scheme for mitigating solar eclipse interference in an ultra-dense low-orbit satellite communication network based on deep reinforcement learning. In our ultra-dense low-orbit satellite communication scenario, the inner satellites are less likely to be affected by solar eclipse interference due to the shadowing of a large number of outer satellites. We select the two outermost satellites that can establish communication with the GS within a fixed timeframe for analysis. The ground-to-satellite physical layer is constructed according to the DVB-S2X standard, and the inter-satellite physical layer is constructed according to the CCSDS standard. We first conduct solar eclipse prediction and analysis in this ultra-dense low-orbit satellite communication scenario, and then evaluate the performance of the adaptive coding and modulation scheme.
[0084] exist Figure 4 In the figure, we obtained the total SNR attenuation of the first-layer satellite and GS communication after considering the solar eclipse interference. It can be seen that the downlink signal transmission is continuously affected by the solar eclipse, causing sudden SNR attenuation. Weak interference causes an SNR attenuation of about 19dB, and strong interference causes an SNR attenuation of about 33dB. Figure 5 In the figure, we get the interference of intersatellite solar transit on the second layer of satellites. By calculating the intersatellite solar transit event in advance, we can avoid the damage of sunlight to the onboard laser receiver when data offloading is required. Figure 6 and Figure 7 In the paper, we obtained the average bit error rate and average spectrum utilization (SU) of the simulated network using the dual deep Q network (DoubleDQN), the dual deep Q network (DuelingDQN), PPO without teacher model guidance, and KD-PPO. Here, the spectrum utilization is defined as SU = R i C i (1-BER i ). It can be seen that KD-PPO reaches convergence the fastest compared to other algorithms, and converges to the best bit error rate of 0.008 and spectrum efficiency of 4.39bps / Hz. Figure 8 In the figure, we obtained the SNR change and the corresponding MCS selection during the whole process. It can be seen that according to our solar eclipse prediction program and ACM strategy, the transmitter switches the MCS in real time and selects the intersatellite link when the solar eclipse interference is strong.
[0085] Based on the same inventive concept, an embodiment of the present invention discloses an adaptive coding and modulation system for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network, comprising:
[0086] The simulation module is used to model the satellite-to-ground and inter-satellite communication links in an ultra-dense low-orbit satellite network, as well as to calculate the timing of solar eclipses and their interference to the low-orbit satellite network. Specifically, this module constructs a simulation system for satellite-to-ground and inter-satellite communication, as well as a communication channel model, within the ultra-dense low-orbit satellite network. Based on the simulation results, the distance and position data of the satellites as they orbit the Earth are obtained. The module simulates solar eclipses within the established ultra-dense low-orbit satellite communication network, determining the links affected by the eclipse and the eclipse angle. Based on this data, the noise level at the receiving end affected by the eclipse is calculated.
[0087] The channel SNR prediction module is used by the ground station to estimate and predict the satellite-to-ground channel status and return the predicted SNR to the satellite.
[0088] The adaptive coding and modulation module is used to select the MCS for downlink data transmission based on the SNR returned by the ground station and the SNR generated by the solar eclipse using a deep reinforcement learning method, and decide whether to select the lower-level satellite for data offload and then transmit it to the ground station.
[0089] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor. When the computer program / instruction is executed by the processor, the steps of the adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network are implemented.
[0090] An embodiment of the present invention also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network.
[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An adaptive coding modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network, characterized in that: The following steps are involved: Modeling of satellite-to-ground and inter-satellite communication links for ultra-dense low-orbit satellite networks, and modeling of the timing of solar eclipses and their interference to low-orbit satellite networks; The ground station estimates and predicts the satellite-to-ground channel state and returns the predicted signal-to-noise ratio (SNR) to the satellite. Based on the SNR returned by the ground station and the SNR generated by the solar eclipse, a deep reinforcement learning method is used to select the coding modulation scheme (MCS) for downlink data transmission, and to decide whether to select a lower-level satellite for data offload before transmitting it to the ground station. The input state of the deep reinforcement learning includes the sum of the SNR returned by the ground station and the SNR generated by the solar eclipse, the satellite number used for communication with the ground station, and the satellite number used to establish the inter-satellite link. The output action is the MCS, and the reward takes into account the system's bit error rate and spectrum utilization.
2. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 1, characterized in that: Considering that a certain layer of low-orbit satellites continuously transmits data downlink to a fixed ground station, and considering that solar eclipse interference mainly affects the receiving end in the downlink of the satellite communication system, the received signal-to-noise ratio during the solar eclipse interference period is predicted based on the downlink channel model, so as to dynamically adjust the MCS to mitigate the interference impact.
3. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 1, characterized in that: The ground station uses the least squares (LS) method to estimate the SNR based on the pilot data in the received data, uses the long short-term memory network (LSTM) to perform time series prediction based on the historical SNR data, and returns the predicted SNR to the satellite.
4. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 1, characterized in that: The deep reinforcement learning method uses the proximal policy optimization (PPO) algorithm based on knowledge distillation (KD) to guide the student model to select the optimal MCS faster in new environments through an offline generated teacher model.
5. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 4, characterized in that: The student model is trained jointly with hard targets and soft targets. The hard target is the original loss function of the student model, ensuring that the student model can extract useful information from the environment; the soft target is the cross entropy loss between the output of the student model and the output of the teacher model, enabling the student model to learn the strategy of the teacher model.
6. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 5, characterized in that: The policy loss function and value loss function of the student model are as follows: Among them, ρ is the weight parameter, Temp is the distillation temperature, θ and They are the policy network and value network parameters, L Pf (θ) and are hard targets, namely the true policy loss function and value loss function of the student model; and They are the soft targets corresponding to the policy network and the value network, respectively, and are defined as the cross entropy loss between the output of the student model and the output of the teacher model; y i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the student model, z i (θ) and is the output of the i-th neuron in the output layer of the strategy network and value network of the teacher model.
7. The adaptive coding and modulation method for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network according to claim 1, characterized in that: The satellite downlinks data to the ground station and performs the following steps: When the satellite in the N-1 layer transmits data to the ground station, if the interference of the solar eclipse on the ground station is less than the set switching threshold, the data is directly downloaded to the ground station; otherwise, the satellite in the N layer is searched for data unloading, and the satellite in the N layer then downloads the data to the ground station; when selecting the satellite in the N layer, the satellite that is not interfered by the solar eclipse and has the smallest SNR for communication with the ground station is selected.
8. An adaptive coding modulation system for mitigating solar eclipse interference in an ultra-dense low-orbit satellite network, characterized in that: include: A simulation module is used to model the satellite-to-ground and inter-satellite communication links of an ultra-dense low-orbit satellite network, as well as to model and calculate the timing of solar eclipses and their interference to the low-orbit satellite network; The channel SNR prediction module is used by the ground station to estimate and predict the satellite-to-ground channel status and return the predicted SNR to the satellite; The adaptive coding and modulation module is used to select the coding and modulation scheme (MCS) for downlink data transmission based on the SNR returned by the ground station and the SNR generated by the solar transit using a deep reinforcement learning method, and decide whether to select a lower-level satellite for data offload and then transmit it to the ground station. The input state of the deep reinforcement learning includes the sum of the SNR returned by the ground station and the SNR generated by the solar transit, the satellite number used for communication with the ground station, and the satellite number used to establish the inter-satellite link. The output action is the MCS, and the reward takes into account the system's bit error rate and spectrum utilization.
9. A computer system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, wherein: When the computer program / instructions are executed by a processor, the steps of the adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the adaptive coding and modulation method for alleviating solar eclipse interference in an ultra-dense low-orbit satellite network according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Low earth orbit satellite communication AMC optimization method based on AI-rule fusion decision
CN121240133A
Multi-constellation satellite signal intelligent identification and adaptive acquisition method and system
CN121348365A