A Satellite Adaptive Coding and Modulation Method Based on Deep Reinforcement Learning

Through deep reinforcement learning and neural networks to realize adaptive coding modulation in satellite communication, the problem of rough channel state division in traditional technology and relying on fixed channel models is solved, and optimized spectrum efficiency and high-quality transmission are achieved.

CN116192227BActive Publication Date: 2025-06-10NANJING UNIV

Patent Information

Application Number
CN202310011797.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-05
Publication Date
2025-06-10
Estimated Expiration
2043-01-05

AI Technical Summary

Technical Problem

Traditional adaptive coding and modulation technology has problems in satellite communication scenarios such as rough channel state division, dependence on fixed channel models, resource waste and throughput performance affected.

Method used

Adaptive coding modulation method based on deep reinforcement learning is adopted, through reinforcement learning agents, learn optimal strategies in continuous interaction and iteration, make optimal decisions, introduce neural networks to avoid excessive state space, and accelerate the convergence of results through dual networks.

Benefits of technology

The goal of optimizing spectrum efficiency under the premise of the lowest bit error rate is achieved, and the high-quality transmission of satellite communications and the optimal use of system resources are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192227B_ABST
    Figure CN116192227B_ABST
Patent Text Reader

Abstract

A satellite adaptive modulation method based on deep reinforcement learning: 1) Perform initialization operations; initialize the state space, action space, and greedy parameter; 2) The signal receiving end on the ground receives the signal from the satellite downlink and extracts the pilot information in the current frame for signal-to-noise ratio (SNR) estimation. After the receiving end calculates the SNR estimation result, it is transmitted to the ground sending end through the ground feedback link; 3) The sending end translates the selected action into the corresponding modulation method and coding rate; 4) Determine whether the current iteration number is an integer multiple of the preset network update step number. If so, go to step 5) for network update; if not, update the SNR state and return to 2) to enter the next round of iteration; 5) Introduce the concept of a dual network and improve the learning effect and accelerate the convergence of the results by optimizing the structure of the neural network. 6) Update the SNR state, increment the greedy parameter, and return to 2) for the next round of iteration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applied to a satellite-ground communication transmission system, and an adaptive coding and modulation method for a satellite communication system is realized based on deep reinforcement learning. Background Art

[0002] In recent years, satellite communication technology has continuously attracted people's attention. With its unique advantages, it can well make up for the deficiencies of the current terrestrial mobile communication system. The characteristics of wide coverage area and being unrestricted by natural geographical conditions of the satellite communication system enable full coverage of communication services in places where ground communication is blocked, such as remote areas. Even when natural disasters such as earthquakes and tsunamis occur and the ground communication system may face the dilemma of being unable to work, satellite communication can still provide reliable communication services for users in its unique communication mode.

[0003] However, due to the particularity of satellite communication methods, satellite communication is subject to interference from many factors. Traditional wireless communication systems use a single modulation and coding method. However, in satellite channels, doing so will cause the performance to vary from good to bad as the satellite channel changes. Adaptive coding modulation technology can well solve this problem. When the communication system senses that the channel environment is poor, it can use a low-order modulation method and a low coding rate; when it senses that the channel quality is good, it can use a high-order modulation method and a high coding rate. However, traditional adaptive modulation and coding technology has great limitations. According to Rajesh et al. (Rajesh M N, Shrisha B K, Rao N, et al. An analysis of BER comparison of various digital modulation schemes used for adaptive modulation[C]. 2016 IEEE International Conference on Recent Trends in Electronics, Information & Communication Technology (RTEICT). IEEE, 2016: 241-245.), we found that traditional adaptive coding modulation technology is mainly based on a lookup index table. By sensing the channel state at the current moment, the corresponding modulation and coding scheme is indexed from the lookup table. This will lead to two problems: (1) The division of the channel state by the index table is often a rough interval division, and it is impossible to make all channel states select the most suitable modulation and coding method; (2) The lookup table method used in traditional adaptive coding modulation technology is based on a fixed channel model. However, due to the time-varying nature of the satellite channel, the lookup table relied on by the transmitter is likely to be mismatched with the current channel. Even if the channel state information fed back by the receiver is real-time and accurate, it will still lead to the selection of the wrong modulation and coding method, resulting in resource waste, seriously affecting the throughput performance of the system, and greatly weakening the effectiveness of traditional adaptive modulation and coding schemes.

[0004] Based on the above problems, according to the unique real-time adaptive learning characteristics of deep reinforcement learning, the present invention proposes an adaptive coding modulation method that does not rely on a fixed channel model. The reinforcement learning agent makes the scheme decision, replacing the traditional rough interval division with an infinite state. At the same time, according to the characteristics of the time-varying satellite channel, the neural network is used to avoid the problem of difficult convergence caused by the too large state space generated in the traditional reinforcement learning process. On this basis, the structure of the neural network is further optimized, and the output layer is divided into an advantage layer and a value layer to improve the learning effect and accelerate the convergence of the results. Finally, on the premise of achieving the lowest bit error rate, the goal of optimizing the spectral efficiency is realized. Summary of the Invention

[0005] The object of the present invention is to provide a satellite communication adaptive modulation and coding method based on deep reinforcement learning, so as to solve the problem of poor communication conditions caused by a single modulation and coding method in the satellite communication scenario, and overcome the defects of traditional lookup table-based adaptive coding and modulation technologies. The reinforcement learning agent learns the optimal strategy and makes optimal decisions through continuous interaction and iteration with the environment; a neural network is introduced to avoid the problem of difficult convergence caused by the too large state space in the traditional reinforcement learning process; considering the characteristic that not all modulation and coding schemes in the communication system are worthy of key attention under a specific signal-to-noise ratio state, the concept of a dual network is introduced to accelerate the convergence of the results. The above measures enable the satellite communication process to achieve the goal of optimizing the spectral efficiency on the premise of achieving the lowest bit error rate, thus providing guarantee for the high-quality transmission of satellite communication and the optimal use of system resources.

[0006] The technical solution adopted by the present invention is: a satellite adaptive coding and modulation method based on deep reinforcement learning, including the following steps:

[0007] Step (1), perform initialization operations. Initialize the state space, action space, greedy parameter, etc. In the satellite-ground communication scenario, the state space in the three elements of reinforcement learning (state space, action space, reward function) is the set of all signal-to-noise ratio values received by the receiving end under the current channel, so as to avoid the drawback of making decisions by the rough division interval of the traditional lookup table method; the action space is the set of all modulation and coding methods in the satellite communication system, and different modulation and coding methods are defined as different actions; the reward function is used to measure the value of different actions in different states. In the current system, the reward function is set with the spectral efficiency as the standard; the exploration probability is the probability of performing random exploration. According to this set probability, random exploration is performed in some cases, and the numerical probability of taking the optimal action based on experience at other times, so as to reduce the correlation of training samples;

[0008] Step (2), the ground signal receiving end receives the signal from the satellite downlink and extracts the pilot information in the current frame for signal-to-noise ratio estimation. After the receiving end calculates the signal-to-noise ratio estimation result, it is transmitted to the ground sending end through the ground feedback link. To avoid the problem of difficult convergence of the Q value in the Q table caused by the too large state space of traditional reinforcement learning through a neural network, the ground sending end inputs the signal-to-noise ratio state received and calculated based on the ground receiving end signal into the evaluation network and uses the greedy algorithm to judge whether to explore with a corresponding probability. If so, a random action is selected from the action space; if not, the action that can obtain the maximum Q value is selected from the output of the evaluation network;

[0009] Step (3): The sending end translates the selected action into the corresponding modulation mode and coding rate according to the selected action, and transmits the signal according to this modulation and coding method. After passing through the satellite uplink channel, the specified modulation and coding method is notified to the satellite for use in the next transmission. After the signal reaches the receiving end, the corresponding reward value is calculated, and the current state, action, reward and other elements are stored in the experience pool as a group of samples. The reward calculation formula is as follows:

[0010]

[0011] where M represents the current modulation order, snr represents the signal-to-noise ratio, and ber mcs (snr) represents the bit error rate of the input signal-to-noise ratio under the current modulation and coding method, and mcs represents the modulation and coding method;

[0012] Step (4): Determine whether the current iteration number is an integer multiple of the preset network update step. If so, go to step (5) for network update; if not, update the signal-to-noise ratio state and return to step (2) to enter the next round of iteration;

[0013] Step (5): Introduce the concept of a dual network, that is, the output layers of the evaluation network and the target network are both split into a value layer and an advantage layer. The input of the value layer is only the signal-to-noise ratio state, which only focuses on the current channel quality. By optimizing the structure of the neural network, the learning effect is improved and the convergence of the result is accelerated. The input of the advantage layer is the signal-to-noise ratio state and the modulation and coding method, which is responsible for focusing on the value of the modulation and coding method under the current signal-to-noise ratio state. Finally, the output items of the two sub-networks are aggregated into the final Q-value output. Extract several groups of samples from the experience pool as the input of the training set. Further determine whether the current iteration number is an integer multiple of the preset target network update step. If not, directly train and update the evaluation network without updating the target network; if so, update the parameters of the evaluation network to the target network, and use the input of the training set and the result calculated by the formula as the new evaluation network. Up to this step is a reinforcement learning round. The iterative update formula of the evaluation network is as follows:

[0014]

[0015] where Q eval is the output of the evaluation network, Q target is the output of the target network, S is the current signal-to-noise ratio state, A is the current action, and γ is the learning rate.

[0016] The final output is aggregated by the advantage layer and the value layer, and its aggregation relationship formula is as follows:

[0017] V(S; θ,β)+Advan(S,A; θ,α)→Q(S,A; θ,α,β)

[0018] Among them, V represents the output of the value layer, Advan represents the output of the advantage layer, θ is the parameter shared by the value layer and the advantage layer, α is the unique parameter of the advantage layer, and β is the unique parameter of the value layer;

[0019] Step (6), update the signal-to-noise ratio state, increment the greedy parameter, and return to step (2) for the next round of iteration.

[0020] Use the reinforcement learning method, and use the reinforcement learning agent to make decisions to replace the method based on the lookup modulation coding combination index table in the traditional adaptive modulation and coding technology.

[0021] Use two neural networks, namely the evaluation network and the target network, to replace the Q-table of the traditional reinforcement learning Q-Learning method, so as to avoid the problem of non-convergence caused by the overly large state space and be applicable to decision-making problems with an infinite state space.

[0022] During the iteration process, the parameters of the evaluation network are updated in each loop, while the parameters of the target network need to be updated at intervals to reduce the correlation of training samples and prevent overfitting.

[0023] Introduce the concept of a dual network, that is, the output layers of the evaluation network and the target network are both split into a value layer and an advantage layer. Among them, the input of the value layer is only the signal-to-noise ratio state, which only focuses on the current channel quality. The input of the advantage layer is the signal-to-noise ratio state and the modulation and coding method, which is responsible for focusing on the value of the modulation and coding method under the current signal-to-noise ratio state. Finally, the output items of the two sub-networks are aggregated into the final Q-value output. By optimizing the structure of the neural network, the learning effect is improved, the convergence of the result is accelerated, and the purpose of optimizing the algorithm is finally achieved.

[0024] The satellite adaptive modulation and coding method based on Deep Reinforcement Learning (DRL) of the present invention combines the perception ability of deep learning and the decision-making ability of reinforcement learning. Using the principle of Q-learning in reinforcement learning, in the satellite communication scenario, by reasonably mapping the three elements of reinforcement learning: state space, action space, and reward, the purpose of selecting the action with the highest value, that is, the optimal modulation and coding method, in different state spaces is achieved through the reward. At the same time, by introducing a neural network, an approximate representation of the value function is carried out, avoiding the problem of excessive state space and thus difficult convergence encountered in the traditional reinforcement learning process. Furthermore, in view of the characteristic that in the satellite communication scenario, at a specific signal-to-noise ratio state, not all modulation and coding methods are worthy of key attention, the concept of a dual network is introduced, and the output is divided into two branches: a value layer and an advantage layer. Among them, the value layer is only responsible for paying attention to the current channel quality, and the advantage layer is responsible for paying attention to the value of the modulation and coding strategy at the current signal-to-noise ratio state. The two are aggregated into the final output layer. By optimizing the structure of the neural network, the learning effect is improved, the convergence of the result is accelerated, and the purpose of optimizing the algorithm is finally achieved.

[0025] Beneficial effects: The present invention formulates an adaptive coding and modulation method for the unique transmission characteristics of satellite communication, solving the problem that the communication effect of the traditional single modulation method is sometimes good and sometimes bad in the satellite communication scenario. At the same time, deep reinforcement learning is adopted to avoid the defect of rough division of the channel state interval in the traditional look-up table method, and the most suitable modulation and coding method can be selected for all channel states through infinite states. Through the interactive learning between the reinforcement learning agent and the environment, a suitable modulation and coding scheme can also be selected when facing the complex satellite channel with channel changes. And on this basis, a neural network is introduced, thus overcoming the defect that the traditional reinforcement learning cannot converge due to the excessive state space. And the output of the neural network is divided into an advantage layer and a value layer to accelerate the convergence of the result. The above measures enable the satellite communication process to achieve the goal of optimizing the spectral efficiency on the premise of achieving the lowest bit error rate, thus providing a guarantee for the high-quality transmission of satellite communication and the optimal use of system resources. Description of the Drawings

[0026] Figure 1 It is a block diagram of an adaptive modulation and coding link model based on the DVB-S2 standard.

[0027] Figure 2 It is a block diagram of an adaptive modulation and coding system based on the DVB-S2 standard.

[0028] Figure 3 It is a flowchart of the execution of the satellite adaptive modulation and coding method based on deep reinforcement learning.

[0029] Figure 4It is a schematic structural diagram of a dual network. Specific implementation manner

[0030] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention all fall within the scope defined by the appended claims of this application.

[0031] A satellite adaptive modulation and coding method based on deep reinforcement learning of the present invention is used for high-benefit transmission in satellite communication scenarios. Taking a satellite link based on the DVB-S2 standard as an example, the block diagrams of its adaptive modulation and coding link model and system can refer to Figure 1 and Figure 2 , the decision-making method proposed by the present invention is the coding rate and modulation mode. According to the DVB-S2 standard (Digital Video Broadcasting Second generation framing structure, channel coding and modulation systems for broadcasting, interactive services, newsgathering and other broadband satellite applications[J]. Final Draft ETSI EN, 2014, 302: 2005-01.); we know that the DVB-S2 standard provides 4 modulation modes and 11 coding rates. The modulation modes are QPSK, 8PSK, 16APSK, and 32APSK respectively, and the coding rates are 1 / 4, 1 / 3, 2 / 5, 1 / 2, 3 / 5, 2 / 3, 3 / 4, 4 / 5, 5 / 6, 8 / 9, and 9 / 10 respectively. These 4 modulation modes and 11 coding rates can be combined into 28 modulation and coding methods in total. These 28 modulation and coding methods constitute the action space. Through the method of deep reinforcement learning, the signal-to-noise ratio state information is input into the reinforcement learning agent, and corresponding action decisions are made according to the reward function and exploration probability, that is, the modulation and coding method is adjusted to affect the satellite communication system, and the corresponding state, action and other parameters are put into the experience pool. The execution flowchart can refer to Figure 3 , and the specific steps are as follows:

[0032] 1. System initialization is required during the first run. The specific operation is as follows: The sending end sends a blank frame, the receiving end performs channel estimation on the received blank frame and returns the information of the channel estimation. After receiving the frame information, the receiving end makes a policy selection and proceeds with the transmission of the next frame. The working process of the sending end module under the DVB-S2 standard can refer to Figure 1As shown in the figure, the transmitting end receives data and adaptive modulation and coding control information through the input interface, and then forms baseband frame data through corresponding pattern recognition, flow adaptation and insertion of baseband signaling. The baseband data frame is subjected to outer coding, inner coding and interleaving to form a forward error correction frame. After that, it goes through frame segmentation, addition of frame header information and signaling blocks, and scrambling addition to become physical layer frame data. The physical layer frame data can be converted into radio frequency signals through baseband signal filtering and quadrature modulation and enter the satellite communication channel through the transmitting antenna;

[0033] 2. Use two neural networks, namely the evaluation (neural) network and the target (neural) network, to replace the traditional Q-value index table, so that it can be applied to decision-making problems with an infinite state space. As Figure 2 shown in the figure, DVB-RCS2 and DVB-S2 respectively represent the uplink and downlink satellite links. The signal receiving end on the ground receives the signal from the satellite downlink and extracts the pilot information in the current frame for signal-to-noise ratio estimation. After the receiving end calculates the signal-to-noise ratio estimation result, it is transmitted to the ground transmitting end through the ground feedback link. The current signal-to-noise ratio state is input into the evaluation network. Using the greedy algorithm, it is judged whether to explore based on the set exploration probability. If so, a random action is selected from the action space; if not, the action that can obtain the maximum Q value is selected from the output of the evaluation network. During the iteration process, the parameters of the evaluation network are updated in each loop, and the parameters of the target network need to be updated after a certain period of time. The update time varies according to different targets and can be freely set according to actual needs. Here, we set it to update once every 200 reinforcement learning rounds to reduce the correlation of training samples and prevent overfitting;

[0034] 3. The transmitting end translates the selected action into the corresponding modulation method and coding rate according to the selected action, and sends the signal according to this modulation and coding method. After passing through the satellite channel, it reaches the receiving end and obtains the corresponding reward value. The current state, action, reward and other elements are stored in the experience pool as a group of samples. The reward calculation formula is as follows:

[0035]

[0036] where M represents the modulation order in the current action, that is, the modulation and coding method, snr represents the signal-to-noise ratio, ber mcs (snr) represents the bit error rate of the input signal-to-noise ratio under the current modulation and coding method, and mcs represents the modulation and coding method;

[0037] 4. Determine whether the current iteration count is an integer multiple of the preset network update steps. If so, proceed to the next step, i.e., step 5 for network update; if not, update the signal-to-noise ratio status and enter the next round of iteration. The number of steps for updating the network may vary according to different objectives and can be freely set according to actual needs. In this example, we set it to update once every 200 reinforcement learning rounds;

[0038] 5. Introduce the concept of a dual network, that is, the output layers of both the evaluation network and the target network are split into a value layer and an advantage layer. The structures of the value layer and the advantage layer in the dual network refer to Figure 4 , where the input of the value layer is only the signal-to-noise ratio status and it only focuses on the current channel quality. The input of the advantage layer is the signal-to-noise ratio status and the modulation and coding scheme, which is responsible for focusing on the value of the modulation and coding scheme under the current signal-to-noise ratio status. Finally, the output items of the two sub-networks are aggregated into the final Q-value output. In theory, the aggregation method is to directly add the two, but in actual operation, we also need to subtract the mean value of the output of the advantage layer to achieve the purpose of decentralized processing. Extract several groups of samples from the experience pool as the input of the training set. Further determine whether the current iteration count is an integer multiple of the preset target network update steps. If not, directly train and update the evaluation network without updating the target network; if so, update the parameters of the evaluation network to the target network, and use the input of the training set and the result calculated by the formula as the new evaluation network to obtain a more accurate Q eval output;

[0039] 6. Update to the signal-to-noise ratio status of the next frame, increment the greedy parameter, and return to step 2 for the next round of iteration. The size of the greedy parameter is set differently according to different experimental conditions and objectives. In this example, the greedy parameter increments by 0.0005 per round, and the initial greedy parameter is 0.1.

[0040] In summary, a satellite adaptive modulation and coding method based on deep reinforcement learning implemented by the present invention solves the problem of the inconsistent communication effect in the satellite communication scenario compared with the traditional single modulation method. At the same time, by using the method of deep reinforcement learning, it avoids the defect of the rough division of the channel state interval in the traditional lookup table method, and can select the most suitable modulation and coding scheme for all channel states through infinite states. Through the reinforcement learning agent, it can also select a suitable modulation and coding scheme when facing the complex satellite channel with changing channels. And on this basis, a neural network is introduced to overcome the defect that the state space of traditional reinforcement learning is too large to converge. And the output of the neural network is divided into an advantage layer and a value layer to accelerate the convergence of the result. The above measures ensure the high-quality transmission of satellite communication and the optimal use of system resources by achieving the goal of optimizing the spectrum efficiency on the premise of the lowest bit error rate.

[0041] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A satellite adaptive modulation method based on deep reinforcement learning, characterized in that, it includes the following steps: Step (1), perform initialization operations; initialize the state space, action space, and greedy parameter; in the satellite-ground communication scenario, the state space in the three elements of reinforcement learning, namely the state space, action space, and reward function, is the set of all signal-to-noise ratios received by the receiving end under the current channel, so as to avoid the drawback of making decisions by roughly dividing intervals in the traditional lookup table method; the action space is the set of all modulation and coding methods in the satellite communication system, and different modulation and coding methods are defined as different actions; the reward function is used to measure the value of different actions in different states. In the current system, the reward function is set based on spectral efficiency; the exploration probability is the probability of performing random exploration. According to this set probability, random exploration is performed in some cases, and the numerical probability of taking the optimal action based on experience at other times is used to reduce the correlation of training samples; Step (2), the signal receiving end on the ground receives the signal from the satellite downlink and extracts the pilot information in the current frame for signal-to-noise ratio estimation. After the receiving end calculates the signal-to-noise ratio estimation result, it is transmitted to the ground sending end through the ground feedback link; the problem of difficult convergence of Q values in the Q table caused by the overly large state space in traditional reinforcement learning is avoided through the neural network method. The ground sending end inputs the signal-to-noise ratio state received and calculated based on the ground receiving end signal into the evaluation network and uses the greedy algorithm to judge whether to explore with a corresponding probability. If so, a random action is selected from the action space; if not, the action that can obtain the maximum Q value is selected from the output of the evaluation network; Step (3), the sending end translates the selected action into the corresponding modulation method and coding rate according to the selected action, and sends the signal according to this modulation and coding method. After passing through the satellite uplink channel, the specified modulation and coding method is notified to the satellite for use in the next transmission; after the signal reaches the receiving end, the corresponding reward value is calculated, and the current state, action, and reward elements are stored as a group of samples in the experience pool; Step (4), judge whether the current iteration number is an integer multiple of the preset network update step number. If so, enter Step (5) for network update; if not, update the signal-to-noise ratio state and return to Step (2) to enter the next round of iteration; Step (5), introduce the concept of a dual network, that is, the output layers of both the evaluation network and the target network are split into a value layer and an advantage layer. The input of the value layer is only the signal-to-noise ratio state, which only focuses on the current channel quality; the learning effect is improved and the convergence of the result is accelerated by optimizing the structure of the neural network; the input of the advantage layer is the signal-to-noise ratio state and the modulation and coding method, which is responsible for focusing on the value of the modulation and coding method under the current signal-to-noise ratio state; finally, the output items of the two sub-networks are aggregated into the final Q value output; several groups of samples are extracted from the experience pool as the input of the training set; further judge whether the current iteration number is an integer multiple of the preset target network update step number. If not, directly train and update the evaluation network without updating the target network; If so, update the parameters of the evaluation network to the target network, and use the training set input and the result calculated by the formula as the new evaluation network. Up to this step, it is one round of reinforcement learning. Step (6): Update the signal-to-noise ratio state, increment the greedy parameter, and return to Step (2) for the next iteration.

2. A satellite adaptive modulation method based on deep reinforcement learning according to claim 1, characterized in that the calculation formula of the reward function is: Where M represents the modulation order of the current action, i.e., the modulation and coding scheme, snr represents the signal-to-noise ratio, and ber mcs (snr) represents the bit error rate of the input signal-to-noise ratio under the current modulation and coding scheme, and mcs represents the modulation and coding scheme.

3. A satellite adaptive modulation method based on deep reinforcement learning according to claim 1, characterized in that: In Step 5, two neural networks, namely the evaluation network and the target network, are used to replace the traditional Q-value index table, so as to be applicable to decision-making problems with an infinite state space; the iterative update formula of the evaluation network is as follows: Among which Q eval is the output of the evaluation network, Q target is the output of the target network, S is the current signal-to-noise ratio state, A is the current action, and γ is the learning rate; The final output is aggregated by the advantage layer and the value layer, and its aggregation formula is as follows: V(S; θ,β)+Advan(S,A; θ,α)→Q(S,A; θ,α,β) where V represents the output of the value layer, Advan represents the output of the advantage layer, θ is the parameter shared by the value layer and the advantage layer, α is the specific parameter of the advantage layer, and β is the specific parameter of the value layer.

4. A satellite adaptive modulation method based on deep reinforcement learning according to claim 1, characterized in that: During the iteration process, the parameters of the evaluation network are updated every time the loop is executed, while the parameters of the target network need to be updated at intervals to reduce the correlation of training samples and prevent overfitting.

Citation Information

Patent Citations

  • Satellite communication system and communication method optimized by adaptive code modulation mode

    CN109525299A

  • Reinforcement learning adaptive coding modulation method, system and device for satellite communication system

    CN114205053A

Cited By

  • Satellite communication anti-interference transmission system

    CN122119753A

  • Satellite communication anti-jamming transmission system

    CN122119753B