NOMA satellite-ground network double-end adaptive power distribution anti-interference method and system based on reinforcement learning

By employing a two-layer reinforcement learning approach involving space and ground, adaptive power allocation and interference suppression for NOMA systems in dynamic and complex environments were achieved, improving system performance and reliability and solving the problem of inaccurate power allocation strategies in traditional methods.

CN121530451APending Publication Date: 2026-02-13STATE GRID CORPORATION OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511701757.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional satellite-to-ground NOMA systems suffer from inaccurate power allocation strategies in dynamic and complex environments, leading to performance degradation. In particular, they struggle to guarantee the high reliability and efficiency of critical services in strong interference scenarios, and one-sided optimization cannot meet the requirements for real-time adaptability and interference suppression.

Method used

A reinforcement learning-based NOMA satellite-ground network dual-end adaptive power allocation method is adopted. Through collaborative decision-making between the satellite side and the terminal side, a two-layer intelligent decision architecture is established to achieve dynamic power adaptation and compound interference suppression. By combining global optimization on the satellite side and local interference suppression on the terminal side, the high-reliability backhaul of critical business data is guaranteed.

Benefits of technology

It significantly improved system throughput and communication reliability, optimized resource utilization efficiency, solved the problem of inaccurate power allocation in dynamic environments, and enhanced anti-interference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121530451A_ABST
    Figure CN121530451A_ABST
Patent Text Reader

Abstract

The invention discloses an NOMA satellite-ground network double-end adaptive power distribution anti-interference method and system based on reinforcement learning, and the method comprises the steps: building a satellite-ground network model and a channel model, and calculating the channel capacity of a ground user according to the channel model; establishing an optimization problem of a satellite side based on the satellite-ground network model and the channel model, modeling the optimization problem of the satellite side as a first Markov decision process, and solving the optimization problem based on a satellite side composite reward function by using a first reinforcement learning algorithm to obtain an optimal power distribution coefficient; and the terminal side establishes a spectrum waterfall anti-interference model, the modeling is a second Markov decision process, and a second reinforcement learning algorithm is used to dynamically select different levels of transmitting power in the terminal side action space according to the current terminal state based on the terminal side composite reward function. According to the invention, through satellite and terminal double-end intelligent agent collaborative learning, joint dynamic optimization of power distribution and interference suppression is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of satellite communication, wireless communication and machine learning, and particularly relates to a NOMA satellite-ground network double-end adaptive power allocation anti-interference method and system based on reinforcement learning for dynamic severe satellite-ground link environment. BACKGROUND

[0002] In recent years, the traditional ground communication infrastructure has gradually shown limitations in coverage, bandwidth and latency, especially in remote rural areas, vast forests and desert mountain areas where it is difficult to deploy base stations. This problem is more prominent. Under this background, the LEO (Low Earth Orbit) satellite-ground communication system has become a key solution to bridge the digital divide and support communication needs in underdeveloped areas due to its low latency, high bandwidth and global coverage. NOMA (Non Othogonal Multiple Access) technology achieves spectrum sharing through power domain signal multiplexing, significantly improving spectrum efficiency in resource-constrained environments, and has been widely used in satellite-ground networks. In infrastructure-poor scenarios such as power tower inclination monitoring, farmland IoT monitoring or emergency communication in remote areas, NOMA can serve users with different channel conditions simultaneously, effectively solving the problem of tight satellite link resources.

[0003] However, the unique transmission environment of satellite-ground links poses a serious challenge to their performance. On the one hand, the signal propagation experiences significant large-scale attenuation and complex small-scale fading, greatly affecting the reliability and stability of the link; on the other hand, the tight spectrum resource makes the CCI (Co-channel Interference) problem more prominent, and the spatial distribution characteristics of the interference source pose a serious threat to the received signal quality. In addition, in the power IoT scenario, satellite user terminals are deployed on towers, transformers, insulators and other devices along the power transmission and distribution lines to realize remote monitoring of the power transmission and distribution scenario using satellite-ground links, which has become a key technology path for power system state perception. However, the strong electromagnetic environment inherent in high-voltage power transmission and distribution systems produces wideband interference noise, significantly raising the noise floor of the receiver, resulting in the inability to timely return the sensing data in the power transmission and distribution environment, directly affecting the fault warning and operation and maintenance decision-making efficiency of the power transmission and distribution line.

[0004] As an effective solution to improve the capacity of satellite-to-ground communication systems, the performance of NOMA technology depends largely on the high-precision power allocation strategy, which is particularly significant in a dynamic satellite environment. On the one hand, the satellite elevation angle is a core variable that determines the quality of the satellite-to-ground link, which directly affects the propagation path length and indirectly relates to the ability to penetrate the atmosphere and avoid obstacles, resulting in significant differences in channel conditions for users at different elevations. On the other hand, the mobility of LEO satellites introduces dynamic channel conditions, with channel gain and the relative geometric position of users and satellites constantly changing. Traditional static power allocation strategies cannot adapt to this dynamic nature, which can easily lead to system performance degradation.

[0005] Existing power allocation strategies for satellite-to-ground NOMA systems typically use static strategies based on historical average channel CSI (Channel State Information) or predefined rules for power allocation. However, the inherent multiple dynamic characteristics of satellite-to-ground links cause the performance of static strategies to decline significantly in real environments, leading to inaccurate power allocation, reduced system throughput, and increased interruption probability, especially in scenarios such as power transmission and distribution with strong interference and high reliability requirements, making it difficult to guarantee critical services. In addition, focusing on single-point optimization on the satellite side ignores the potential of ground terminals in local spectrum sensing and adaptive interference suppression. The satellite side cannot obtain accurate dynamic interference information at the terminal, and the terminal also cannot use local information for rapid adaptive adjustment to compensate for the lag or inaccuracy of satellite-side decisions, significantly weakening the overall system's interference suppression capability.

[0006] Therefore, in the satellite-to-ground network, facing the multiple challenges of satellite high-speed movement, dynamic changes in elevation angle, and complex ground interference, traditional single-sided static optimization strategies cannot meet the core needs of real-time adaptability of power allocation, precision of interference suppression, and satellite-to-ground collaborative closed-loop optimization. SUMMARY

[0007] To address the deficiencies in the prior art, the present application provides a NOMA satellite-to-ground network double-end adaptive power allocation anti-interference method and system based on reinforcement learning, which combines global power optimization on the satellite side with local interference suppression on the terminal side by designing a satellite-to-ground collaborative double-layer intelligent decision-making architecture, synchronously solving the problems of dynamic power adaptation and composite interference suppression, and ensuring the high-reliability backhaul of critical business data such as power transmission and distribution monitoring. In a complex and dynamic NOMA satellite-to-ground communication system and an environment with power transmission and distribution electromagnetic interference and co-channel interference, the method ensures user power allocation fairness, achieves efficient resource utilization, improves system communication reliability, and optimizes system performance.

[0008] The technical scheme adopted by the present application is as follows.

[0009] The first aspect of the application provides a NOMA satellite-to-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning, comprising: A NOMA satellite-to-ground network model is established, the NOMA satellite-to-ground network model comprising one low earth orbit satellite, two ground users with different channel states and multiple access points, the multiple access points being used to generate same-frequency interference to the ground users; the satellite side and the terminal side cooperatively implement respective adaptive power allocation under different intensity of interference, specifically: A channel model corresponding to the link between the low earth orbit satellite and the ground user is established, and the channel capacity of the ground user is calculated according to the channel model; Based on the NOMA satellite-to-ground network model and the channel model, an optimization problem of the satellite side is established, the optimization problem being to dynamically adjust the power allocation coefficients corresponding to different ground users with the maximum total throughput of the two different channels as the target; The optimization problem of the satellite side is modeled as a first Markov decision process, and a satellite side state space, a satellite side action space and a satellite side composite reward function are set; wherein the state space comprises the instantaneous channel gain of different ground users, the elevation angle of the satellite, the instantaneous rate of the user, and the interference information reported by the terminal side user; The first reinforcement learning algorithm is used to solve the optimization problem based on the satellite side composite reward function, to obtain the optimal power allocation coefficient, and the satellite side allocates the power allocation coefficient to the two ground users with different channel states to realize the satellite side adaptive power allocation; The terminal side establishes a spectrum waterfall anti-interference model, which is modeled as a second Markov decision process, and a terminal side state space, a terminal side action space and a terminal side composite reward function are set, the terminal state space comprising real-time spectrum information, a strategic guidance signal issued by the satellite and a received signal interference noise ratio of the user; The second reinforcement learning algorithm is used to dynamically select different levels of transmit power in the terminal side action space based on the terminal side composite reward function according to the current terminal state to correct the guiding power corresponding to the strategic guidance signal issued by the satellite, to realize the terminal side adaptive power allocation.

[0010] Optionally, the satellite side composite reward function is used to guide the satellite agent to maximize the total throughput of the two different channels under the constraints of power allocation fairness, energy consumption and minimum rate.

[0011] Optionally, the satellite side composite reward function is represented as follows:

[0012] wherein, is the satellite side composite reward function, is the user is the throughput of the user at time . for users at time throughput, denotes the instantaneous rate of user at time throughput, denotes the instantaneous rate of user at time throughput, denotes the total transmit power of the satellite, denotes the maximum value of the transmit power of the satellite, denotes the minimum rate of user i, denotes the rate of user i; is a fairness penalty term for power allocation, is a weight coefficient of the fairness penalty term, which penalizes the excessive rate difference between users; is an energy consumption constraint penalty term, is a weight coefficient of the energy consumption constraint penalty term, which imposes a penalty when the total transmit power exceeds the maximum value ; is a minimum rate constraint penalty term, is a weight coefficient of the minimum rate constraint penalty term, which imposes a penalty when the instantaneous rate of any user is lower than its minimum required rate , wherein . Optionally, the terminal-side composite reward function is used to guide the terminal agent to learn an anti-interference strategy based on communication reliability, energy efficiency and policy stability, so that the terminal state dynamically selects different levels of transmit power.

[0013] Optionally, the terminal-side composite reward function is as follows:

[0014] wherein, denotes the terminal-side composite reward function, denotes the received signal-to-interference signal-to-noise ratio at the current time t; denotes the demodulation threshold; denotes the normalized value of the actual transmit power of the terminal at time , which is used to represent the energy efficiency reward, to penalize high power consumption behavior and encourage energy saving; denotes the level of the uplink transmit power selected by the terminal at time t; denotes the level of the uplink transmit power selected by the terminal at time t-1; is an indicator function, which is used to represent the communication reliability reward, and if the received signal to interference signal noise ratio of the terminal at the current time reaches the demodulation threshold , the value is 1, successful communication, otherwise 0, communication failure, which is used to directly stimulate the terminal agent to maintain reliable connection; is an indicator function, which is used to represent the strategy stability reward, and if the current action is different from the last action, that is, the power level switching occurs, the value is 1. is a weight coefficient of the communication reliability reward, is a weight coefficient of the energy efficiency reward, is a weight coefficient of the strategy stability reward.

[0015] Optionally, the satellite side state space is represented as follows:

[0016] wherein, represents the system state observed by the satellite agent at the decision time , and represents the instantaneous channel gain of the user , and represents the instantaneous channel gain of the user , and represents the elevation angle of the satellite at the decision time , and represents the instantaneous speed of the user at the decision time , and represents the instantaneous speed of the user at the decision time , and represents the interference environment feature vector reported by the user , and represents the interference environment feature vector reported by the user , and the interference feature vector is used to represent the interference environment of the user at the location.

[0017] Optionally, the interference environment feature vector reported by the user is represented as follows:

[0018] wherein, represents the average interference power perceived by the terminal user in the communication bandwidth, which is used to quantify the interference intensity, is an interference type feature vector, wherein, represents the existence probability of noise interference, represents the existence probability of pulse interference, represents the existence probability of sweep interference, represent the timing characteristics of the interference.

[0019] Optionally, the satellite-side action space is expressed as:

[0020] wherein, is a continuous scalar representing the change in the power coefficient allocated to the user , and are discrete strategic guidance signals, issued to the user through a downlink control channel, issued to the user through a downlink control channel, wherein the strategic guidance signal is a user-reported interference environment feature vector setting used to guide the terminal-side to select a transmit power level matching the strategic guidance signal.

[0021] Optionally, the terminal state is expressed as:

[0022] wherein, represents a terminal state vector, is a normalized spectral waterfall matrix representing real-time spectral information, the spectral waterfall matrix being composed of a power spectral density vector, represents a strategic guidance signal issued by the satellite to the terminal at the current time t, represents the received signal-to-interference signal-to-noise ratio of the user at time t-1.

[0023] Optionally, the second reinforcement learning algorithm is a deep Q-network-based reinforcement learning algorithm, and a preset lightweight recurrent convolutional neural network is used to approximate the Q function.

[0024] Optionally, the preset lightweight recurrent convolutional neural network structure includes: an input layer for receiving a spectral waterfall matrix; a recurrent convolutional layer composed of one-dimensional or two-dimensional convolutional kernels, which uses a recurrent calculation method to extract local features of the spectrum in the frequency dimension and the time dimension, and flattens the output of the recurrent convolutional layer into a one-dimensional vector; a recurrent fully connected layer for mapping the one-dimensional vector extracted by the recurrent convolutional layer to the Q value corresponding to each action; an output layer with an output dimension consistent with the size of the terminal action space, each output node corresponding to the Q value of an action.

[0025] The second aspect of the application provides a NOMA satellite-to-ground network double-end adaptive power allocation anti-interference system based on reinforcement learning, which is used to implement the NOMA satellite-to-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning. The first establishing module is configured to establish a NOMA satellite-to-ground network model, the NOMA satellite-to-ground network model comprising one low earth orbit satellite, two ground users with different channel states, and a plurality of access points, the plurality of access points being configured to generate co-channel interference to the ground users; based on the first establishing module, the satellite-side adaptive power allocation module and the terminal-side adaptive power correction module are configured to cooperatively implement respective adaptive power allocation under different intensities of interference; wherein, The satellite-side adaptive power allocation module comprises: The second establishing module is configured to establish a channel model corresponding to a link between the low earth orbit satellite and the ground users, and calculate channel capacity of the ground users according to the channel model. The third establishing module is configured to establish an optimization problem at the satellite side based on the NOMA satellite-to-ground network model and the channel model, the optimization problem being configured to dynamically adjust power allocation coefficients corresponding to different ground users with a target of maximizing total throughput of the two different channels. The power allocation module is configured to model the optimization problem at the satellite side into a first Markov decision process, and set a satellite-side state space, a satellite-side action space, and a satellite-side composite reward function, wherein the state space comprises instantaneous channel gains of different ground users, an elevation angle of the satellite, an instantaneous rate of the user, and interference information reported by the terminal-side user; the power allocation module is configured to solve the optimization problem based on the satellite-side composite reward function using a first reinforcement learning algorithm, and obtain optimal power allocation coefficients. The terminal-side adaptive power correction module comprises: The first establishing module is configured to establish a spectrum waterfall anti-interference model at the terminal side, model the spectrum waterfall anti-interference model into a second Markov decision process, and set a terminal-side state space, a terminal-side action space, and a terminal-side composite reward function, wherein the terminal state space comprises real-time spectrum information, a strategic guidance signal issued by the satellite, and a received signal interference noise ratio of the user. The power correction module is configured to dynamically select different levels of transmission power in the terminal-side action space based on the terminal-side composite reward function according to a current terminal state using a second reinforcement learning algorithm, to correct a guiding transmission power corresponding to the strategic guidance signal issued by the satellite, and implement terminal-side adaptive power allocation.

[0026] The third aspect of the application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being loaded into the processor to implement the NOMA satellite-to-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning.

[0027] The fourth aspect of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program is executed by a processor to realize the NOMA satellite-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning.

[0028] Compared with the prior art, the beneficial effects of the present application at least include: The present application overcomes the problem of serious performance decline of static power allocation strategy and single-sided optimization difficult to cope with composite interference in the prior art dynamic harsh environment of satellite-ground link, and provides a satellite-ground cooperative double-layer reinforcement learning anti-interference method. Through satellite and terminal double-end agent cooperative learning, joint dynamic optimization of power allocation and interference suppression is realized, and system throughput and communication reliability are significantly improved. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them: Figure 1 is a flowchart of a NOMA satellite-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning provided by an embodiment of the present application; Figure 2 is another flowchart of a NOMA satellite-ground network double-end adaptive power allocation anti-interference method based on reinforcement learning provided by an embodiment of the present application; Figure 3 is a schematic diagram of a constructed satellite-ground network model provided by an embodiment of the present application; Figure 4 is a schematic diagram of a satellite-side optimized DDPG training process provided by an embodiment of the present application; Figure 5 is a schematic diagram of a terminal-side DQN training process provided by an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical scheme and advantages of the present application more clear, the technical scheme of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. The embodiments described in the present application are only some embodiments of the present application, not all embodiments. Based on the spirit of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0031] In combination with Figure 1 andFigure 2 As shown, Embodiment 1 of this application provides a reinforcement learning-based dual-end adaptive power allocation anti-interference method for NOMA satellite-to-ground networks, comprising the following steps: Step 1: Establish a satellite-to-ground network model, defining a low Earth orbit satellite, two ground users with different channel states, and defining several access points as interference models.

[0032] Combination Figure 3 As shown, in step 1, the low Earth orbit satellite in the satellite-to-ground network model utilizes non-orthogonal multiple access (NOMA) technology to simultaneously serve two ground users in different channel conditions. The satellite allocates different transmit powers to the two users based on their channel conditions. Furthermore, multiple access points deployed near the users can cause co-channel interference, as these access points communicate with the satellite on the same frequency band, thus interfering with signal transmission from the satellite to the users.

[0033] The specific model establishment of the satellite-ground network is as follows: Define a low Earth orbit satellite Ground users ,in This indicates a user with good channel conditions and stable reception quality. This indicates users with poor channel conditions and unstable reception quality. (Two ground users) and The midpoint is defined as Low Earth orbit satellites and The distance between the links is calculated using the following formula: (1) in, For low Earth orbit satellites and The distance between the links, For Earth's radius, The distance from a low Earth orbit satellite to the Earth's surface is called the satellite's altitude. for The position observes the satellite's elevation angle. , Maximum elevation angle, This is the minimum elevation angle.

[0034] According to the law of cosines, the distance between a low Earth orbit satellite and two ground users can be calculated, expressed as: (2) in, The distance between the satellite and the ground terminal user. for and The horizontal distance between them For low Earth orbit satellites and The distance between the links, α for The position is used to observe the satellite's elevation angle.

[0035] Step 2: Establish a channel model, using the Nakagami-m fading model to characterize the link between low Earth orbit satellites and users, and define the channel capacity for the two ground users.

[0036] It should be noted that the Nakagami-m fading model is a statistical model used to describe small-scale fading in wireless channels, and it is an existing technology.

[0037] In step 2, low Earth orbit satellite ground communication is affected by both large-scale attenuation and small-scale fading, as well as co-channel interference on the satellite-to-ground network, satellite elevation angle on the NOMA satellite-to-ground system, dynamic channel conditions caused by the high-speed movement of LEO satellites on the performance of the satellite-to-ground network, and electromagnetic interference from the power transmission and distribution environment.

[0038] Step 2.1: Considering the simultaneous existence of LOS (Line of Sight) and NLOS (Non-Line of Sight) components in the satellite-to-ground link, the Nakagami-m fading model is used to characterize the multipath fading characteristics in the satellite-to-ground model. Ground users and The emitted signals are defined as follows: and Low Earth orbit satellites and ground users The channel coefficients of the links between them are expressed by the following formula:

[0039] in, For low Earth orbit satellites and ground users Channel coefficients of the link between them This is the path loss factor. For antenna gain of low Earth orbit satellites, For ground users Antenna gain, For low Earth orbit satellites and ground users The channel fading factor of the link between them.

[0040] Specifically, the channel fading factor follows the following gamma distribution:

[0041] The channel fading factor satisfies the shape parameter as follows The scale parameter is The gamma distribution, when represents the severity of the fading.

[0042] The path loss factor is calculated as:

[0043] wherein, is the speed of light, is the carrier frequency, is the distance between the satellite and the ground terminal user.

[0044] Step 2.2: The received signal at the ground user may be represented as: (3) wherein, represents the received signal at the ground user , represents the power coefficient allocated to the user , represents the power coefficient allocated to the user , satisfying the conditions and , represents the transmission power of the low earth orbit satellite, represents the transmission power of the access point, represents the number of access points that interfere with the user , represents the number of access points that interfere with the user , represents the additive white Gaussian noise at the user , and satisfies , represents the noise variance, represents the average power of the noise, represents the channel coefficient of the link between the th interfering access point and the user with a better channel state, represents the channel coefficient of the link between the th interfering access point and the user with a worse channel state. The parameters , are used to describe the signal amplitude variation caused by small-scale fading in the signal transmission process from the access point to the user.

[0045] Step 2.3: From equation (3), the received signal-to-interference-and-noise ratio at the users and may be represented as: (4) (5) wherein, is the received signal to interference noise ratio at the user , is the received signal to interference noise ratio at the user , denotes the channel coefficient of the link between the low earth orbit satellite and the terrestrial user , denotes the average power of the noise representing the user , denotes the number of access points causing interference to the user , denotes the number of access points causing interference to the user , denotes the channel coefficient of the link between the low earth orbit satellite and the terrestrial user , denotes the average power of the noise representing the user .

[0046] denotes the co-channel interference experienced by the user , denotes the co-channel interference experienced by the user . Due to the presence of dense shadowing in the terrestrial network, the line of sight component of the signal almost vanishes. Therefore, it is proposed in the present application that the fading of the link between the terrestrial user and the access point follows the Rayleigh model.

[0047] Step 2.4: The instantaneous channel capacity of the user and the user may be represented as: (6) (7) wherein, is the instantaneous channel capacity of the user , is the instantaneous channel capacity of the user .

[0048] Step 3: Establish an optimization problem to maximize the total throughput of the two different channels while ensuring the constraints of power allocation fairness, energy consumption constraints and minimum rate constraints to adjust the power allocation coefficients dynamically.

[0049] The step 3 optimization problem can be specifically described as maximizing the total throughput of two terrestrial users in a given satellite-terrestrial communication model, i.e. the user Users with poor channel conditions The sum of the throughputs of users with poor channel conditions, using the channel capacity formula above: (8) C1:

[0050] C2:

[0051] C3:

[0052] C4:

[0053] where, denotes the total throughput, B denotes the bandwidth of the communication channel, denotes the total transmit power of the satellite, denotes the maximum value of the transmit power of the satellite, denotes the minimum rate requirement of the user , denotes the minimum rate requirement of the user .

[0054] C1 denotes the fairness constraint, which ensures that the difference in user throughput does not exceed the tolerance threshold ; C2 denotes the energy consumption constraint, which ensures that the total transmit power of the satellite does not exceed the maximum value; C3 denotes the minimum rate constraint, which ensures that the throughput of the user cannot be lower than the respective guaranteed rate; and C4 denotes the NOMA power allocation principle, which ensures that users with good channel conditions are allocated less power.

[0055] Step 4: Establish the optimization problem on the satellite side as a Markov decision process, and set up the state space, action space, and reward function respectively.

[0056] In step 4, to achieve dynamic optimization of the power allocation coefficient on the satellite side, the optimization problem on the satellite side is modeled as a Markov decision process, which consists of three core elements: state space, action space, and reward function.

[0057] Step 4.1: Create a state space, use the Nakagami-m fading model to generate the channel gain between the low Earth orbit satellite and the user, and at each decision time , the system state observed by the satellite agent contains the following key information: the instantaneous channel gain of the user , the instantaneous channel gain of the user , the current satellite elevation angle , and the instantaneous rate of the user at the decision time . ​​,user Decision-making moment instantaneous rate and by ground terminal and The real-time interference feature vectors reported via a dedicated feedback channel accurately describe the interference environment at the terminal's location, enabling the satellite to perceive dynamic ground interference and make collaborative decisions. Therefore, the state vector is defined as: (9) in, Indicates the moment of decision-making The system state observed by the satellite intelligent agent, Indicates user Instantaneous channel gain, Indicates user Instantaneous channel gain, Indicates the moment of decision Satellite elevation angle, Indicates user Decision-making moment instantaneous rate, Indicates user Decision-making moment The instantaneous rate. Indicates user The reported feature vector of the interfering environment, Indicates user The reported interference environment feature vector is used to characterize the interference environment at the user's location.

[0058] Specifically, the feature vector of the interference environment reported by the user is represented by the following formula:

[0059] in, This represents the average interference power perceived by end users within the communication bandwidth, measured in dBm, and is used to quantify interference intensity. It is a feature vector of the interference type. ,in This indicates the probability of the presence of noise interference. This indicates the probability of the presence of impulse interference. The probability of the presence of frequency sweep interference can be generated by a lightweight online classifier; It represents the timing characteristics of the interference; for pulse interference, it captures the duty cycle; for sweep interference, it captures the sweep period.

[0060] Step 4.2: Create the action space. The action performed by the agent is the adjustment of the power allocation coefficient. Define the action space as follows: ,in, is a continuous scalar, representing the power allocation coefficient for user at time t. is the change of the power allocation coefficient for user , and is set in the range of , after the action is performed, the power allocation coefficient for user at time t is updated to , and satisfies . Wherein, represents the power allocation coefficient for user before the action is performed, represents the power allocation coefficient for user after the action is performed, represents the power allocation coefficient for user after the action is performed.

[0061] and are discrete strategic guidance signals, taking values in the range of {0, 1, 2}, where 0 corresponds to low level of transmit power, 1 corresponds to medium level of transmit power, and 2 corresponds to high level of transmit power, which are respectively sent to terminal and through the downlink control channel. The strategic guidance signal is generated based on the satellite's evaluation of the global interference state, i.e., the interference environment features reported by the user, and is used to guide the terminal side to select the transmit power level that matches it, realizing the linkage and cooperation of satellite-ground strategy for anti-jamming.

[0062] Step 4.3: Set the reward function to guide the agent to maximize the total system throughput under the premise of meeting all constraint conditions, and the reward function is represented as follows: (10) Wherein, is the reward function, is the throughput of user at time t, is the throughput of user at time t, is the total system throughput term, which directly encourages the target maximization; represents the instantaneous rate of user at the decision-making moment , represents the instantaneous rate of user at the decision-making moment , is the active power allocation fairness penalty term, is the weight coefficient of the fairness penalty term, which penalizes the excessive rate difference between users and promotes fairness; As an energy consumption constraint penalty term, The weighting coefficients for the power constraint penalty term are applied when the total transmit power exceeds the maximum value. Punishment should be imposed at the appropriate time; This is a minimum rate constraint penalty term. The weighting coefficients for the constraint penalty term; when any user's instantaneous rate is lower than its minimum required rate. Punishment is imposed at the time, among which, Weighting coefficient , , It is used to balance the importance of throughput, fairness, and constraint violations.

[0063] Step 5: Use reinforcement learning algorithms to solve the optimization problem on the satellite side.

[0064] Combination Figure 4 As shown, in step 5, a deep reinforcement learning algorithm based on the actor-critic framework is used, which utilizes DDPG (deep deterministic policy gradient) to learn the optimal power allocation policy on the satellite side.

[0065] The DDPG algorithm comprises four neural networks: an actor network (Actor), a critic network (Critic), and their respective target networks: the target actor network and the target critic network. The actor network generates actions, while the critic network evaluates the actor network's performance.

[0066] The training of the satellite-side DDPG algorithm aims to collaboratively optimize the actor network and critic network through interactive experience data, ultimately enabling the actor network to output an action policy that maximizes long-term cumulative rewards. The specific training process includes the following steps: Since the actor network learns a deterministic policy, in order to fully explore the environment in the early stages of training, noise needs to be explicitly added to the actions output by the agent. This application uses Ornstein-Uhlenbeck process noise, so the actual actions executed by the agent are: (11) in, The action actually performed by the agent at time t. This represents the system state observed by the satellite agent at decision time t. The parameter weights represent the actor network parameters. The actor network indicates the current status. and parameters The calculated original motion output, denotes the Ornstein-Uhlenbeck process noise.

[0067] The magnitude of the noise is gradually decayed as the training proceeds, so that the policy gradually transitions from exploration-dominant to exploitation-dominant.

[0068] At each time step t, the satellite agent interacts with the environment and obtains an experience tuple:

[0069] wherein, denotes the experience tuple obtained by the satellite agent after interacting with the environment at time t, is the action actually performed by the agent at time t, denotes the system state observed by the satellite agent at time t after performing the action is the reward function, denotes the system state observed by the satellite agent at time t after performing the action at time t.

[0070] The experience tuples are stored in an experience replay buffer of fixed size When training the network, a small batch of experiences is uniformly sampled from the buffer at random. For each experience sampled from the experience replay buffer, the target Q value is calculated as: (12) wherein, denotes the target Q value corresponding to experience i, which is calculated from the expected cumulative reward of the current experience, denotes the immediate reward at time t, denotes the state at the next time step, denotes the target action calculated by the target actor network according to the state at the next time step, denotes the parameters of the target actor network, denotes the parameters of the target critic network, denotes the Q value estimate of the target critic network for the state-action pair. The satellite agent selects the corresponding action according to the size of the target Q value, is the discount factor, which is used to weigh the importance of the immediate reward and the long-term reward. The loss function of the critic network is defined as the mean squared error between the target value and the predicted value:

[0071] ​​ (13) wherein, represents the loss function of the critic network, represents the batch size, represents the target Q value corresponding to experience i, represents the predicted Q value corresponding to experience i.

[0072] In some embodiments, the loss function can be minimized by using a stochastic gradient descent method to update the parameters of the online critic network .

[0073] The update of the parameters of the online actor network is achieved by maximizing the performance of the strategy evaluated by the critic network, and the performance objective function is defined as: (14) wherein, represents the performance objective function, represents the parameters of the online actor network, and D represents the experience sample distribution sampled from the experience replay buffer X, represents the expectation of the predicted Q value of the target critic for the experience sample sampled from the experience replay buffer X, and the parameters of the online actor network are obtained by maximizing the performance objective function .

[0074] By the chain rule, the gradient of the performance objective with respect to the actor network parameters, i.e. the strategy gradient, can be calculated as: (15) In some embodiments, the parameters of the online actor network can be updated along this gradient direction by using a gradient ascent method .

[0075] The target network provides a stable supervisory signal for the learning process, and the network parameters change slowly through soft updating, i.e. after each iteration, the target network parameters approach the online network parameters according to the following rule: (16) wherein, is a soft update coefficient, which is a constant much smaller than 1.

[0076] When the training is completed, the satellite side optimization algorithm enters the execution phase. At this time, after all the training is completed, only the state needs to be input to the trained neural network model for forward propagation, and there is no need to explore the action through noise again, i.e. the exploration noise is removed. The satellite agent observes the system state in real time, inputs it into the trained online actor network, directly outputs the optimal power allocation coefficient adjustment action and executes it, so as to realize the high-precision, self-adaptive dynamic optimization of the satellite-ground link power allocation.

[0077] Step 6: The terminal side establishes a spectrum waterfall anti-interference model, which is modeled as a Markov decision process, and the power spectral density (PSD) is used to represent the spectrum state.

[0078] In step 6, the terminal side aims to solve the problem of dynamic interference suppression of ground user terminals in a complex electromagnetic environment. The terminal obtains real-time spectrum information through local spectrum sensing and uses reinforcement learning algorithms to make autonomous decisions based on this information to achieve adaptive adjustment of transmission power and interference avoidance. Specifically, the terminal side inputs the pre-processed spectrum waterfall graph, strategic guidance signals from satellites, and historical communication quality data into a reinforcement learning model based on deep Q-network to comprehensively consider the reward function of communication reliability, energy efficiency, and policy stability to guide training. Finally, the optimal discrete transmission power level decision is output to achieve adaptive adjustment of transmission power and interference avoidance. Step 6 specifically includes: Step 6.1: At each decision-making time , the terminal agent scans the entire communication frequency band through its built-in spectrum sensing module to obtain the power spectral density vector at the current time:

[0079] where is the number of subcarriers divided in the frequency band, represents the received signal power at the center frequency of the th subcarrier, with units of dBm.

[0080] To capture the time-domain dynamic characteristics of interference, the terminal agent will retain the spectrum vectors of the last time points to form a spectrum waterfall matrix as its observed environment state : (17) where represents the environment state observed by the terminal agent in the last time points, represents the environment state observed by the terminal agent at time t. This matrix is a two-dimensional matrix of , and its heat map visually displays the power intensity of different frequencies at different times, containing the complete characteristics of interference in the time and frequency domains.

[0081] Further, the complete terminal state vector is defined as: , where is the terminal state vector, is the spectrum waterfall matrix after normalization and noise filtering, is the strategic guidance signal sent by the satellite to the terminal at the current time t, which makes the terminal learn the action selection of the satellite as the state of the terminal, and combines the spectrum to guide the anti-interference behavior of the terminal, is the signal-to-noise ratio of the terminal at the last time.

[0082] In this embodiment, the terminal state space is designed by the present application, and the complete state of the terminal intelligent agent contains the processed spectrum waterfall information, the strategic guidance signal from the satellite and the local communication quality information. The strategic guidance from the satellite is considered, so that the terminal side and the satellite side can work cooperatively to realize adaptive power allocation anti-interference. In the complex dynamic NOMA satellite-ground communication system and the environment with power distribution electromagnetic interference and co-channel interference, the user power allocation fairness is guaranteed, and the resource efficient utilization is realized, and the system throughput and communication reliability are significantly improved.

[0083] Step 6.2: Action of the terminal intelligent agent is defined as the adjustment strategy of the uplink transmission power. The action space is a discrete set, which contains multiple alternative power level selections: (18) wherein, denotes the set of alternative power levels, denotes the first level transmission power, denotes the second level transmission power, denotes the third level transmission power, wherein, < < .

[0084] The terminal dynamically selects the transmission power of the low, medium and high levels according to the current interference condition. When the interference is weak, the high power is selected to improve the transmission rate, when the interference is strong, the low power is selected to avoid conflict or save energy, and when the interference is moderate, the compromise power is selected.

[0085] Step 6.3: Reward function of the terminal side is used to guide the intelligent agent to learn an effective anti-interference strategy, which needs to consider the communication quality, energy efficiency and strategy stability at the same time, and is defined as follows: (19) wherein, denotes the composite reward function of the terminal side, denotes the received signal-to-interference signal-to-noise ratio at the current time t; denotes the demodulation threshold; denotes the terminal at time the normalized value of the actual transmit power of the terminal, used to represent the energy efficiency reward, which is used to punish high power consumption behavior and encourage energy saving; represents the level selection strategy of the uplink transmit power of the terminal at time t; represents the level selection strategy of the uplink transmit power of the terminal at time t-1; is an indicator function, used to represent the communication reliability reward, if the received signal-to-interference signal-to-noise ratio of the terminal at the current time reaches the demodulation threshold , the value is 1, i.e. successful communication, otherwise 0, i.e. communication failure, which is used to directly encourage the terminal agent to maintain reliable connection; is an indicator function, used to represent the strategy stability reward, if the current action is different from the last action, i.e. power level switching occurs, the value is 1, which is used to punish too frequent strategy switching, maintain the relative stability of the communication link, and avoid unnecessary switching overhead; is the weight coefficient of the communication reliability reward, is the weight coefficient of the energy efficiency reward, is the weight coefficient of the strategy stability reward, used to balance the importance of communication reliability, energy efficiency and strategy stability.

[0086] Step 7: using reinforcement learning to process the anti-interference problem of the terminal side.

[0087] As shown in Figure 5 , in step 7, the terminal side uses a deep Q network-based reinforcement learning algorithm to solve its anti-interference decision problem. A lightweight convolutional neural network is designed as a function approximator of the Q network for the high-dimensional state input of the spectrum waterfall diagram.

[0088] Step 7.1: preprocessing of network input, the environment state observed by the terminal agent is the spectrum waterfall diagram matrix . In order to improve learning efficiency and stability, the state is preprocessed before inputting into the neural network. First, normalize each element of the matrix, linearly transform it to the interval [0, 1]. Then, add preprocessing, set a noise threshold , and set the spectrum points with normalized values below to zero to filter out background noise and reduce unnecessary calculation. The processed state matrix is denoted as .

[0089] Step 7.2: Network structure design. Considering the limited computing resources of terminal devices and the spatio-temporal characteristics of the spectrum waterfall diagram, a lightweight RCNN (Regional Convolutional Neural Network) structure is designed for approximating the Q function wherein is the network weight parameter, which is specifically optimized for the extraction of spectral features, and includes the following layers: Input layer, used to receive the pre-processed state matrix , which is treated as a single-channel image.

[0090] Recursive convolution layer, composed of one-dimensional or two-dimensional convolution kernels, used to automatically extract local features of the spectrum in the frequency and time dimensions. Considering the time correlation of the state sequence:

[0091] wherein is the new matrix without the oldest frame, and the recursive convolution layer uses recursive calculation. Let the input, output, and weight of the th recursive convolution layer at time be , respectively, then the recursive calculation formula is: (20) wherein represents the output of the th recursive convolution layer at time , represents the new original data frame input to this layer at time , represents the weight of the th recursive convolution layer, represents the output of the th recursive convolution layer at time , represents the convolution operation, , and is the convolution kernel size. This recursive structure significantly reduces the computational complexity, which is about .

[0092] Recursive fully connected layer, used to map the flattened features extracted by the convolution layer to the Q value corresponding to each action, which is the value function of the action. Let the input, output, and weight of the recursive fully connected layer be , respectively, and the recursive update formula is: (21) wherein It reflects the change in spectral features between adjacent time steps. This structure only performs calculations when features change, further improving efficiency. The flattened feature vector, which is the input to the fully connected layer and the output of the convolutional layer, flattens the output of the convolutional layer into a one-dimensional vector.

[0093] Output layer, output dimension and action space The sizes are consistent, and each output node corresponds to one action. Q-value estimation.

[0094] Step 7.3 Training of the reinforcement learning algorithm based on the deep Q-network on the terminal side aims to optimize the network parameters by minimizing the temporal difference error. This allows the learner to acquire the optimal anti-interference strategy. The specific training process includes the following steps: In the early stages of training, the following methods were adopted: A greedy strategy is used to select actions with a certain probability. Explore by randomly selecting actions. The probability is used to select the action with the highest current Q value. As training progresses, Gradually decrease, increase the proportion of utilization.

[0095] The agent will generate experience tuples from each interaction. The experience is stored in a fixed-size experience replay buffer E. During training, a small batch of experiences is randomly sampled from buffer E to break the correlation between data and improve learning stability.

[0096] Use a structure with the same parameters A slower-updating target network is used to compute the target value for Q-learning, mitigating the oscillation problem in Q-value estimation. The parameters of the target network are periodically copied from an online network. ,in, Indicates the target network parameters. This indicates online network parameters.

[0097] Based on the sampling experience, the target value for Q-learning is... The calculation is as follows: (twenty two) in, This is a discount factor used to balance immediate rewards with long-term returns. The loss function of an online network is defined as the mean squared error between the target value and the current predicted value: (twenty three) in, Describes the loss function of an online network. This represents the target value for Q-learning. This represents the current predicted value. MSE, target Q value generated by the target network and the predicted Q value generated by the online network are iterated constantly until convergence.

[0098] The gradient descent method is used to minimize the loss function to update the online network parameters The gradient calculation is as follows: (24) wherein, represents the gradient of the loss function.

[0099] Embodiment 2 of the present application provides a reinforcement learning-based anti-interference system for adaptive power allocation of a dual-end of a NOMA satellite-ground network, which runs the reinforcement learning-based anti-interference method for adaptive power allocation of a dual-end of a NOMA satellite-ground network as described in Embodiment 1, and the system comprises: A first establishing module is configured to establish a NOMA satellite-ground network model, wherein the NOMA satellite-ground network model comprises one low earth orbit satellite, two ground users with different channel states, and multiple access points, and the multiple access points are configured to generate co-channel interference to the ground users; based on the first establishing module, a satellite-side adaptive power allocation module and a terminal-side adaptive power correction module cooperatively implement adaptive power allocation of each under different intensity of interference; wherein, The satellite-side adaptive power allocation module comprises: A second establishing module is configured to establish a channel model corresponding to a link between a low earth orbit satellite and a ground user, and calculate a channel capacity of the ground user according to the channel model; A third establishing module is configured to establish an optimization problem of the satellite side based on the NOMA satellite-ground network model and the channel model, wherein the optimization problem is to dynamically adjust power allocation coefficients corresponding to different ground users with the objective of maximizing the total throughput of the two different channels; A power allocation module is configured to model the optimization problem of the satellite side into a first Markov decision process, and set a satellite-side state space, a satellite-side action space, and a satellite-side compound reward function, wherein the state space comprises instantaneous channel gains of different ground users, an elevation angle of the satellite, an instantaneous rate of the user, and interference information reported by the terminal-side user; the optimization problem is solved based on the satellite-side compound reward function using a first reinforcement learning algorithm to obtain optimal power allocation coefficients; The terminal-side adaptive power correction module comprises: A first establishing module is configured to establish a spectrum waterfall anti-interference model on the terminal side, and model it into a second Markov decision process, and set a terminal-side state space, a terminal-side action space, and a terminal-side compound reward function, wherein the terminal state space comprises real-time spectrum information, a strategic guidance signal issued by the satellite, and a received signal interference noise ratio of the user; The power correction module is configured to dynamically select the transmission power of different levels in the terminal-side action space based on the terminal-side composite reward function according to the current terminal state by using the second reinforcement learning algorithm, so as to correct the guiding transmission power corresponding to the strategic guidance signal transmitted by the satellite, and realize the terminal-side adaptive power allocation.

[0100] As to the system in the above-mentioned embodiments, the specific manner in which each unit performs the operation has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0101] Embodiment 3 of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when loaded into the processor, implements the method for anti-interference of adaptive power allocation of both ends of a NOMA satellite-ground network based on reinforcement learning according to Embodiment 1.

[0102] Embodiment 4 of the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program, when executed by a processor, implements the method for anti-interference of adaptive power allocation of both ends of a NOMA satellite-ground network based on reinforcement learning according to Embodiment 1.

[0103] It should be understood that the size of the serial number of each step in the above-mentioned embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0104] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit the same, and although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific implementation of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A reinforcement learning-based dual-end adaptive power allocation anti-interference method for NOMA satellite-to-ground networks, characterized in that, include: A NOMA satellite-to-ground network model is established, comprising a low Earth orbit satellite, two ground users with different channel states, and multiple access points. These access points are used to generate co-channel interference to the ground users. The satellite and terminal sides collaborate to achieve adaptive power allocation under different interference intensities, specifically as follows: Establish a channel model corresponding to the link between low Earth orbit satellites and ground users, and calculate the channel capacity of ground users based on the channel model; Based on the NOMA satellite-ground network model and channel model, an optimization problem is established on the satellite side. The optimization problem is to dynamically adjust the power allocation coefficients corresponding to different ground users with the objective of maximizing the total throughput of two different channels. The optimization problem on the satellite side is modeled as a first Markov decision process, establishing a satellite-side state space, a satellite-side action space, and a satellite-side composite reward function. The state space includes the instantaneous channel gain of different ground users, the satellite elevation angle, the instantaneous rate of the users, and the interference information reported by the terminal users. The first reinforcement learning algorithm is used to solve the optimization problem based on the satellite-side composite reward function to obtain the optimal power allocation coefficient. The satellite side allocates the power to two ground users with different channel states according to the power allocation coefficient to achieve satellite-side adaptive power allocation. A spectrum waterfall anti-interference model is established on the terminal side, which is modeled as a second Markov decision process. A terminal-side state space, a terminal-side action space, and a terminal-side composite reward function are established. The terminal state space includes real-time spectrum information, strategic guidance signals transmitted by satellites, and the interference-to-noise ratio of the user's received signals. The second reinforcement learning algorithm uses a terminal-side composite reward function to dynamically select different levels of transmit power within the terminal-side action space based on the current terminal state to correct the guidance power corresponding to the strategic guidance signal issued by the satellite, thereby achieving adaptive power allocation on the terminal side.

2. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 1, characterized in that: The satellite-side composite reward function is used to guide the satellite agent to maximize the total throughput of two different channels under constraints of power allocation fairness, energy consumption, and minimum rate.

3. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 2, characterized in that: The satellite-side composite reward function is expressed as follows: in, For satellite-side composite reward function, For users At any moment throughput, For users At any moment throughput, Indicates user At any moment instantaneous rate, Indicates user At any moment instantaneous rate, This indicates the total transmission power of the satellite. This indicates the maximum value of the satellite's transmission power. This represents the minimum rate for user i. Represents the rate of user i; There is a penalty for fairness in power allocation. This is the weighting coefficient for the fairness penalty item, used to penalize excessive rate differences between users; As an energy consumption constraint penalty term, The weighting coefficient for the energy consumption constraint penalty term is applied when the total transmit power exceeds the maximum value. Punishment should be imposed at the appropriate time; This is a minimum rate constraint penalty term. The weighting coefficient for the minimum rate constraint penalty term; when any user's instantaneous rate is lower than its minimum required rate. Punishment is imposed at the time, among which, .

4. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 1, characterized in that: The terminal-side composite reward function is used to guide the terminal agent to learn anti-interference strategies based on communication reliability, energy efficiency, and policy stability, so that the current terminal state can dynamically select different levels of transmit power.

5. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 4, characterized in that: The terminal-side composite reward function is expressed as follows: in, This represents the composite reward function on the terminal side. This represents the signal-to-noise ratio (SNR) of the received signal at time t. Indicates the demodulation threshold; Indicates the time of the terminal The normalized value of the actual transmit power is used to characterize energy efficiency rewards, to penalize high-power consumption behavior, and to encourage energy conservation. This indicates the uplink transmit power level selected by the terminal at time t; This indicates the uplink transmit power level selected by the terminal at time t-1; It is an indicator function used to characterize the communication reliability reward, which is given if the signal-to-noise ratio (SNR) of the received signal at the terminal at the current moment reaches the demodulation threshold. If the value is 1, communication is successful; otherwise, it is 0, communication has failed. This is used to directly incentivize the terminal agent to maintain a reliable connection. It is an indicator function used to characterize the reward for policy stability. If the current action is different from the previous action and a power level switch has occurred, the value is 1. It is the weighting coefficient for communication reliability rewards. It is the weighting coefficient for energy efficiency rewards. It is the weighting coefficient for the strategy stability reward.

6. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 1, characterized in that: The satellite-side state space is represented as follows: in, Indicates the moment of decision-making The system state observed by the satellite intelligent agent, Indicates user Instantaneous channel gain, Indicates user Instantaneous channel gain, Indicates the moment of decision Satellite elevation angle, Indicates user Decision-making moment instantaneous rate, Indicates user Decision-making moment instantaneous rate, Indicates user The reported feature vector of the interfering environment, Indicates user The reported interference environment feature vector is used to characterize the interference environment at the user's location.

7. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 6, characterized in that: The feature vector of the interference environment reported by the user is represented by the following formula: in, This represents the average interference power perceived by end users within the communication bandwidth, used to quantify interference intensity. It is a feature vector of the interference type. ,in, This indicates the probability of the presence of noise interference. This indicates the probability of the presence of impulse interference. This indicates the probability of the presence of frequency sweep interference. This indicates the timing characteristics of the interference.

8. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 7, characterized in that: The satellite-side motion space is expressed as follows: in, It is a continuous scalar, representing the user. Power coefficient of allocation The change and These are discrete strategic guidance signals. Send to the user via the downlink control channel , Send to the user via the downlink control channel The strategic guidance signal is set based on the interference environment feature vector reported by the user, and is used to guide the terminal to select the transmission power level that matches the strategic guidance signal.

9. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 8, characterized in that: Terminal status is expressed as follows: in, Represents the terminal state vector. The normalized spectral waterfall plot matrix represents real-time spectral information. This matrix is ​​composed of power spectral density vectors. This indicates the strategic guidance signal sent from the satellite to the terminal at the current time t. This represents the signal-to-noise ratio (SNR) of the received signal at time t-1.

10. The NOMA satellite-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 9, characterized in that: The second reinforcement learning algorithm is a reinforcement learning algorithm based on a deep Q-network, and it approximates the Q-function through a pre-defined lightweight recurrent convolutional neural network.

11. The NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method based on reinforcement learning according to claim 10, characterized in that: Pre-defined lightweight recurrent convolutional neural network architectures include: The input layer is used to receive the spectral waterfall plot matrix; Recursive convolutional layers, composed of one-dimensional or two-dimensional convolutional kernels, use recursive calculation to extract local features of the spectrum in the frequency and time dimensions, and flatten the output of the recursive convolutional layer into a one-dimensional vector. A recursive fully connected layer is used to map the one-dimensional vector extracted by the recursive convolutional layer to the Q-value corresponding to each action; The output layer has the same output dimension as the terminal action space, and each output node corresponds to the Q value of an action.

12. A reinforcement learning-based NOMA satellite-to-ground network dual-end adaptive power allocation anti-jamming system, used to implement the reinforcement learning-based NOMA satellite-to-ground network dual-end adaptive power allocation anti-jamming method as described in any one of claims 1-11, characterized in that, include: The first establishment module is used to establish a NOMA satellite-to-ground network model. This model includes a low Earth orbit satellite, two ground users with different channel states, and multiple access points. These access points are used to generate co-channel interference to the ground users. Based on the first establishment module, a satellite-side adaptive power allocation module and a terminal-side adaptive power correction module collaboratively implement adaptive power allocation under different interference intensities. The satellite-side adaptive power allocation module includes: The second module is used to establish a channel model corresponding to the link between low Earth orbit satellites and ground users, and to calculate the channel capacity of ground users based on the channel model. The third module is used to establish an optimization problem on the satellite side based on the NOMA satellite-ground network model and channel model. The optimization problem is to dynamically adjust the power allocation coefficients corresponding to different ground users with the goal of maximizing the total throughput of two different channels. The power allocation module is used to model the optimization problem on the satellite side as a first Markov decision process, establishing a satellite-side state space, a satellite-side action space, and a satellite-side composite reward function. The state space includes the instantaneous channel gain of different ground users, the satellite elevation angle, the instantaneous rate of the users, and the interference information reported by the terminal users. The first reinforcement learning algorithm is used to solve the optimization problem based on the satellite-side composite reward function to obtain the optimal power allocation coefficient. The terminal-side adaptive power correction module includes: The fourth module is used to establish a spectrum waterfall anti-interference model on the terminal side, which is modeled as a second Markov decision process. It establishes a terminal-side state space, a terminal-side action space, and a terminal-side composite reward function. The terminal state space includes real-time spectrum information, strategic guidance signals transmitted by satellites, and the interference-to-noise ratio of the user's received signals. The power correction module is used to dynamically select different levels of transmit power in the terminal action space based on the current terminal state using a second reinforcement learning algorithm and the terminal-side composite reward function, so as to correct the guidance transmit power corresponding to the strategic guidance signal issued by the satellite, thereby realizing adaptive power allocation on the terminal side.

13. An electronic device, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the reinforcement learning-based NOMA satellite-ground network dual-end adaptive power allocation anti-interference method according to any one of claims 1-11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the reinforcement learning-based NOMA satellite-to-ground network dual-end adaptive power allocation anti-interference method as described in any one of claims 1-11.