Multi-beam hopping satellite beam level co-frequency interference multi-parameter adjustment method
By optimizing multiple parameters of co-frequency interference at the satellite beam level through a multi-agent deep reinforcement learning network, the problem of co-frequency interference between beams in multi-beam hopping satellite communication systems is solved, achieving interference avoidance with low complexity and high real-time performance, and improving the signal-to-interference-plus-noise ratio and communication quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
- Filing Date
- 2026-05-22
- Publication Date
- 2026-07-10
AI Technical Summary
In multi-beam hopping satellite communication systems, co-channel interference between beams severely affects communication quality and spectrum utilization efficiency. Existing technologies have failed to effectively optimize the multiple parameters of satellite beam-level co-channel interference, resulting in high computational complexity, poor real-time performance, and failure to fully exploit airspace resources.
By employing a multi-agent deep reinforcement learning network and combining it with a satellite beam-level co-channel interference multi-parameter optimization model, and through a centralized training and distributed execution framework, the number of satellite beams, beamwidth, beam center deviation from the satellite normal angle, and beam isolation angle are optimized to reduce the level of co-channel interference and improve the signal-to-interference-plus-noise ratio.
It achieves low-complexity, high-real-time beam co-frequency interference avoidance, improves the signal-to-interference-plus-noise ratio at the system receiver, meets the millisecond-level decision generation requirements, and significantly improves communication quality.
Smart Images

Figure CN122372061A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication technology. The proposed method for adjusting multiple parameters of co-channel interference at the beam level in a multi-beam hopping satellite system can, based on the level of co-channel interference between beams (i.e., the signal-to-interference-plus-noise ratio at the receiver), and under the limited bandwidth and power resources of the multi-beam hopping satellite communication system, jointly optimize the activation state of each beam, the width of each beam, the beam center deviation angle of each beam from the satellite normal, and the inter-beam isolation angle of each beam. With a decision generation time in milliseconds, this method reduces the level of co-channel interference between beams, increases the peak signal-to-interference-plus-noise ratio at the receiver, and improves the system's communication quality. Background Technology
[0002] In multi-beam hopping satellite communication systems employing full-frequency reuse strategies and beam-hopping technology, inter-beam co-channel interference severely restricts the communication quality and spectrum utilization efficiency of downlink user links. Effectively avoiding inter-beam co-channel interference and improving the signal-to-interference-plus-noise ratio (SNR) at the receiver is crucial for optimizing satellite communication networks. Existing research primarily manages inter-beam co-channel interference through beamforming technology and resource scheduling. However, this approach still relies on optimizing physical layer antenna shaping parameters and scheduling network layer resources, including beam-hopping patterns, satellite transmit power, and transmission bandwidth, to mitigate interference. This approach lacks sufficient exploration of the spatial resource freedom at the multi-parameter level of satellite beam-level co-channel interference. Furthermore, some studies only adjust single beam parameters such as beamwidth or beam center position to alleviate interference. The mathematical programming methods used in these studies suffer from high computational complexity and poor real-time performance. Moreover, they fail to conduct a detailed analysis of the coupling effect mechanism of multiple satellite beam-level co-channel interference parameters, including the number of hopping beams, beamwidth, beam center deviation from the satellite normal angle, and inter-beam isolation angle, on the system's SNR. Furthermore, the aforementioned studies lack sufficient analysis in complex channel environments. Summary of the Invention
[0003] The purpose of this invention is to overcome the research limitations and gaps in the aforementioned background technology. It comprehensively considers the complex satellite-to-ground channel environment, including free-space propagation loss, atmospheric attenuation, Ricean fading, and co-channel interference between beams. It explores the mechanism of co-channel interference coupling effects of the number of satellite hopping beams, beam activation state, beamwidth, beam center deviation from the satellite normal angle, and inter-beam isolation angle on the system's signal-to-interference-plus-noise ratio (SIR). The invention proposes a satellite beam-level co-channel interference multi-parameter joint optimization model and algorithm that comprehensively considers the complex channel environment of the downlink user link in a multi-beam hopping satellite communication system. This model aims to avoid co-channel interference between satellite beams in the system, thereby improving the SIR at the system receiver and achieving low-complexity, high-real-time decision-making requirements.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] A method for adjusting multiple parameters of co-frequency interference at the beam level for multi-beam hopping satellites includes the following steps:
[0006] Step 1: Mathematically characterize the multiple parameters of satellite beam-level co-frequency interference and their adjustment range, model the satellite-to-ground path loss under complex channel conditions, and then establish a joint optimization problem model for multiple parameters of satellite beam-level co-frequency interference.
[0007] Step 2: Define multiple agents, global state, action space, and global reward based on the satellite beam-level co-frequency interference multi-parameter joint optimization problem model;
[0008] Step 3, Construct a multi-agent deep reinforcement learning network structure: Using a centralized training and distributed execution framework, construct a current network and a target network; both the current network and the target network contain four Actor networks based on parameter sharing and one global hybrid Critic network; the current network and the target network have the same structure and number of parameters, only the network parameters are different; then train the current network;
[0009] Step 4: After training is completed, only the four trained Actor networks in the current network are deployed; during real-time decision-making, the current global state is input, and a joint action is generated through a single forward propagation to complete the multi-beam hopping satellite beam-level co-frequency interference multi-parameter adjustment.
[0010] Furthermore, the specific method of step 1 is as follows:
[0011] Step 101, Define the satellite beam set : Total number of beams generated for multi-beam satellites:
[0012]
[0013] Define beam activation vector :when When, it indicates the beam In the time slot Not activated, when When, it indicates the beam In the time slot Activated:
[0014]
[0015] Define beamwidth vector Using beams In the time slot Half-power beamwidth As beamwidth:
[0016]
[0017] when hour, ,when hour, ;
[0018] Define the angle vector of beam center offset from satellite normal. Define beam In the time slot The angle by which the beam center deviates from the satellite normal is :
[0019]
[0020] when hour, ,when hour, ;
[0021] Define the inter-beam isolation angle matrix Define time slots At that time, beam and The beam separation angle is ,but The inter-beam isolation angle between each beam is one. The matrix:
[0022]
[0023] Specifically, within the same time slot, ; , ,when hour, ;
[0024] In addition, when and When both are 1, ,when and When any one or all are 0, ;
[0025] Step 102, Modeling satellite-to-ground path loss in complex channel environments:
[0026]
[0027]
[0028]
[0029]
[0030]
[0031] In the current time slot , For satellite to beam Path loss corresponding to wave position; For free space propagation loss; The attenuation is due to atmospheric effects, including gas absorption attenuation. Cloud and fog attenuation Rainfall attenuation and tropospheric scintillation , The values of the four quantities are the same for all beams; For the first Riceian fading complex channel gain for each beam; The carrier frequency of the beam is the same for all beams, and the unit is GHz. For satellite to beam Distance of service points The average radius of the Earth The altitude is the satellite's orbital altitude, in kilometers (km). For satellites and beams The geocentric angle between service points;
[0032] Step 103: Construction of a multi-parameter joint optimization model for co-frequency interference at the beam level of multi-beam hopping satellites. The model is as follows:
[0033] :
[0034]
[0035] st
[0036]
[0037]
[0038]
[0039]
[0040] exist middle, Total observation period for long-term performance; It is the noise power spectral density; The total bandwidth of the satellite; Target beam in a multi-beam hopping system In the time slot At that time, the total interference power received in the system, For time slots At that time, target beam Interference beam The power of the co-channel interference signal is calculated using the following formula:
[0041]
[0042] Specifically, for each active beam, the total on-satellite power is shared equally among all beams. ,Right now:
[0043]
[0044] In the formula, in the current time slot , For target beam The transmission power; To interfere with the beam The transmission power; The number of active beams; Indicates beam The desired signal antenna gain; To interfere with the beam exist and Under these conditions, when serving its own target users, the signal radiates to the target beam. The transmit antenna gain; Target beam The receiving antenna gain at the receiving end in the satellite direction; and Target beams Interference beam The active state;
[0045] in,
[0046]
[0047]
[0048]
[0049]
[0050] The number of phased array elements; The aperture of the phased array antenna element; For antenna efficiency; and These are first-order and third-order Bessel functions, respectively.
[0051] Furthermore, the specific method for step 2 is as follows:
[0052] Defining multi-agent systems: A multi-beam hopping satellite is equipped with a total of A smart agent, among which, the beamwidth vector is adjusted. Multi-agent Adjust the beam center offset from the satellite normal angle vector. Multi-agent Control beam activation vector Multi-agent In addition, a global intelligent agent Functions to control the satellite Inter-beam isolation angle matrix between beams ;
[0053] Define global state Each agent in the system can obtain this state, and the state space is defined as follows: ;
[0054] Define action space The action space in the system is a concatenation of all agent actions, defined as follows:
[0055]
[0056] in:
[0057] for Inside The actions of an intelligent agent In the formula, Indicates in time slot Beam Beamwidth adjustment amount:
[0058]
[0059] Inside The actions of an intelligent agent, In the formula, Indicates in time slot Beam The amount of adjustment by which the beam center deviates from the satellite normal angle:
[0060]
[0061] Inside The actions of an intelligent agent, In the formula, Indicates in time slot Beam Beam activation status:
[0062]
[0063] Inside The actions of an intelligent agent ,exist , and After these three groups of agents finish performing their actions, the actions... Executed, in the formula, Indicates in time slot Beam and The amount of beam isolation angle adjustment, , , , ,when hour, :
[0064]
[0065] in, within Each agent first generates an action. If a certain beam is in an active state, the beamwidth adjustment amount, the beam center deviation angle from the satellite normal adjustment amount, and the beam isolation angle adjustment amount related to the beam are respectively taken within the above range. If a certain beam is in an inactive state, the beamwidth adjustment amount, the beam center deviation angle from the satellite normal adjustment amount, and the beam isolation angle adjustment amount related to the beam are all 0.
[0066] Define global reward Its corresponding of :
[0067] .
[0068] Furthermore, the specific method for step 3 is as follows:
[0069] The current network consists of four actors based on parameter sharing and a global hybrid Critic. The four actors are satellite beamwidth adjustment actors. Beam center deviation from satellite normal angle adjustment Actor Beam Activation Decision Adjustment Actor And beam isolation angle adjustment Actor All beams share the same set of Actor network parameters, with the first three shared Actors embedded via beam embedding vectors. Differentiating different beams, beam embedding vector The components are based on a uniform distribution of independent and identically distributed components. Perform initialization. Beam embedding vector The dimension;
[0070] for , and The intelligent agents in the game all have inputs from each other. and corresponding splicing; and The input is The four actors output their corresponding actions and concatenate them to obtain the motion space. ; This indicates flattening and concatenating a one-dimensional vector;
[0071] Global Hybrid Critic is based on global state and joint actions The joint feature vector is obtained. As input, the output Q-value is used to evaluate the policy;
[0072] Then, the network parameters of the current network are updated by combining the empirical replay pool, mean squared error loss calculation and gradient algorithm with the target network. After the update is completed, the target network is softly updated and the training process is repeated. Each training round includes T time slots. After each training round, the global reward in this training round is counted. When the global reward of multiple consecutive training rounds converges, the training ends and the network parameters are saved.
[0073] Compared with the prior art, the present invention has the following advantages:
[0074] 1. This invention is the first to jointly optimize the number of hopping beams, beamwidth, beam center deviation from satellite normal angle, and beam isolation angle in a multi-beam hopping satellite communication system, so as to improve the signal-to-interference-plus-noise ratio of the system and fully exploit the airspace resources at the multi-beam geometry level of the multi-beam hopping satellite to avoid co-frequency interference between beams.
[0075] 2. The multi-parameter joint optimization problem model for beam-level co-frequency interference of multi-beam hopping satellites in this invention, combined with the satellite orbital altitude and the electronic scanning capability of the onboard phased array antenna, more objectively clarifies the adjustment range of the number of hopping beams, beamwidth, beam center deviation from the satellite normal angle, and beam isolation angle of multi-beam hopping satellites when avoiding co-frequency interference between beams.
[0076] 3. This invention comprehensively analyzes the coupled effects of different values of the number of satellite beams, beamwidth, beam center deviation from the satellite normal angle, and inter-beam isolation angle on the signal-to-interference-plus-noise ratio and inter-beam interference level of a multi-beam hopping satellite communication system in a complex channel environment. It clarifies the influence mechanism of changes in the values of multiple parameters at the beam level of a multi-beam hopping satellite on the avoidance of co-channel interference between beams, providing an objective and scientific theoretical basis for the avoidance of co-channel interference between beams in a multi-beam hopping satellite communication system.
[0077] 4. The beam multi-parameter optimization method based on parameter sharing and multi-agent deep deterministic policy gradient effectively solves the problem of mixed high-dimensionality action space, reduces the complexity of algorithm training, significantly improves the stability of the algorithm, and can achieve millisecond-level decision generation requirements. It can also adjust satellite beam-level co-frequency interference parameters in real time to avoid co-frequency interference between beams.
[0078] 5. This invention demonstrates, through an example verification of the StarLink shell 1 satellite communication system, that the method can achieve a balance between improving communication quality and reducing algorithm decision time, providing a powerful optimization method for multi-beam hopping satellite communication systems. Attached Figure Description
[0079] Figures 1(a)-1(d) show the results under a complex channel environment that comprehensively considers large-scale fading, small-scale fading, and co-channel interference between beams, with a fixed beam spacing angle of 60°. They are respectively The figure shows the effect of co-frequency interference coupling between different values of beamwidth and beam center deviation from the satellite normal and the system signal-to-interference-plus-noise ratio (SINR) (dB).
[0080] Figures 2(a)-2(d) show the results under a complex channel environment that comprehensively considers large-scale fading, small-scale fading, and co-channel interference between beams, with a fixed beamwidth of 5°. They are respectively The figure shows the effect of co-frequency interference coupling between different values of the beam center deviation from the satellite normal angle and the beam isolation angle and the system signal-to-interference-plus-noise ratio (SINR) (dB).
[0081] Figure 3 In complex channel environments, At that time, the convergence performance of the beam multi-parameter optimization algorithm based on parameter sharing and multi-agent deep deterministic policy gradient during the training process is evaluated.
[0082] Figure 4 In complex channel environments, At that time, the performance of the beam multi-parameter optimization algorithm based on parameter sharing and multi-agent deep deterministic policy gradient as the signal-to-interference-plus-noise ratio changes with the number of test rounds during the execution test phase.
[0083] Figure 5 In complex channel environments, At that time, the average signal-to-interference-plus-noise ratio peak value obtained by the beam multi-parameter optimization algorithm based on parameter sharing multi-agent deep deterministic policy gradient during the execution test phase.
[0084] Figure 6 In complex channel environments, At that time, the average decision time of the beam multi-parameter optimization algorithm based on parameter sharing and multi-agent deep deterministic policy gradient during the execution test phase was measured. Detailed Implementation
[0085] The specific embodiments of the present invention will be described in detail with reference to the accompanying drawings and examples.
[0086] A method for adjusting multiple parameters of co-frequency interference at the beam level for multi-beam hopping satellites includes the following steps:
[0087] 1. Establish a multi-parameter joint optimization model for co-frequency interference at the beam level of multi-beam hopping satellites.
[0088] Taking Starlink Shell 1 satellites as the object, and combining the satellite orbital altitude, the minimum elevation angle of the ground terminal, and the electronic scanning capability of the onboard phased array antenna, this paper mathematically characterizes the number of hopping beams, beamwidth, beam center deviation from the satellite normal angle, and inter-beam isolation angle of multi-beam satellites, and clarifies the value range of each parameter, establishing an optimization problem. A multi-parameter joint optimization mathematical model for multi-beam hopping satellite beam-level co-frequency interference.
[0089] (1) Mathematical characterization of multiple parameters of satellite beam-level co-frequency interference and their adjustment range
[0090] Based on an in-depth analysis of satellite-cell relationships in multi-beam hopping satellite communication systems, this invention defines multiple parameters for satellite beam-level co-channel interference and their adjustment ranges. These parameters, combined with typical satellite orbital altitude, minimum elevation angle of ground terminals, and the electronic scanning capability of onboard phased array antennas, balance coverage area and inter-beam interference. The mathematical representation of these multiple parameters and their adjustment ranges in this invention is detailed below:
[0091] Satellite Beam Assembling: Multi-beam Satellite Generation Each beam covers various ground cells using spatiotemporal multiplexing. The satellite beam set is defined as follows:
[0092]
[0093] Beam activation state: using binary decision variables Describes the activation state of the beam. When At that time, beam In the time slot Not activated; otherwise, beam In the time slot Activated. Furthermore, the number of satellite beams activated by a multi-beam hopping satellite communication system at any given time should not exceed the total number of beams the satellite can generate; therefore, in each time slot... The activation vector of the satellite beam is:
[0094]
[0095] Beamwidth: In this invention, the half-power beamwidth is used as the beamwidth. To balance the coverage and inter-beam interference of multi-beam hopping satellite communication systems, a beam is defined... In the time slot The beamwidth at that time is ,but The width of each beam in the time slot It is described as:
[0096]
[0097] when hour, ,when hour, ;
[0098] Beam center deviation angle from satellite normal: The deviation angle between the beam center and the satellite normal direction; this angle is also the off-axis angle. In this invention, the satellite normal is defined as perpendicular to the ground. In a multi-beam hopping satellite communication system, the beam is defined as... In the time slot The angle by which the beam center deviates from the satellite normal is A typical satellite orbital altitude of 550 km and the minimum elevation angle of the ground terminal are used. Calculate the amount of information the satellite can provide. The value is approximately To ensure satellite coverage margin and adapt to the actual scanning capability of the phased array antenna, time slots are defined. hour, The angle vector of the beam center of each beam from the satellite normal is:
[0099]
[0100] when hour, ,when hour, ;
[0101] Inter-beam isolation angle: The angle between the centers of different satellite beams. Time slot definition. At that time, beam and The beam separation angle is ,but The inter-beam isolation angle between each beam is one. The matrix:
[0102]
[0103] Specifically, within the same time slot, ; , ,when hour, ;
[0104] In addition, when and When both are 1, ,when and When any one or all are 0, ;
[0105] (2) Modeling of satellite-to-ground path loss in complex channel environments
[0106] Channel attenuation is defined as the combined superposition of large-scale fading and small-scale fading. Furthermore, although the desired beam and interfering beam of the same satellite are spatially homogeneous, the ground areas served by different beams are different, resulting in different local scattering environments around these service areas. Therefore, the small-scale fading in different beams is considered independent. The satellite-to-ground path loss model under complex channel conditions is established as follows:
[0107]
[0108]
[0109]
[0110]
[0111]
[0112] In the current time slot , For satellite to beam Path loss corresponding to wave position; For free space propagation loss; The attenuation is due to atmospheric effects, including gas absorption attenuation. Cloud and fog attenuation Rainfall attenuation and tropospheric scintillation , The values of the four quantities are the same for all beams; For the first Riceian fading complex channel gain for each beam; The carrier frequency of the beam is the same for all beams, and the unit is GHz. For satellite to beam Distance of service points The average radius of the Earth The altitude is the satellite's orbital altitude, in kilometers (km). For satellites and beams The geocentric angle between service points.
[0113] (3) Modeling of Multi-parameter Joint Optimization Problem of Co-frequency Interference at the Beam Level of Multi-beam Jumping Satellites
[0114] By jointly optimizing and adjusting multiple parameters of beam-level co-channel interference in a multi-beam hopping satellite communication system, the level of inter-beam co-channel interference can be reduced, thereby improving the system's signal-to-interference-plus-noise ratio (SNR). The joint optimization problem model for multiple parameters of beam-level co-channel interference in multi-beam hopping satellites is as follows:
[0115] :
[0116]
[0117] st
[0118]
[0119]
[0120]
[0121]
[0122] Depend on It can be seen that, To optimize the vector, when At that time, no further calculations will be performed. The values in the constraint Indicates beam In the time slot The active state, and the total number of active beams at this time must not exceed the total number of satellite beams; constraint To constraint Each beam was limited In the time slot At that time, its beam center deviates from the satellite normal angle, its beamwidth, and its relationship with other beams The range of values for the isolation angle between them. Among them, due to the setting... The maximum value is lower than Therefore, considering the electronic scanning range of the spaceborne phased array antenna, the limitations are... .exist middle, This is the pre-set total duration of long-term performance observation; It is the noise power spectral density; The total bandwidth of the satellite; For target beam in multi-beam hopping system In the time slot At that time, the total interference power received in the system, of which express ,and , For time slots At that time, target beam Interference beam The power of the co-channel interference signal is calculated using the following formula:
[0123]
[0124] Specifically, for each active beam, the total on-satellite power is shared equally among all beams. ,Right now:
[0125]
[0126] In the formula, in the current time slot , For target beam The transmission power; To interfere with the beam The transmission power; The number of active beams; Indicates beam The desired signal antenna gain; To interfere with the beam exist and Under these conditions, when serving its own target users, the signal radiates to the target beam. The transmit antenna gain; Target beam The receiving antenna gain at the receiving end in the satellite direction; and Target beams Interference beam The active state;
[0127] The antenna gain of the beam is calculated as follows:
[0128]
[0129]
[0130]
[0131]
[0132] The number of phased array elements; The aperture of the phased array antenna element; For antenna efficiency; and These are first-order and third-order Bessel functions, respectively.
[0133] 2. Simulation investigation of the coupling effects of different values of multiple parameters of co-frequency interference at the beam level of multi-beam hopping satellite on the peak signal-to-interference-plus-noise ratio of the system.
[0134] Taking Starlink Shell 1 satellites as the target, in The analysis of the changes in the system signal-to-interference-plus-noise ratio surface under different values of multiple parameters in satellite beam-level co-frequency interference under the definition of the range of each parameter value includes two steps: (1) fixing the inter-beam isolation angle ,set up In a channel environment that comprehensively considers free space propagation loss, atmospheric attenuation, Ricean fading and co-channel interference between beams, the effects of different numbers of satellite beams, different beamwidths, and different beam center deviations from the satellite normal angle on the peak signal-to-interference-plus-noise ratio of the system are simulated and analyzed; (2) Fixed beamwidth ,set up In a channel environment that comprehensively considers free-space propagation loss, atmospheric attenuation, Ricean fading, and co-channel interference between beams, the effects of different numbers of satellite beams, different beam isolation angles, and different beam center deviations from the satellite normal angle on the peak signal-to-interference-plus-noise ratio of the system are simulated and analyzed.
[0135] 3. Design a beam multi-parameter optimization algorithm model based on parameter sharing and multi-agent deep deterministic policy gradient.
[0136] Defining multi-agent systems: A multi-beam hopping satellite has a total of One intelligent agent. Among them, adjusting beamwidth... Multi-agent Adjust the angle of the beam center off from the satellite normal. Multi-agent The multi-agent system that controls the number of activated beams is... In addition, a global intelligent agent Functions to control the satellite Inter-beam isolation angle between beams .
[0137] Define global state, action space, and global reward:
[0138] Global state Each agent in the system can obtain this state, and the state space is defined as follows: .
[0139] Action space The action space in the system is a concatenation of all agent actions, defined as follows:
[0140]
[0141] in:
[0142] for Inside The actions of an intelligent agent In the formula, Indicates in time slot beam Beamwidth adjustment amount:
[0143]
[0144] Inside The actions of an intelligent agent, In the formula, Indicates in time slot beam The amount of adjustment by which the beam center deviates from the satellite normal angle:
[0145]
[0146] Inside The actions of an intelligent agent, In the formula, Indicates in time slot beam Beam activation status:
[0147]
[0148] Inside The actions of an intelligent agent ,exist , and After these three groups of agents finish performing their actions, the actions... Executed, in the formula, Indicates in time slot beam and The amount of beam isolation angle adjustment, , , ,when hour, :
[0149]
[0150] Global Rewards All agents share a global reward. First, define the global instant reward. For the current time slot The sum of the SINR of all active beams directly reflects the quality of single-step decision-making in a multi-agent system, and its expression is as follows:
[0151]
[0152] when At that time, no further calculations will be performed. The values in;
[0153] Furthermore, there is a systematic, global, long-term cumulative reward system. , which corresponds to of :
[0154]
[0155] A multi-agent deep reinforcement learning network structure is constructed: A centralized training and distributed execution framework is adopted to build the current network, which includes four actors based on parameter sharing and a global hybrid Critic. The four actors are: a satellite beamwidth adjustment actor... Beam center deviation from satellite normal angle adjustment Actor Beam Activation Decision Adjustment Actor And beam isolation angle adjustment Actor All beams share the same set of Actor network parameters, with the first three shared Actors embedded via beam embedding vectors. Distinguish between different beams As a learnable vector, its components follow a uniform distribution of independent and identically distributed components. Perform initialization. Let be the dimension of the embedding vectors. The embedding vectors of all beams form a learnable lookup table. This lookup table serves as the trainable parameters for the parameter-sharing Actor network. During training, The parameters of the Actor network are updated via gradient descent. In this embodiment, .
[0156] The structure and parameters of each network are now explained. For , and The intelligent agents in the game all have inputs from each other. and corresponding The concatenation, with all input dimensions being... In this embodiment, ;and The input dimension is only The input dimension is ,Right now =16; This indicates flattening and concatenating a one-dimensional vector;
[0157] Specifically, In the network The network structure of each agent is consistent, containing a hidden layer with 64 neurons (using the ReLU activation function) and an output layer with 1 neuron, which uses the Tanh activation function to map the output to... Then multiply by a learnable scaling factor (initialized to 0.1) to obtain the optimal beamwidth adjustment.
[0158] In the network The network structure of each agent is consistent, containing two hidden layers with 128 and 64 neurons respectively, both using the ReLU activation function, and an output layer with one neuron, which uses the Tanh activation function to map the output to... Then multiply by a learnable scaling factor (initialized to 0.1) to output the optimal beam center offset from the satellite normal angle adjustment;
[0159] In the network The network structure of each agent is consistent, containing a hidden layer with 32 neurons (using the ReLU activation function) and an output layer with 1 neuron, which uses the Sigmoid activation function to map the output, with the output range being... , which represents the activation probability of each beam, and a beam is activated when its activation probability is greater than 0.5. In addition, the network has no learnable scaling coefficients.
[0160] The agent in the model contains two hidden layers with 256 and 128 neurons respectively, both using the ReLU activation function, and also includes an output layer. Each neuron corresponds to an upper triangular element of the inter-beam isolation angle adjustment matrix. This layer has no activation function, and the output is symmetric. The final inter-beam isolation angle adjustment matrix is obtained by multiplying the beams by a scaling factor of 0.1.
[0161] in, The network first outputs the beam activation status. If a beam is not activated, then the corresponding beam... , The agent in the network outputs 0 directly; otherwise, it outputs according to the network structure.
[0162] If a certain beam is not activated, then The agent in the network outputs 0 directly to the neuron corresponding to the beam; otherwise, it outputs according to the network structure.
[0163] In addition, when the satellite beamwidth is adjusted, Actor Beam center deviation from satellite normal angle adjustment Actor Actor and beam isolation angle adjustment When the output results cause the satellite beamwidth, beam center deviation from the satellite normal angle, and inter-beam isolation angle to exceed the constraint range, the corresponding agent will readjust the output.
[0164] Adjust the beamwidth of the four satellites (Actor) Beam center deviation from satellite normal angle adjustment Actor Beam Activation Decision Adjustment Actor And beam isolation angle adjustment Actor The outputs are concatenated to obtain the motion space. ;
[0165] Global Hybrid Critic receives global state and joint actions The Q-value is output to evaluate the policy. The calculation process for the Q-value is as follows: First, ... and Flatten each feature vector into a one-dimensional vector and concatenate them along the feature dimension to form a joint feature vector. , here and These represent the state space dimension and action space dimension, respectively. The agent in the global hybrid Critic network contains three fully connected layers with 256, 128, and 1 neurons, respectively. The first two are hidden layers using the ReLU activation function, and the last is the output layer without an activation function. The global hybrid Critic network receives a joint feature vector. Directly output the scalar Q value:
[0166] .
[0167] Experience gathering and network updates: in every time slot Current strategy Based on global state Generate joint actions After the environment performs this action, it returns to the next state. Instant rewards for the current time slot In this embodiment, all agents share the same instant reward value. ,Right now .
[0168]
[0169] Subsequently, the empirical tuple Store the data in the experience replay pool. Randomly sample a small batch of data from the experience pool (in the early stages of training, if the number of samples in the experience pool is less than the batch size). If the number of samples in the experience pool is not less than a certain threshold, the update step is skipped, and experience collection continues. At that time, according to (Random sampling is performed), and the target Q-value of the sampled empirical tuples is calculated using the target network. The target network includes a target Critic network. and four target Actor networks (respectively) , , , The target network, with its identical structure and number of parameters to the current network, differs only in its network parameters. In this process, since the current network parameters are updated via gradient descent at each step, directly using the current network to calculate the target value would cause drastic fluctuations, making training difficult to converge. The target network, as a lagging copy of the current network, has its parameters changed slowly through soft updates, effectively avoiding training oscillations and thus providing a smooth and stable regression target for the current Critic network. For the sampled... Let there be a set of empirical tuples, and denote their state in the current time slot as... The corresponding combined action is Instant rewards are The next state is The network consists of four target actors based on... Generate corresponding joint actions The four target actor networks generate corresponding joint actions. It also meets the constraints of beam activation state (i.e., the adjustment amount corresponding to the inactive beam is 0), as well as the constraints of beamwidth, beam center deviation from satellite normal angle and beam isolation angle. A joint feature vector is formed and input into the target Critic network, outputting Q. Value, further calculation The expression is as follows:
[0170]
[0171] This formula, based on the Bellman optimality equation, indicates that the expected cumulative reward of the current state-action pair equals the immediate reward. The sum of future discounted returns. Wherein, the discount factor... Used to balance immediate and long-term rewards. The larger the value, the more emphasis is placed on future returns. In this embodiment, , These are the parameters of the target Critic network. Based on this, the current Critic network... The update is performed by minimizing the following mean squared error loss:
[0172]
[0173] The loss is backpropagated through the Adam optimizer, updating only the parameters of the current Critic network. Furthermore, the four Actors are updated sequentially, in the following order: , , , When updating one Actor, the outputs of the other three Actors are fixed, and only the output of the current Actor is included in the joint action construction, with the goal of maximizing the Q-value of the current Critic network output. This goal is achieved through gradient ascent, equivalent to the loss function. , ( ) represents the expected value; each update only improves the Q-value of the current state-action pair in a single step, rather than iterating repeatedly until convergence. Subsequently, the Adam optimizer is used to update the parameters of the current Actor network. The parameters of the four current Actor networks are as follows: , , , .
[0174] In particular, for , , Three actors, whose trainable parameters also include a beam embedding lookup table. In response, during the training process, The update method is as follows: During each forward computation, since the above three Actors need to... and their respective The concatenated sequence is used as input; therefore, Directly involved in action generation; during backpropagation, the gradient of the Actor network's optimization objective (i.e., maximizing the Q-value of the current Critic output) with respect to the network parameters is backpropagated to the input layer, because... Also as part of the input, then regarding each The gradient will also be further propagated to The corresponding vector in, and because Shared by three Actors, each Actor has access to... The gradients need to be accumulated sequentially to avoid subsequent updates overwriting previous gradients. Then, the Adam optimizer updates them synchronously with each Actor update. .
[0175] Target network soft update: After each training step (i.e., after completing a mini-batch sampling, Critic update, and sequential updates of the four Actors), a soft update method is used to gradually blend the parameters of the current network into the parameters of the target network at a small proportion, making the target value change smoother and enhancing the algorithm's training stability. Specifically, since the target Actor network also needs to maintain a corresponding beam embedding lookup table... It will also use soft updates to ensure the stability of the algorithm training. The specific formula for soft updates is as follows:
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182] The parameters of the four target actor networks are as follows: (Corresponding beam activation decision adjustment) (Corresponding satellite beamwidth adjustment) (Corresponding to the adjustment of the beam center's deviation from the satellite normal angle) (Corresponding to the beam isolation angle adjustment); In this embodiment, the soft update coefficient is used. This process does not involve the Adam optimizer; it is purely a linear interpolation operation.
[0183] The training process is repeated, with each training round consisting of T time slots. After each training round, the global long-term cumulative reward for that training round is calculated. When the global long-term cumulative reward of multiple consecutive training rounds Upon convergence, training ends, and the network parameters and beam embedding lookup table are saved. .
[0184] Algorithm Execution and Real-Time Decision-Making: After training, only the four trained Actor networks from the current network are deployed. During real-time decision-making, the current global state is input, combined with the beam embedding lookup table obtained at the end of training. Joint actions can be generated with a single forward propagation. This process involves only a small number of matrix multiplications and activation function calculations, resulting in low computational complexity and meeting the real-time control requirements of satellite systems with millisecond-level time slots.
[0185] Experimental verification
[0186] Taking Starlink Shell 1 satellites as an example, the specific results are as follows after following the above steps.
[0187] The coupling effects of different values of multiple parameters for co-channel interference at the beam level of multi-beam hopping satellites on the peak signal-to-interference-plus-noise ratio (SNR) of the system are shown in Figures 1(a)-1(d) and 2(a)-2(d). The results indicate that in multi-beam hopping satellite communication systems: 1) In complex channel environments considering both large-scale and small-scale fading, the nonlinear changes in the SNR surface are significant under different values of the multiple parameters for co-channel interference at the beam level; 2) Increasing the number of hopping beams and the angle of beam center deviation from the satellite normal will significantly enhance co-channel interference between beams, thus reducing the peak SNR of the system; 3) A beamwidth greater than... At that time, the peak signal-to-interference-plus-noise ratio of the system exhibits a typical nonlinear change, and the beamwidth is less than At this time, the peak value of the system signal-to-interference-plus-noise ratio tends to stabilize, and the level of co-channel interference is significantly reduced; 4) the inter-beam isolation angle is less than At this time, the peak value of the signal-to-interference-plus-noise ratio exhibits a significant non-linear change; when this angle is smaller than... The level of co-frequency interference between beams in the system is significantly enhanced, and the system signal-to-interference-plus-noise ratio drops sharply.
[0188] For the number of beams In the embodiment, the beam multi-parameter optimization algorithm model based on parameter sharing and multi-agent deep deterministic policy gradient, during the training process, exhibits the following convergence performance for the system signal-to-interference-plus-noise ratio: (see attached figure) Figure 3As shown in the figure. The results show that the beam multi-parameter optimization algorithm model based on parameter sharing and multi-agent deep deterministic policy gradient proposed in this invention can effectively learn stable and effective policies. The multi-agents can learn from each other and cooperate. The peak signal-to-interference-plus-noise ratio in the optimized multi-beam hopping satellite communication system can reach about 21 dB.
[0189] For the above embodiments, the performance of the algorithm model in avoiding co-frequency interference during test execution, compared with traditional stochastic optimization and greedy strategy optimization methods, is shown in the appendix. Figure 4 and attached Figure 5 As shown in the figure. The results show that the beam multi-parameter optimization algorithm model based on parameter sharing and multi-agent deep deterministic policy gradient proposed in this invention exhibits the best performance. Its average system signal-to-interference-plus-noise ratio (SINR) peak value during the test execution phase is improved by 23.9% and 20.7% respectively compared to the two traditional methods, and its average SINR peak value reaches 21.53 dB, while also demonstrating significant stability. Furthermore, the decision-time performance results of the algorithm model proposed in this invention during the test execution phase are attached. Figure 6 As shown in the figure. The results show that, through a sufficient offline training phase, it can learn the direct mapping relationship from the system state to the optimal action, enabling it to generate multi-beam jump satellite beam-level co-frequency interference multi-parameter optimization adjustment decisions within 2ms while maintaining optimal decision quality, and to avoid inter-beam co-frequency interference in the system in real time to improve the system's communication quality.
Claims
1. A method for adjusting multiple parameters of co-frequency interference at the beam level for multi-beam hopping satellites, characterized in that, Includes the following steps: Step 1: Mathematically characterize the multiple parameters of satellite beam-level co-frequency interference and their adjustment range, model the satellite-to-ground path loss under complex channel conditions, and then establish a joint optimization problem model for multiple parameters of satellite beam-level co-frequency interference. Step 2: Define multiple agents, global state, action space, and global reward based on the satellite beam-level co-frequency interference multi-parameter joint optimization problem model; Step 3, Construct a multi-agent deep reinforcement learning network structure: Using a centralized training and distributed execution framework, construct a current network and a target network; both the current network and the target network contain four Actor networks based on parameter sharing and one global hybrid Critic network; the current network and the target network have the same structure and number of parameters, only the network parameters are different; then train the current network; Step 4: After training is completed, only the four trained Actor networks in the current network are deployed; during real-time decision-making, the current global state is input, and a joint action is generated through a single forward propagation to complete the multi-beam hopping satellite beam-level co-frequency interference multi-parameter adjustment.
2. The method for adjusting multiple parameters of multi-beam hopping satellite beam-level co-frequency interference according to claim 1, characterized in that: The specific method for step 1 is as follows: Step 101, Define the satellite beam set : Total number of beams generated for multi-beam satellites: ; Define beam activation vector :when When, it indicates the beam In the time slot Not activated, when When, it indicates the beam In the time slot Activated: ; Define beamwidth vector Using beams In the time slot Half-power beamwidth As beamwidth: ; when hour, ,when hour, ; Define the angle vector of beam center offset from satellite normal. Define beam In the time slot The angle by which the beam center deviates from the satellite normal is : ; when hour, ,when hour, ; Define the inter-beam isolation angle matrix Define time slots At that time, beam and The beam separation angle is ,but The inter-beam isolation angle between each beam is one. The matrix: ; Specifically, within the same time slot, ; , ,when hour, ; In addition, when and When both are 1, ,when and When any one or all are 0, ; Step 102, Modeling satellite-to-ground path loss in complex channel environments: ; ; ; ; ; In the current time slot , For satellite to beam Path loss corresponding to wave position; For free space propagation loss; The attenuation is due to atmospheric effects, including gas absorption attenuation. Cloud and fog attenuation Rainfall attenuation and tropospheric scintillation , The values of the four quantities are the same for all beams; For the first Riceian fading complex channel gain for each beam; The carrier frequency of the beam is the same for all beams, and the unit is GHz. For satellite to beam Distance of service points The average radius of the Earth The altitude is the satellite's orbital altitude, in kilometers (km). For satellites and beams The geocentric angle between service points; Step 103: Construction of a multi-parameter joint optimization model for co-frequency interference at the beam level of multi-beam hopping satellites. The model is as follows: : ; st ; ; ; ; exist middle, Total observation period for long-term performance; It is the noise power spectral density; The total bandwidth of the satellite; Target beam in a multi-beam hopping system In the time slot At that time, the total interference power received in the system, For time slots At that time, target beam Interference beam The power of the co-channel interference signal is calculated using the following formula: ; Specifically, for each active beam, the total on-satellite power is shared equally among all beams. ,Right now: ; In the formula, in the current time slot , For target beam The transmission power; To interfere with the beam The transmission power; The number of active beams; Indicates beam The desired signal antenna gain; To interfere with the beam exist and Under these conditions, when serving its own target users, the signal radiates to the target beam. The transmit antenna gain; Target beam The receiving antenna gain at the receiving end in the satellite direction; and Target beams Interference beam The active state; in, ; ; ; ; The number of phased array elements; The aperture of the phased array antenna element; For antenna efficiency; and These are first-order and third-order Bessel functions, respectively.
3. The method for adjusting multiple parameters of multi-beam hopping satellite beam-level co-frequency interference according to claim 1, characterized in that: The specific method for step 2 is as follows: Defining multi-agent systems: A multi-beam hopping satellite is equipped with a total of A smart agent, among which, the beamwidth vector is adjusted. Multi-agent Adjust the beam center offset from the satellite normal angle vector. Multi-agent Control beam activation vector Multi-agent In addition, a global intelligent agent Functions to control the satellite Inter-beam isolation angle matrix between beams ; Define global state Each agent in the system can obtain this state, and the state space is defined as follows: ; Define action space The action space in the system is a concatenation of all agent actions, defined as follows: ; in: for Inside The actions of an intelligent agent In the formula, Indicates in time slot Beam Beamwidth adjustment amount: ; Inside The actions of an intelligent agent, In the formula, Indicates in time slot Beam The amount of adjustment by which the beam center deviates from the satellite normal angle: ; Inside The actions of an intelligent agent, In the formula, Indicates in time slot Beam Beam activation status: ; Inside The actions of an intelligent agent ,exist , and After these three groups of agents finish performing their actions, the actions... Executed, in the formula, Indicates in time slot Beam and The amount of beam isolation angle adjustment, , , , ,when hour, : ; in, within Each agent first generates an action. If a certain beam is in an active state, the beamwidth adjustment amount, the beam center deviation angle from the satellite normal adjustment amount, and the beam isolation angle adjustment amount related to the beam are respectively taken within the above range. If a certain beam is in an inactive state, the beamwidth adjustment amount, the beam center deviation angle from the satellite normal adjustment amount, and the beam isolation angle adjustment amount related to the beam are all 0. Define global reward Its corresponding of : 。 4. The method for adjusting multiple parameters of multi-beam hopping satellite beam-level co-frequency interference according to claim 1, characterized in that: The specific method for step 3 is as follows: The current network consists of four actors based on parameter sharing and a global hybrid Critic. The four actors are satellite beamwidth adjustment actors. Beam center deviation from satellite normal angle adjustment Actor Beam Activation Decision Adjustment Actor And beam isolation angle adjustment Actor All beams share the same set of Actor network parameters, with the first three shared Actors embedded via beam embedding vectors. Differentiating different beams, beam embedding vector The components are based on a uniform distribution of independent and identically distributed components. Perform initialization. Beam embedding vector The dimension; for , and The intelligent agents in the game all have inputs from each other. and corresponding splicing; and The input is The four actors output their corresponding actions and concatenate them to obtain the motion space. ; This indicates flattening and concatenating a one-dimensional vector; Global Hybrid Critic is based on global state and joint actions The joint feature vector is obtained. As input, the output Q-value is used to evaluate the policy; Then, the network parameters of the current network are updated by combining the empirical replay pool, mean squared error loss calculation and gradient algorithm with the target network. After the update is completed, the target network is softly updated and the training process is repeated. Each training round includes T time slots. After each training round, the global reward in this training round is counted. When the global reward of multiple consecutive training rounds converges, the training ends and the network parameters are saved.