A method for resource allocation in a multi-beam satellite communication system
By designing user clustering strategies, beam illumination, and power allocation in a multi-beam satellite communication system, and utilizing a phased strategy gradient (PPG) network to optimize resource allocation, the performance improvement problem of multi-beam low-Earth orbit satellite networks under dynamic changes and resource constraints was solved, achieving a comprehensive improvement in system performance.
Patent Information
- Application Number
- CN202510036096.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Multi-beam low-Earth orbit satellite networks struggle to serve all users simultaneously and optimize system performance when faced with dynamic changes in network topology caused by satellite motion, uneven user service demands, and resource constraints.
By modeling a satellite communication system, user clustering strategies, beam illumination, and power allocation are designed. Resource allocation is optimized using a phased policy gradient (PPG) network, and beamforming strategies are combined to maximize the system's cumulative reward.
This achieves improved overall system performance, optimized resource utilization, and enhanced overall system performance while ensuring users' communication needs are met.
Smart Images

Figure CN119893684B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of wireless communication, relates to the field of satellite communication, and particularly relates to a resource allocation method for a multi-beam satellite communication system. BACKGROUND
[0002] In recent years, multi-beam low-orbit satellite networks have become an effective supplement to ground cellular networks due to their ability to provide global coverage, meet diversified service demands, and resist major natural disasters. In order to improve the resource utilization rate and overall system performance of multi-beam low-orbit satellite networks, beam illumination optimization and wireless resource allocation are effective solutions, which allocate different time-frequency resources and power resources to users to maximize the satisfaction of different service demands of users. However, due to the dynamic changes of network topology caused by satellite motion, the uneven geographical distribution of user service demands, and the contradiction between the growing user demands and the limited resources, the beam illumination and resource allocation problems of multi-beam low-orbit satellite systems are challenging, and how to design an efficient beam illumination, beam power allocation, and intra-cluster beamforming strategy to improve the system performance has become an important research topic.
[0003] Existing research has considered the resource allocation problem of multi-beam satellite communication systems, but few studies have considered the optimization of the inability to simultaneously serve all users and the average performance of the system in the multi-beam low-orbit satellite scenario, resulting in limited performance of the resource allocation scheme. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a resource allocation method for a multi-beam satellite communication system, which models the cumulative reward of the system as an optimization target for a communication system containing multiple multi-beam low-orbit satellites and multiple ground users, optimizes the determination of beam illumination, beam power allocation, and intra-cluster beamforming strategy, and improves the overall system performance.
[0005] To achieve the above purpose, the present application provides the following technical scheme:
[0006] A resource allocation method for a multi-beam satellite communication system, specifically comprising the following steps:
[0007] S1: modeling a satellite communication system model;
[0008] S2: determining a user clustering strategy;
[0009] S3: modeling beam illumination variables and power allocation variables;
[0010] S4: modeling a user rate model;
[0011] S5: modeling a user service model and a satellite queue model;
[0012] S6: modeling the amount of data to be transmitted by the user cluster and a system cost function;
[0013] S7: modeling system resource allocation constraints;
[0014] S8: modeling system state, action, and reward;
[0015] S9: constructing and training a periodic policy gradient (PPG) network;
[0016] S10: determining a system resource allocation strategy using the trained PPG network.
[0017] Further, in step S1, a satellite communication system model is modeled, specifically including: the system includes a plurality of multi-beam low-orbit satellites and a plurality of ground users, let M represent the number of low-orbit satellites, S m represents the mth satellite, each satellite is equipped with a uniform planar array with a size of N s = N x × N y , each satellite can simultaneously generate N beams, let K represent the number of ground users, U k represents the kth user, each user is a single antenna user; let P tot represent the total power of a single satellite, P max represents the maximum transmission power of a single satellite beam; the system time is divided into T consecutive equal length time slots, and the time slot length is τ; let q m (t) = [x m (t), y m (t), H m (t)] T represent the coordinates of S m at time t, q k = [x k , y k , 0] T represent the coordinates of U k ; let B represent the total bandwidth of the satellite beam, and assume that all beams of a single satellite occupy N time-frequency resource blocks in an orthogonal manner;
[0018] A satellite channel model is modeled, specifically including: let h m,k (t) represent the channel gain of the link between S m and U k at time t, which can be modeled as where g m,k (t), υ m,k , τ m,k , and v m,k are the complex gain, Doppler shift, propagation delay, and array response vector of the link from S m to U k at time t, fc is the carrier frequency; g m,k (t) can be modeled as where δ m,k (t) denotes the Rician fading factor between the t-th time slot S m and U k , c and d m,k (t) denote the transmit antenna gain of S m , the receive antenna gain of U k , the speed of light and the distance between S m and U k , respectively; v m,k can be modeled as where, and denote the array response vectors in x-axis and y-axis directions, respectively, denotes the Kronecker product.
[0019] Further, in step S2, the user clustering strategy is determined, specifically comprising: letting R max denote the satellite beam coverage radius, and applying the mean shift algorithm to design the user clustering strategy, the specific steps of which are as follows:
[0020] (1) Initialization: Let Φ denote the set of unclustered users, i.e., Φ = {U k ,1≤k≤K}, and C i denote the user set of the i-th cluster, 1≤i<K, let i = 1;
[0021] (2) Select an initial center point: randomly select an unclustered user, and take its position as the initial center point, for example, select U k ∈Φ, take q k as the initial center point of C i , denoted as
[0022] (3) Determine the initial cluster members: calculate the distance between the unclustered users and , if , then add U k′ to C i , i.e., C i =C i ∪{U k′}, and let L i denote the number of users in C i ;
[0023] (4) Update the cluster center point: let M i denote the mean shift vector of C i , which can be modeled as: where Let σ be the Gaussian kernel function, and σ be the parameters of the Gaussian kernel function; based on Update C i The center point;
[0024] (5) Update cluster members: Repeat steps (3) and (4) until M. i =0; Update C i Delete C i Zhongyu The distance is greater than R max The user, that is, if C i =C i \{U k′}; Update Φ, delete clustered users, if U k′ ∈C i Then Φ=Φ\{U k′};
[0025] (6) Determine if the initial clustering algorithm has terminated: Determine if there are any unclustered U cells. k′ ,like Then let i = i + 1 and return to step (2); otherwise, execute step (7) and record the total number of user clusters as I.
[0026] (7) Calculate the clustering strategy evaluation function: Determine the clustering performance evaluation function based on intra-cluster and inter-cluster distances, let S i C represents i The dispersion is defined as C i The average distance from all users within the network to the center point can be modeled as follows: make C represents i With C j The distance between them can be modeled as Let F i,j C i With C j The similarity measure between them can be modeled as Let F i C represents i The evaluation function can be modeled as Let F denote the clustering policy evaluation function, which can be modeled as
[0027] (8) Determine the clustering strategy update condition: Let F th For a predefined evaluation function threshold, if F≤F th The algorithm terminates, and the current clustering result is the final result. The user's clustering strategy β is output. i,k ∈{0,1}, that is, if U k ∈C i , then β i,k =1, otherwise βi,k = 0; if F > F th , let σ = ασ, 0 < α < 1, return to step (1).
[0028] Further, in step S3, the beam illumination variable is modeled, specifically including: modeling α m,n,i (t) ∈ {0, 1} as the illumination variable between beam n of time slot S m and C i , if beam n of time slot S m illuminates C i , then α m,n,i (t) = 1, otherwise α m,n,i (t) = 0.
[0029] The power allocation variable is modeled, specifically including: modeling the beam power allocation vector of time slot S m as P m (t) = [p m,1 (t), p m,2 (t), … p m,N (t)] T , where p m,n (t) represents the transmission power of beam n of time slot S m .
[0030] Further, in step S4, the user rate model is modeled, specifically including: dividing the message sent by the satellite to the user into a public part and a private part, uniformly encoding the public part of all user messages in each cluster as a public stream, and independently encoding the private part of each user message as a user private stream; the satellite designs precoding for the public stream and multiple private streams of the users in the cluster respectively, and multiplexes and transmits, after receiving the message from the satellite, the users in the cluster decode the public stream and apply successive interference cancellation (SIC) technology, and after removing the public stream from the received signal, each user decodes its corresponding private stream respectively;
[0031] Let s i,c (t) represent the public stream of all user messages in time slot C i , s i,k (t) represent the private stream of user U i in time slot C k , let represent the beamforming matrix when S m communicates with C i , where w m,i,c (t) and w m,i,k (t) are the precoding vectors of the public stream and the private stream of U k respectively; the signal sent by S m to C i in time slot t can be modeled as t time slot Ci inner U k Upon receiving the signal from S m , the signal can be modeled as where denotes the additive white Gaussian noise;
[0032] Let γ m,i,k,c (t) and γ m,i,k (t) denote the signal-to-interference-plus-noise ratio (SINR) of the common stream and the private stream, respectively, in the signal received by U i inner U k from S m , which can be modeled as and Let R m,i,c (t) denote the transmission rate of the common stream from S m to C i at time slot t, which can be modeled as Let R m,i,k (t) denote the transmission rate of the private stream from S m to U i inner U k at time slot t, which can be modeled as R m,i,k (t) = Blog2(l + γ m,i,k (t)); let denote the sum rate from S m to U i inner U k at time slot t, which can be modeled as
[0033] Further, in step S5, the user traffic model is modeled, specifically including: assuming that the user traffic streams arrive randomly and dynamically, and the amount of traffic arriving at each user in each time slot follows a Poisson distribution; let A k (t) denote the amount of traffic arriving at U k at time slot t, and the expectation is E[A k (t)] = λ k τ, where λ k denotes the average arrival rate of the traffic streams of U k ;
[0034] The satellite queue model is modeled, specifically including: each satellite is equipped with a user data buffer server, and the data streams randomly arriving at each user can be stored in the corresponding queue; let Q max denote the maximum capacity of the satellite buffer; let Q m,k (t) denote the queue length of the data streams of U m buffered in S k at the end of time slot t, which can be modeled as:
[0035]
[0036] Further, in step S6, the amount of data to be transmitted of the user cluster is modeled, specifically including: let O m,i (t) denote the amount of data to be transmitted of the user cluster in time slot t m to C i , which can be modeled as Let r(t) denote the system cost function in time slot t, which can be modeled as
[0037] Further, in step S7, the system resource allocation constraints are modeled, specifically including:
[0038] (1) Beam illumination constraint
[0039] Each single beam of each satellite can serve at most one user cluster in any time slot, then we have:
[0040]
[0041] Each user cluster can be served by at most one beam of one satellite in any time slot, then we have:
[0042]
[0043] (2) Beam transmit power constraint
[0044] The total transmit power of multiple beams of each satellite in any time slot cannot exceed the total transmit power of the satellite, then we have:
[0045]
[0046] There is a maximum transmit power limit for each single beam in any time slot, then we have:
[0047]
[0048] The power that can be allocated to the data stream of each user in a cluster cannot exceed the transmit power of the beam, then we have:
[0049]
[0050] In order to successfully implement SIC while decoding the common stream at the receiving end, the signal strength of the common stream should be greater than that of the private stream and noise, then we have:
[0051]
[0052] where θ th is the minimum power difference between the common stream and the private stream and noise.
[0053] Further, in step S8, the system state, action and reward are modeled, specifically including: defining the global state space in time slot t as s t = {s t,m|1≤m≤M}, where s t,m ={Q m,k (t),h m,k (t)1≤k≤K} represents time slot S of time t. m The state; define the joint action space of time slot t as a t ={α m,n,i (t),p m,n (t),W m,i (t)1≤m≤M,1≤n≤N,1≤i≤I}, which includes beam illumination, beam power allocation and intra-cluster beamforming strategies; let r(t) be the reward function of the t-slot system.
[0054] Further, in step S9, a phased policy gradient (PPG) network is constructed and trained, specifically including: the PPG network comprises a policy network and a value network, where θ and φ represent the parameters of the policy network and the value network, respectively; the training process of the PPG network alternates between a policy training phase and a knowledge distillation-assisted training phase, wherein the policy training phase uses the proximal policy optimization (PPO) algorithm to train the agent, and the knowledge distillation-assisted training phase can send useful information from the value network to the policy network; the parameters of the policy network and the value network, as well as the experience replay cache, are initialized. Given an initial state s t The agent, based on the output π(a) of the policy network, t |s t ;θ t ) Perform action a t Rewards are obtained after interacting with the environment. t The system transitions to the next state s t+1 , the quadruple (s t ,a t ,r t ,s t+1 Deposit During each network parameter update process, it is from Training samples are extracted from the data, and the parameters of the policy network and the value network are updated alternately.
[0055] Define the loss function of the policy network as follows: in This indicates the difference between the current policy and the previous policy in state s. t Take action a at the time t The probability ratio, π(a) t |s t ;θ t ) indicates the state s under the current policy. t Take action a at that time t The probability, π(a) t |s t ;θ old ) indicates the state s under the previous policy.t Take action a at that time t The probability; A(t) = δ t +(γλ)δ t+1 +…+(γλ) X-t+1 δ X-1 δ represents the dominance function for time slot t. t =r t +γV φ (s t+1 )-V φ (s t ) represents the time difference error, V φ (s t ) represents state s t The value function is defined as follows: γ∈(0,1) is the discount factor, λ is a constant, and X is the trajectory length; ε∈(0,1) is the cutoff coefficient, and clip(μ) is the value function. t (θ t (),1-ε,1+ε) are cutoff functions, representing the reduction of μ t (θ t The policy update magnitude is limited to [1-ε, 1+ε] to ensure it is not too large. The stochastic gradient ascent method is used to update the policy network parameters θ, and the update formula can be modeled as θ... t+1 =θ t +η·▽ θ L(θ t ), among which ▽ θ L(θ t ) represents L(θ) t The gradient with respect to parameter θ, where η is the learning rate;
[0056] Define the loss function of the value network as follows: in Let be the target value function for time slot t; the parameters φ of the value network are updated using stochastic gradient descent, and the update formula can be modeled as φ t+1 =φ t -η·▽ φ L(φ t );
[0057] In the knowledge distillation-assisted training phase, the loss function of the policy network is defined as follows: Where L(θ) t ) aux To supplement the loss function, the auxiliary value function V of the policy network is utilized. θ (s t The useful information from the learning value network can be modeled as follows: L(θ t ) joint The second term is the behavioral cloning loss component, β clone For cloning parameters, π(·s)t ;θ old ) represents the strategy before the start of the auxiliary phase, π(·s) t ;θ t ) represents the current strategy, KL(π(·s) t ;θ old ),π(·s t ;θ t )) is used to calculate π(·s t ;θ old ) and π(·s t ;θ t The relative entropy between the policy network and the value network is used; the parameters θ and φ of the policy network are updated separately using stochastic gradient descent, and their update formulas can be modeled as θ t+1 =θ t -η·▽ θ L(θ t ) joint and φ t+1 =φ t -η·▽ φ L(φ t ).
[0058] Furthermore, in step S10, the trained PPG network is used to determine the beam illumination, beam power allocation, and intra-cluster beamforming strategies. Specifically, this includes: under the condition of satisfying the beam illumination and beam transmit power constraints, optimizing and determining the resource allocation strategy with the objective of maximizing the system's cumulative reward, i.e.:
[0059]
[0060] in and These are the optimal beam illumination, beam power allocation, and intra-cluster beamforming strategies, respectively.
[0061] The beneficial effects of this invention are as follows: the method of this invention can maximize the cumulative reward of the system and improve the overall performance of the system by based on beam illumination, beam power allocation and intra-cluster beamforming strategies, while ensuring the communication needs of different users.
[0062] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0064] Figure 1 This is a schematic diagram of a multi-beam satellite communication system scenario.
[0065] Figure 2 This is a flowchart illustrating the resource allocation method for the multi-beam satellite communication system of the present invention. Detailed Implementation
[0066] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0067] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0068] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0069] Please see Figures 1-2 , Figure 1 This is a schematic diagram of a multi-beam satellite communication system scenario, such as... Figure 1 As shown, the communication system includes multiple multi-beam low-orbit satellites and multiple ground users. By jointly designing optimal beam illumination, beam power allocation and intra-cluster beamforming strategies, the system's cumulative reward can be maximized.
[0070] Figure 2 This is a flowchart illustrating the resource allocation method for the multi-beam satellite communication system of the present invention, as shown below. Figure 2 As shown, the method specifically includes the following steps:
[0071] Step 1: Modeling the satellite communication system model.
[0072] The system contains multiple multi-beam low-orbit satellites and multiple ground users, let M represent the number of low-orbit satellites, S m represents the mth satellite, each satellite is equipped with a uniform planar array with a size of N s = N x × N y , each satellite can generate N beams at the same time, let K represent the number of ground users, U k represents the kth user, each user is a single antenna user; let P tot represent the total power of a single satellite, P max represents the maximum transmission power of a single satellite beam; divide the system time into T consecutive equal length time slots, the time slot length is τ; let q m (t) = [x m (t), y m (t), H m (t)] T represent the coordinates of S m at time t, q k = [x k , y k , 0] T represent the coordinates of U k ; let B represent the total bandwidth of the satellite beam, assume that all beams of a single satellite occupy N time-frequency resource blocks in an orthogonal manner;
[0073] Modeling the satellite channel model, specifically including: let h m,k (t) represent the channel gain of the link between S m and U k at time t, which can be modeled as where g m,k (t), υ m,k , τ m,k and ν m,k are the complex gain, Doppler shift, propagation delay and array response vector of the link from S m to U k at time t, f c is the carrier frequency; g m,k (t) can be modeled as where δ m,k (t) represents the Rician fading factor between S m and U k at time t, c and d m,k (t) represent the transmit antenna gain of S m and U kThe receiving antenna gain, the speed of light, and the t time slot S m The distance between U k The distance between U m,k The distance between U Wherein, and respectively represent the array response vectors in the x-axis and y-axis directions, represents the Kronecker product.
[0074] Step 2: Determine the user clustering strategy.
[0075] Let R max represent the satellite beam coverage radius, and the mean shift algorithm is applied to design the user clustering strategy, and the specific steps are as follows:
[0076] (1) Initialization: Let Φ represent the set of unclustered users, that is, Φ = {U k ,1≤k≤K}, C i represent the user set of the i-th cluster, 1≤i<K, let i = 1;
[0077] (2) Select the initial center point: randomly select an unclustered user, and take its position as the initial center point, such as selecting U k ∈Φ, taking q k as the initial center point of C i , denoted as
[0078] (3) Determine the initial cluster members: calculate the distance between the unclustered user and , if , then U k′ is added to C i , that is, C i =C i ∪{U k′}, and let L i represent the number of users in C i ;
[0079] (4) Update the cluster center point: let M i represent the mean shift vector of C i , which can be modeled as: Wherein is a Gaussian kernel function, and σ is a Gaussian kernel function parameter; update the center point of C i based on ;
[0080] (5) Update the cluster members: repeat steps (3) and (4) until M i =0; update C i , and delete C i from C with a distance greater than Rmax a user, i.e., if C i = C i \{U k′} ; update Φ, delete the clustered user, if U k′ ∈ C i , then Φ = Φ \ {U k′} ;
[0081] (6) Determine whether the initial clustering algorithm is terminated: determine whether there is an unclustered U k′ , if , then let i = i + 1, return to step (2), otherwise, execute step (7), and record the total number of user clusters at this time as I;
[0082] (7) Calculate the clustering strategy evaluation function: determine the clustering performance evaluation function based on the intra-cluster and inter-cluster distance, let S i represent the dispersion of C i , defined as the average distance from all users in C i to the center point, which can be modeled as Let represent the distance between C i and C j , which can be modeled as Let F i,j be the similarity measure between C i and C j , which can be modeled as Let F i represent the evaluation function of C i , which can be modeled as Let F represent the clustering strategy evaluation function, which can be modeled as
[0083] (8) Determine the clustering strategy update condition: let F th be a predefined evaluation function threshold, if F ≤ F th , the algorithm is terminated, and the current clustering result is the final result, output the user clustering strategy β i,k ∈ {0, 1}, i.e., if U k ∈ C i , then β i,k = 1, otherwise β i,k = 0; if F > F th , let σ = ασ, 0 < α < 1, return to step (1).
[0084] Step 3: Model the beam illumination variable and power allocation variable.
[0085] Model α m,n,i (t) ∈ {0, 1} as the beam n of S m at time slot t and Ci illumination variable between t time slot S m beam n illuminates C i , then α m,n,i (t) = 1, otherwise α m,n,i (t) = 0;
[0086] modeling power allocation variable, specifically including: modeling beam power allocation vector P m (t) for t time slot S m is P m,1 (t) = [p m,2 (t), p m,N (t), … p T (t)] m,n , where p m (t) represents the transmit power of beam n for t time slot S i,c .
[0087] Step 4: modeling user rate model.
[0088] The message sent by the satellite to the user is divided into public and private parts, and the public part of all user messages in each cluster is uniformly encoded as a public stream, and the private part of each user message is independently encoded as a user private stream; the satellite designs precoding for the public stream and multiple private streams of users in the cluster respectively, and multiplexes and transmits, after receiving the message from the satellite, the users in the cluster decode the public stream and apply successive interference cancellation (SIC) technology, and after removing the public stream from the received signal, each user decodes its corresponding private stream respectively;
[0089] Let s i,c (t) represent the public stream of all user messages in t time slot C i , s i,k (t) represent the private stream of user U i in t time slot C k , let represent the beamforming matrix when S m communicates with C i in t time slot, where w m,i,c (t) and w m,i,k (t) are the precoding vectors of the public stream and the private stream of U k respectively; the signal sent by S m to C i in t time slot can be modeled as The signal received by U i from S k in t time slot C m can be modeled as where represents additive white Gaussian noise;
[0090] Let γ m,i,k,c (t) and γm,i,k (t) represents the t-th time slot C i inner U k received from S m The signal-to-interference-plus-noise ratio of the public stream and the private stream in the signal received from S and Let R m,i,c (t) represent the t-th time slot S m to C i The public stream transmission rate, which can be modeled as Let R m,i,k (t) represent the t-th time slot S m to C i inner U k The private stream transmission rate, which can be modeled as R m,i,k (t) = Blog2(1 + γ m,i,k (t)); Let represent the t-th time slot S m to C i inner U k The sum rate, which can be modeled as
[0091] Step 5: Model the user traffic model and the satellite queue model.
[0092] Assume that the user traffic stream is randomly dynamic, and the traffic amount of each user arriving at each time slot follows a Poisson distribution; let A k (t) represent the traffic amount of U k arriving at the t-th time slot, and the expectation is E[A k (t)] = λ k τ, where λ k represents the average arrival rate of the U k traffic stream;
[0093] Model the satellite queue model, which specifically includes: each satellite is equipped with a user data buffer server, and the data stream randomly arriving from each user can be stored in the corresponding queue; let Q max represent the maximum capacity of the satellite buffer; Q m,k (t) represents the queue length of the U m data stream buffered in S k at the end of the t-th time slot, which can be modeled as:
[0094]
[0095] Step 6: Model the amount of data to be transmitted by the user cluster and the system cost function.
[0096] Let O m,i (t) represent the amount of data to be transmitted by S m to C i at the t-th time slot, which can be modeled as Let r(t) denote the t-th time slot system cost function, which can be modeled as
[0097] Step 7: Model system resource allocation constraints.
[0098] (1) Beam illumination constraint
[0099] Each satellite single beam can serve at most one user cluster in any time slot, then we have:
[0100]
[0101] Each user cluster can be served by at most one satellite single beam in any time slot, then we have:
[0102]
[0103] (2) Beam transmit power constraint
[0104] The total transmit power of multiple beams of each satellite cannot exceed the total transmit power of the satellite in any time slot, then we have:
[0105]
[0106] There is a maximum transmit power limit for each single beam in any time slot, then we have:
[0107]
[0108] The data stream of each user in a cluster cannot be allocated more power than the transmit power of the beam, then we have:
[0109]
[0110] In order to successfully implement SIC while decoding the public stream at the receiving end, the signal strength of the public stream should be greater than the private stream and noise, then we have:
[0111]
[0112] where θ th is the minimum power difference between the public stream and the private stream and noise.
[0113] Step 8: Model system state, action and reward.
[0114] Define the global state space of the t-th time slot as s t = {s t,m |1≤m≤M}, where s t,m = {Q m,k (t), h m,k (t)1≤k≤K} represents the S mThe state; define the joint action space of time slot t as a t ={α m,n,i (t),p m,n (t),W m,i (t)1≤m≤M,1≤n≤N,1≤i≤I}, which includes beam illumination, beam power allocation and intra-cluster beamforming strategies; let r(t) be the reward function of the t-slot system.
[0115] Step 9: Construct and train the phased policy gradient (PPG) network.
[0116] The PPG network comprises a policy network and a value network, where θ and φ represent the parameters of the policy and value networks, respectively. The training process of the PPG network alternates between a policy training phase and a knowledge distillation-assisted training phase. The policy training phase uses the Proximal Policy Optimization (PPO) algorithm to train the agent, while the knowledge distillation-assisted training phase sends useful information from the value network to the policy network. The parameters of the policy and value networks, as well as the experience replay cache, are initialized. Given an initial state s t The agent, based on the output π(a) of the policy network, t |s t ;θ t ) Perform action a t Rewards are obtained after interacting with the environment. t The system transitions to the next state s t+1 , the quadruple (s t ,a t ,r t ,s t+1 Deposit During each network parameter update process, it is from Training samples are extracted from the data, and the parameters of the policy network and the value network are updated alternately.
[0117] Define the loss function of the policy network as follows: in This indicates the difference between the current policy and the previous policy in state s. t Take action a at that time t The probability ratio, π(a) t |s t ;θ t ) indicates the state s under the current policy. t Take action a at that time t The probability, π(a) t |s t ;θ old ) indicates the state s under the previous policy. t Take action a at that time t The probability; A(t) = δ t +(γλ)δ t+1+…+(γλ) X-t+1 δ X-1 δ represents the dominance function for time slot t. t =r t +γV φ (s t+1 )-V φ (s t ) represents the time difference error, V φ (s t ) represents state s t The value function is defined as follows: γ∈(0,1) is the discount factor, λ is a constant, and X is the trajectory length; ε∈(0,1) is the cutoff coefficient, and clip(μ) is the value function. t (θ t (),1-ε,1+ε) are cutoff functions, representing the reduction of μ t (θ t The policy update magnitude is limited to [1-ε, 1+ε] to ensure it is not too large. The stochastic gradient ascent method is used to update the policy network parameters θ, and the update formula can be modeled as θ... t+1 =θ t +η·▽ θ L(θ t ), among which ▽ θ L(θ t ) represents L(θ) t The gradient with respect to parameter θ, where η is the learning rate;
[0118] Define the loss function of the value network as follows: in Let be the target value function for time slot t; the parameters φ of the value network are updated using stochastic gradient descent, and the update formula can be modeled as φ t+1 =φ t -η·▽ φ L(φ t );
[0119] In the knowledge distillation-assisted training phase, the loss function of the policy network is defined as follows: Where L(θ) t ) aux To supplement the loss function, the auxiliary value function V of the policy network is utilized. θ (s t The useful information from the learning value network can be modeled as follows: L(θ t ) joint The second term is the behavioral cloning loss component, β clone For cloning parameters, π(·s) t ;θ old ) represents the strategy before the start of the auxiliary phase, π(·s) t ;θ t) is the current policy, KL(π(·s t ; θ old ), π(·s t ; θ t )) is used to calculate the relative entropy between π(·s t ; θ old ) and π(·s t ; θ t ); the parameters θ and φ of the policy network and the value network are updated respectively by using the stochastic gradient descent method, and the update formulae thereof can be modeled as θ t+1 = θ t - η·▽ θ L(θ t ) joint and φ t+1 = φ t - η·▽ φ L(φ t ).
[0120] Step 10: determining the system resource allocation strategy by using the trained PPG network.
[0121] Under the conditions of meeting the beam illumination and beam transmission power constraints, the resource allocation strategy is optimized and determined to maximize the system cumulative reward, that is:
[0122]
[0123] wherein π(·s ; θ ) is the optimal beam illumination, beam power allocation and intra-cluster beamforming strategy.
[0124] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the same, and although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those of ordinary skill in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the purpose and scope of the technical solutions, and all of them should be covered in the scope of the claims of the present application.
Claims
1. A method of resource allocation for a multi-beam satellite communication system, characterized by, The method specifically comprises the following steps: S1: Modeling a satellite communication system, specifically including: the system comprises multiple multi-beam low-Earth orbit satellites and multiple ground users, making... Indicates the number of low-orbit satellites. Indicates the first Each satellite is equipped with a uniform planar array, with a size of [missing information]. Each satellite can generate simultaneously A beam, making Indicates the number of ground users. Indicates the first There are [number] users, each with a single antenna; let [number] users [become] users. This represents the total power of a single satellite. Indicates the maximum transmit power of a single satellite beam; divides system time into... A series of equal-length time slots, with a time slot length of... ;make express Time slot coordinates express The coordinates; let This represents the total bandwidth of the satellite beam, assuming that all beams of a single satellite occupy bandwidth in an orthogonal manner. One time-frequency resource block; modeling a satellite channel model; S2: determine the user clustering strategy, let denote the satellite beam coverage radius, denote the user set of the th cluster, determine the user clustering strategy based on the mean shift algorithm , if , then , otherwise ; S3: Modeling beam illumination variables and power distribution variables, Let express Time slot beam and The lighting variables between, if Time slot beam illuminate ,but ,on the contrary Modeling Time slot The beam power allocation vector is ,in express Time slot beam The transmission power; S4: modeling a user rate model, specifically comprising: dividing the message sent by the satellite to the user into a public part and a private part, uniformly encoding the public part of all user messages in each cluster as a public stream, and independently encoding the private part of each user message as a user private stream; the satellite designs precoding for the public stream and multiple private streams of the users in the cluster respectively, and performs multiplexing transmission, after receiving the message from the satellite, the users in the cluster decode the public stream and apply successive interference cancellation (SIC) technology, remove the public stream from the received signal, and then each user decodes the corresponding private stream, and determines the user rate based on the signal-to-interference-and-noise ratio (SINR) at the user receiving end; S5: Model the user service model and satellite queue model, specifically including: assuming that user service flows arrive randomly and dynamically, and that the service volume arriving at each user in each time slot follows a Poisson distribution; let... express exist The expected traffic volume arriving in the time slot is ,in express Average arrival rate of service flows; modeling a satellite queue model, specifically including: each satellite is equipped with a user data cache server, which can store the randomly arriving data flows of each user in the corresponding queue, allowing... Indicates the maximum capacity of the satellite buffer; express End of time slot medium cache The queue length of a data stream can be modeled as: ,in express Time slot arrive Internal users The sum and rate; S6: Modeling the amount of data transmitted by user clusters and the system utility function, specifically including: Let express Time slot arrive The amount of data transmitted can be modeled as ;make express The utility function of a time-slotted system can be modeled as follows: ; S7: modeling system resource allocation constraint conditions, including beam illumination constraints, in any time slot, a single beam of each satellite can serve at most one user cluster, and in any time slot, each user cluster can be served by at most one beam of one satellite; beam transmission power constraints, in any time slot, the total transmission power of multiple beams of each satellite cannot exceed the total transmission power of the satellite, in any time slot, the power of a single beam cannot exceed its maximum transmission power, and the power that can be allocated to the data stream of each user cluster cannot exceed the transmission power of the beam, and the signal power of the public stream should be greater than the sum of the powers of the private stream and the noise; S8: modeling system state, action and reward; S9: constructing and training a phase policy gradient (PPG) network; S10: determining the satellite beam illumination, beam power allocation and in-cluster beam forming strategy by using the trained PPG network, specifically comprising: under the conditions of satisfying the beam illumination and beam transmission power constraints, optimizing and determining the resource allocation strategy to maximize the system cumulative reward.
2. The method of claim 1, wherein: In step S1, the satellite channel model is modeled, specifically comprising: Let denote the channel gain of the link between and , which can be modeled as where , , and are the complex gain, Doppler shift, propagation delay and array response vector of the link between and , respectively, is the carrier frequency; can be modeled as where denotes the Ricean fading factor between and , , , and denote the transmit antenna gain of , the receive antenna gain of ,the speed of light and the distance between and , respectively, can be modeled as where denote the array response vectors in x- and y-axis directions, denotes the Kronecker product.
3. The method of claim 2, wherein: In step S2, the user clustering strategy is determined, specifically comprising the following steps: (1) Initialization: Let denote the set of unclustered users, , , , let ; (2) Select initial center point: randomly select a user who is not clustered, and take its position as the initial center point, such as selecting , as the initial center point of , denoted as . ; (3) Determine initial cluster members: compute the distance between unclustered users and , if , add to , , let denote the number of users within ; (4) Update cluster center: Let denote the mean shift vector, which can be modeled as where is a Gaussian kernel function, is a Gaussian kernel function parameter; update the center point of based on (5) Update cluster members: repeat steps (3) and (4) until ; update , delete users in with distance greater than from , if , ; update , delete clustered users, if then ; (6) Determine whether the initial clustering algorithm is terminated: determine whether there is an unclustered user , if , let , return to step (2), otherwise, execute step (7), and record the total number of user clusters at this time as ; (7) Calculate the clustering strategy evaluation function: Determine the clustering performance evaluation function based on intra-cluster and inter-cluster distances, let express The dispersion is defined as The average distance from all users within the network to the center point can be modeled as follows: ;make express and The distance between them can be modeled as ; Let be the similarity measure between ; Let denote the evaluation function of the cluster strategy, which can be modeled as ; let denote the evaluation function of the cluster strategy, which can be modeled as ; let denote the evaluation function of the cluster strategy, which can be modeled as (8) Determine the clustering strategy update condition: Let be the predefined evaluation function threshold. If , the algorithm terminates, and the current clustering result is the final result; if , let , , return to step (1).
4. The method of claim 3, wherein: In step S4, the user rate model is modeled, specifically comprising: make express Time slot A public stream of all user messages within the system. express Time slot Users within Private flow, making express Time slot and The beamforming matrix during communication, where and Public flow and The precoded vector of the private stream; Time slot Send to The signal can be modeled as ; Time slot Inside Received from The signal can be modeled as ,in This represents additive white Gaussian noise. Noise power; Let and denote the public and private flow transmission rates in slot , respectively, and model them as and , respectively. Let denote the public and private flow transmission rates in slot , respectively, and model them as and , respectively. Let denote the public and private flow transmission rates in slot , respectively, and model them as and , respectively. Let denote the public and private flow transmission rates in slot , respectively, and model them as and .
5. The method of claim 4, wherein: In step S7, the system resource allocation constraint conditions are modeled, specifically comprising: (1) beam illumination constraints In any time slot, a single beam of each satellite can serve at most one user cluster, so that: , , In any time slot, each user cluster can be served by at most one beam of one satellite, so that: , ; (2) beam transmission power constraints In any time slot, the total transmission power of multiple beams of each satellite cannot exceed the total transmission power of the satellite, so that: , , In any time slot, a single beam has a maximum transmission power limit, so that: , , The power that can be allocated to the data stream of each user in the cluster cannot exceed the transmission power of the beam, so that: , , In order to decode the public stream at the receiving end and successfully implement SIC, the signal strength of the public stream should be greater than the private stream and the noise, so that: , , wherein is the minimum power difference between public and private streams and noise.
6. The method of claim 5, wherein, In step S8, the modeling of system states, actions, and rewards specifically includes: defining... The global state space of the time slot is ,in, express Time slot State; definition The time-slot joint action space is It includes beam illumination, beam power allocation, and intra-cluster beamforming strategies; for Reward function for time-slotted systems.
7. The method of claim 6, wherein, In step S9, the periodic policy gradient PPG network is constructed and trained, specifically including: the PPG network includes a policy network and a value network, let and represent the parameters of the policy network and the value network respectively; the training process of the PPG network is alternately performed by a policy training stage and a knowledge distillation auxiliary training stage, wherein the policy training stage trains the agent using a proximal policy optimization PPO algorithm, and the knowledge distillation auxiliary training stage can send useful information of the value network to the policy network; the parameters of the policy network and the value network and an experience replay buffer pool are initialized , given an initial state , the agent performs an action according to the output of the policy network , obtains a reward after interacting with the environment, and the system moves to a next state , and a four-tuple is stored in ; during each network parameter update process, training samples are extracted from , and the parameters of the policy network and the value network are alternately updated; The loss function of the policy network is defined as wherein represents the probability ratio of the action taken by the current policy and the previous policy at state represents the probability of taking action at state under the current policy, represents the probability of taking action at state under the previous policy; represents the advantage function at time slot represents the time difference error, represents the value function of state is a discount factor, is a constant, is the length of the trajectory; is a truncation coefficient, is a truncation function, which limits to to ensure that the policy update is not too large; the parameters of the policy network are updated using the stochastic gradient ascent method , and the update formula can be modeled as wherein represents the gradient of the parameter with respect to the parameter , and is a learning rate; The loss function of the value network is defined as wherein is the target value function of the time slot; the parameters of the value network are updated by using a stochastic gradient descent method , and the update formula can be modeled as ; The loss function of the policy network in the knowledge distillation assisted training stage is defined as wherein is the auxiliary loss part, and the auxiliary value function of the policy network is utilized to learn the useful information of the value network, which can be modeled as ; The second term is the behavior cloning loss part, is the cloned parameter, is the policy before the auxiliary stage, is the current policy, is used to calculate the relative entropy between and and ; the parameters of the policy network and the parameters of the value network are updated respectively by using the stochastic gradient descent method and , and the update formula can be modeled as and .
8. The method of claim 7, wherein, In step S10, the beam illumination, beam power allocation and in-cluster beam forming strategy are determined by using the trained PPG network, specifically comprising: under the conditions of satisfying the beam illumination and beam transmission power constraints, optimizing and determining the resource allocation strategy to maximize the system cumulative reward: , wherein , and are optimal beam illumination, beam power allocation and intra-cluster beamforming strategies, respectively.
Citation Information
Patent Citations
End-To-End Beamforming Systems And Satellites
CN107636985A
Beam scheduling and resource allocation method for satellite system
CN114553299A