A two-stage multi-dimensional resource joint allocation method for a satellite communication system

CN116707610BActive Publication Date: 2026-10-09CHINA ACADEMY OF SPACE TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310620139.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2026-10-09
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

该方法在求解最优资源分配策略时,容易陷入局部最优解,且耗时十分严重

Benefits of technology

[0067] (1) This invention achieves dynamic joint adjustment of beam position and beam size through self-supervised learning method. Compared with the supervised learning method in the prior art, it does not require additional manual marking under the premise of meeting the beam request capacity, thus greatly saving computational and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116707610B_ABST
    Figure CN116707610B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of two-stage multidimensional resource joint allocation method of satellite communication system, comprising the following steps: collecting on-board information and user information;U user is clustered into B user cluster, B user cluster is covered by B beam;Self-supervised learning network model is constructed, loss function is established to train model, the user information collected is input to the optimized self-supervised learning network model, and beam-level resource allocation scheme is output;Action network and evaluation network are established based on proximal policy optimization algorithm, user-level resource allocation model is formed and trained, and the optimal user channel allocation decision under the current beam-level resource allocation scheme configuration is obtained.The present application can greatly improve the flexibility of resource allocation, while meeting the dynamic flow demand of each beam, improve the resource utilization rate of satellite, realize "on-demand coverage, on-demand allocation".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite communication technology and relates to a two-stage multi-dimensional resource joint allocation method for satellite communication systems. Background Technology

[0002] With the continuous development of global informatization, satellite communication, with its wide coverage, high transmission rate, and strong resilience, plays an increasingly important role in military, civilian, and commercial fields. In maritime, aviation, and remote land areas with insufficient infrastructure, providing multimedia broadband services to fixed and mobile users via satellite has become an important technological choice. Today, especially high-throughput satellites, employ flexible onboard payloads and multi-beam configurations, and frequency reuse can improve system capacity to some extent. However, given the uneven temporal and spatial distribution of user services, developing more flexible multi-dimensional resource allocation schemes to further improve system resource utilization and user service quality is essential. Existing satellite resource allocation technologies mainly fall into three categories:

[0003] (1) Fixed resource allocation method. This method divides resources equally among beams and maintains the same resource allocation scheme over time. Because it does not take into account the dynamic changes in user service needs, it cannot adapt to the time-varying characteristics of services, resulting in poor system performance. In addition, it does not take into account the uneven distribution of user service needs, resulting in some beams having "resource surplus" while others have "resource shortage", making it difficult to coordinate the utilization of resources among beams.

[0004] (2) Iterative optimization resource allocation methods, such as brute force and combinatorial optimization algorithms. This method usually searches for the global optimum. Because the optimization parameters are often interconnected, it is difficult to adapt to the dynamically changing satellite network environment. In addition, it needs to consider multiple constraints such as channel conditions, co-channel interference, power usage limits, bandwidth usage limits, and beam antenna limitations. The optimization process requires a lot of iterative calculations, which is very computationally expensive and has no practical value.

[0005] (3) Resource allocation methods based on genetic algorithms, such as simulated annealing and hill climbing. These methods are prone to getting trapped in local optima when solving for optimal resource allocation strategies, and are extremely time-consuming. Furthermore, they are ill-suited for large-scale satellite communication scenarios, as their computational complexity generally increases exponentially with the number of satellite beams, resulting in high optimization costs. Summary of the Invention

[0006] The technical problem solved by this invention is to overcome the shortcomings of the prior art and propose a two-stage multi-dimensional joint resource allocation method for satellite communication systems. By using machine learning technology to solve the two-stage resource allocation problem of satellite communication systems, the flexibility of resource allocation can be greatly improved. While meeting the dynamic traffic requirements of each beam, the resource utilization rate of the satellite can be improved, and "on-demand coverage and on-demand allocation" can be achieved.

[0007] The solution of this invention is: a two-stage multi-dimensional resource joint allocation method for a satellite communication system. The satellite communication system has K orthogonal channels, where each beam can use all orthogonal channels within the system, but users within the same beam cannot share the same channel, where K > 1. The method of this invention includes the following steps:

[0008] Step 1: The ground network control center collects satellite information and user information; the user information includes user location coordinates and user requested capacity; the satellite information includes the maximum transmit power of the beam, the maximum transmit gain, and the upper limit of the equivalent isotropic radiated power.

[0009] Step 2: Based on the user location coordinates, use the K-means clustering method to automatically cluster U users into B user clusters. Use one beam to cover users within the same user cluster. The B user clusters are covered by B beams, resulting in the number of beams B, where B > 1.

[0010] Step 3: In the beam-level allocation stage, a self-supervised learning network model is constructed. The collected user information is input into the self-supervised learning network model to obtain the decision factors that determine the beam transmission power and transmission gain. Based on the decision factors, the maximum transmission power of the beam, the maximum transmission gain, and the upper limit of the equivalent isotropic radiation power, the transmission power and transmission gain of B beams are calculated.

[0011] Step 4: In the user-level allocation phase, based on the state s obtained at the target time... t By adopting a user-level resource allocation model, the optimal user channel allocation decision is obtained;

[0012] The user-level resource allocation model is a reinforcement learning network model based on the near-end policy optimization algorithm. In this model, the reward r is defined as... t Define state s to represent the actual allocated capacity of B beams. t This includes the transmit power and transmit gain of B beams, the user request capacity of U users, user location information, and resource reservation information for K orthogonal channels, defining action a. t This is a user channel allocation decision to allocate K orthogonal channels to U users.

[0013] Furthermore, the self-supervised learning network model constructed in step three is a fully connected network model, including an input layer, N hidden layers, and an output layer;

[0014] Input layer: Input data is Among them, u Lon u Lat u alt These represent the user's longitude, latitude, and altitude, respectively. B For the beams accessed by the user, Let B be the user's request capacity, B be the number of beams, and U be the number of users. The input layer has a dimension of U*5, and the output of the input layer is:

[0015]

[0016] δ IN W is the Tanh activation function of the input layer. IN B represents the weights of the input layer. IN For the bias of the input layer;

[0017] Hidden Layer: The first hidden layer takes H1 as input, calculates the weights and biases of each hidden layer network connection, and outputs the result to the next layer, ultimately outputting H. N+1 This process is represented as:

[0018] H2=δ1(W1·H1+B1)

[0019] H3=δ2(W2·H2+B2) ...

[0021] H N+1 =δ N (W N ·H N +B N )

[0022] Among them, W1, W2, ..., W N Let B1, B2, ..., B be the weights of hidden layers 1 through N. N The bias values ​​for hidden layers 1 to N are δ1, δ2, ..., δ. N The Tanh activation function is used for hidden layers 1 to N.

[0023] Output layer: The input data of the output layer is H N+1 The output layer is based on the network connection weights W. OUT and bias B OUT The input data is calculated to obtain the output data as follows:

[0024]

[0025] Where δ OUTThe output layer uses the sigmoid activation function, with the dimension of the output layer being the number of beams B, and the output consists of decision factors for B beams.

[0026] Furthermore, the process of training a self-supervised learning network model includes:

[0027] Establish the loss function J(θ) S ), in the form of:

[0028]

[0029] Where, θ S For parameters of a self-supervised learning network model; The requested capacity of beam b is set to the user request capacity within the coverage area of ​​beam b. The sum of; C aloc The allocation capacity of beam b, where b = 1, ..., B, constitutes the allocation capacity of B beams.

[0030] The parameters of the self-supervised learning network model are iteratively updated using the gradient descent method. The training of the self-supervised learning network model is complete when the loss function value is less than the expected threshold.

[0031] Furthermore, the allocation capacity C of the B beams aloc The calculation process is as follows:

[0032] Step 301: Utilize the transmission power P = {P1, P2, ..., P} of B beams. B} and the transmit gain G={G1,G2,...,G B Based on the link budget, the useful power values ​​E = {E1, E2, ..., E} received by the B beams are calculated. B} and the interference power values ​​I={I1,I2,...,I2} experienced by B beams. B};

[0033] Step 302: Calculate the allocation capacity of B beams according to Shannon's formula; where the allocation capacity of beam b is... W b N is the bandwidth of beam b. σ This represents noise power.

[0034] Furthermore, the transmit power and transmit gain of the B beams are calculated as follows:

[0035] transmit gain

[0036] Transmit power

[0037] in EIRP is the maximum transmit gain of beam b. up This is the upper limit of the equivalent isotropic radiated power, and ensures... Constraints Let ξ be the maximum transmit power of beam b, where b = 1, ..., B; ξ is the scaling factor of the decision factor; [] indicates that the calculation is performed in dB units.

[0038] Furthermore, in step 301, the useful power value E = {E1, E2, ..., E...} B Interference power value I = {I1, I2, ..., I} B The calculation method for} is as follows:

[0039] Useful power E b =P b *H b ;

[0040] In the above formula, P b H is the transmit power of beam b. b H is the channel gain of beam b. b =G b *G R *L b Among them, G b G is the transmit gain of beam b. R L represents the receiving gain of beam b. b For the path loss of beam b;

[0041] Interference power

[0042] Among them, P j H j Let be the transmit power and channel gain of the adjacent beam j of beam b, respectively; β = {β1, β2, ..., β...} B} represents the proportion of edge users in B beams.

[0043] Furthermore, the edge user ratio β of the B beams is β = {β1, β2, ..., β...} B The method for determining} is as follows:

[0044] Using the soft frequency reuse principle, edge users and center users are distinguished. Users located within 3dB of the beam angle are considered center users, while users located within 3dB-1dB of the beam angle are considered edge users. This yields the edge user ratio β = {β1, β2, ..., β} for B beams. B}

[0045] Furthermore, the onboard information also includes the busy / idle status of each channel;

[0046] The state s described in step fourt The principle for determining the user's location is as follows:

[0047] If the user is a central user, the user's location information is 0; if the user is a peripheral user, the user's location information is 1.

[0048] state s t The principle for determining the resource reservation status in the satellite information is as follows: based on the busy / idle status of each channel in the satellite information, if the channel is in a "busy" state, the resource reservation information of the channel is 1; if the channel is in an "idle" state, the resource reservation information of the channel is 0.

[0049] Action a t The output format is: if user u occupies channel k, then a u,k =1, otherwise, a u,k =0; u=1,2,…,U, k=1,2,…,K;

[0050] Reward r t The method for determining the actual allocated capacity of beam b in the image is as follows:

[0051]

[0052] Among them W k W represents the bandwidth of channel k. k =W b / K,

[0053] E k This represents the useful power of channel k. Where P b,u H represents the transmit power of the beam b where user u is located. sat→u This represents the channel gain from the satellite to user u;

[0054] I k This represents the interference power of channel k. P b',u' H represents the transmit power of beam b' where other users u' occupying channel k are located. sat→u' This represents the channel gain from the satellite to user u'.

[0055] Furthermore, the reinforcement learning network model established based on the proximal policy optimization algorithm includes an action network and an evaluation network. The action network is a fully connected network model, with the input state s. t Using the forward propagation method, the user resource allocation strategy p(a) is output. t |s t ), for user resource allocation strategy p(a t |s t The user channel allocation decision is obtained through sampling and used as action a. tThe evaluation network is a fully connected network model, with input state s. t User request capacity, action a t Using the forward propagation method, output the current joint state s. t With action a t The evaluation value.

[0056] Furthermore, the process of training the user-level resource allocation model includes:

[0057] The state s at time t t Input action network, collect action a t The data will show the state s at time t. t User request capacity, action a t Input the evaluation network, collect evaluation data, and calculate the reward r at time t. t Update time t, repeat the process at least a preset number of times, and store the collected state, action, reward, and evaluation data as training data in the data cache;

[0058] Construct the action network loss function, in the form of:

[0059]

[0060] Where, θ A For action network model parameters; The ratio of resource allocation strategies for old and new users; This represents the user resource allocation strategy before the action network model parameters are updated; This represents the user resource allocation strategy after the network model parameters are updated; the specific form of the clip(·) function is clip(ρ(θ)). A ),1-ε,1+ε), representing ρ(θ) A The parameter is reduced to the range [1-ε, 1+ε], where ε is a hyperparameter. The advantage function is calculated using the GAE method in the near-end policy optimization algorithm;

[0061] Construct an evaluation network loss function in the following form:

[0062]

[0063] Where, θ C To evaluate the network model parameters, To evaluate the evaluation value of the network output, Let γ be the target value to be achieved, where γ is the discount factor, t' is the target time, and r is the target value. t' The reward at time t';

[0064] Input training data, train the action network using the action network loss function, and follow... Iteratively update the action network model parameters; train the evaluation network using the evaluation network loss function, according to... Iterative updates to the evaluation network model parameters;

[0065] When the reward is satisfied At that time, the user-level resource allocation model training is complete, where λ user The preset user-level resource allocation precision. For the actual allocated capacity of beam b, The requested capacity for beam b.

[0066] The advantages of this invention compared to the prior art are:

[0067] (1) This invention achieves dynamic joint adjustment of beam position and beam size through self-supervised learning method. Compared with the supervised learning method in the prior art, it does not require additional manual marking under the premise of meeting the beam request capacity, thus greatly saving computational and time costs.

[0068] (2) The present invention combines reinforcement learning and soft frequency reuse methods to constrain sub-channel allocation strategies. Compared with reinforcement learning methods in the prior art, it can avoid co-channel interference and meet user request capacity while significantly reducing the action space size, thereby accelerating the convergence of reinforcement algorithms. Attached Figure Description

[0069] Figure 1 This is a flowchart of the two-stage multi-dimensional resource joint allocation method according to an embodiment of the present invention;

[0070] Figure 2 This is a schematic diagram of a self-supervised learning network model according to an embodiment of the present invention;

[0071] Figure 3 This is a schematic diagram of a user-level resource allocation model according to an embodiment of the present invention. Detailed Implementation

[0072] The present invention will be further described below with reference to the embodiments.

[0073] Example 1

[0074] Figure 1 This is a flowchart of the two-stage multi-dimensional resource joint allocation method for the satellite communication system of the present invention. The satellite communication system has a total of K orthogonal channels, where each beam can use all orthogonal channels within the system, but users within the same beam cannot share the same channel. The method of the present invention includes the following steps:

[0075] Step 1: The ground network control center collects satellite information and user information. The user information includes the user's location coordinates and requested capacity. The satellite information includes the satellite's total transmit power, maximum transmit power of the beam, maximum transmit gain, upper limit of equivalent isotropic radiated power, and busy / idle status of each channel.

[0076] Step 2: Based on the user location coordinates, use the K-means clustering method to automatically cluster U users into B user clusters. The number of user clusters is the number of beams. Use one beam to cover users within the same user cluster. The B user clusters are covered by B beams, resulting in the number of beams B.

[0077] Step 3: In the beam-level resource allocation stage, a self-supervised learning network model is constructed. The user location coordinates and user requested capacity within the coverage area of ​​B beams are collected and input into the self-supervised learning network model. Through the forward propagation method, B decision factors that can determine the power and gain of the B beams are output.

[0078] like Figure 2 As shown, the self-supervised learning network model is a fully connected network model, including an input layer, N hidden layers, and an output layer;

[0079] Input layer: Input data is Among them, u Lon u Lat u alt These represent the user's longitude, latitude, and altitude, respectively. B For the beams accessed by the user, Let B be the user's request capacity, B be the number of beams, and U be the number of users. The input layer has a dimension of U*5, and the output of the input layer is:

[0080]

[0081] δ IN W is the Tanh activation function of the input layer. IN B represents the weights of the input layer. IN For the bias of the input layer;

[0082] Hidden Layers: There are N hidden layers in total. The first hidden layer takes H1 as input and calculates its value according to the weights and biases of each hidden layer's network connections, then outputs it to the next layer, finally outputting H. N+1 This process is represented as:

[0083] H2=δ1(W1·H1+B1)

[0084] H3=δ2(W2·H2+B2) ...

[0086] HN+1 =δ N (W N ·H N +B N )

[0087] Among them, W1, W2, ..., W N Let B1, B2, ..., B be the weights of hidden layers 1 through N. N The bias values ​​for hidden layers 1 to N are δ1, δ2, ..., δ. N The Tanh activation function is used for hidden layers 1 to N.

[0088] Output layer: The input data of the output layer is H N+1 The output layer is based on the network connection weights W. OUT and bias B OUT The input data is calculated to obtain the output data as follows:

[0089]

[0090] Where δ OUT The output layer uses the sigmoid activation function, with the dimension of the output layer being the number of beams B, and the output consists of decision factors for B beams.

[0091] Step 4: Train the self-supervised learning network model:

[0092] First, sum the user request capacity within the coverage area of ​​each of the B beams to obtain the total request capacity of the B beams. The allocation capacity of B beams is calculated based on on-board information and output decision factors. A loss function J(θ) is established based on the mean square error between the allocated capacity and requested capacity of the B beams. S ).

[0093] Then, the gradient is calculated based on the backpropagation method, and the parameters of the self-supervised learning network model are iteratively updated using the gradient descent method until the loss function value is less than the expected threshold, at which point the training of the self-supervised learning network model is complete.

[0094] Finally, the collected user information is input into the trained self-supervised learning network model to obtain the decision factors. Based on the decision factors, the maximum transmit power of the beam, the maximum transmit gain, and the upper limit of the equivalent isotropic radiated power, the transmit power and transmit gain of B beams are calculated.

[0095] Specifically, calculate the beam allocation capacity. include:

[0096] Step 401: Based on the decision factors of the B beams output in Step 3, and using the constraints of the maximum transmit power, maximum transmit gain, and upper limit of the equivalent isotropic radiated power of each beam, calculate the transmit power P = {P1, P2, ..., P...} of each beam. B} and the transmit gain G={G1,G2,...,G B}, as shown below:

[0097] transmit gain

[0098] Transmit power

[0099] in EIRP is the maximum transmit gain of beam b. up This is the upper limit of the equivalent isotropic radiated power, and ensures... Constraints Let b be the maximum transmit power of beam b, where b = 1, ..., B; ξ is the scaling factor of the decision factor; [·] indicates that the calculation is performed in dB units.

[0100] Step 402: Utilize the transmission power P = {P1, P2, ..., P} of the B beams. B} and the transmit gain G={G1,G2,...,G B Based on the link budget, the useful power values ​​E = {E1, E2, ..., E} received by the B beams are calculated. B} and the interference power values ​​I={I1,I2,...,I2} experienced by B beams. B}:

[0101] (1) First, calculate the channel gain: H b =G b *G R *L b ;

[0102] Among them, H b G is the channel gain of beam b. b G is the transmit gain of beam b. R L represents the receiving gain of beam b. b Let L be the path loss of beam b. In calculating the path loss L of beam b... b First, the path length can be obtained based on the beam pointing, and then the path loss L of beam b can be obtained based on the path length. b In this invention, the beam direction is determined to be from the satellite to the center of the corresponding user cluster.

[0103] (2) Next, calculate the useful power: E b =P b *H b ;

[0104] Among them, P b Let be the transmit power of beam b.

[0105] (3) Then, the interference power is calculated. The specific method is as follows: the soft frequency reuse principle is used to distinguish between edge users and center users. That is, users located within 3dB angle of the beam are considered center users, and users located within 3dB-1dB angle of the beam are considered edge users. Thus, the proportion of edge users of B beams β={β1,β2,...,β B Assuming the interference originates from edge users of other beams under the same satellite, the interference power is calculated as follows:

[0106]

[0107] Among them, P j H j , respectively, are the transmit power and channel gain of beam j, which are adjacent beams of beam b.

[0108] Step 403: Utilize the useful power values ​​E = {E1, E2, ..., E} received by each beam B Interference power value I = {I1, I2, ..., I} B The allocation capacity of B beams is calculated using Shannon's formula. Among them, the allocation capacity of beam b is W b N is the bandwidth of beam b. σ This represents noise power.

[0109] Request capacity based on B beams The allocation capacity of B beams The loss function is constructed as follows: Where, θ S These are the parameters for a self-supervised learning network model.

[0110] The network model parameters are updated using gradient descent to update the parameters of the self-supervised learning network model.

[0111] Finally, the beam-level resource allocation scheme obtained at time t includes: the number of beams, beam pointing, and the transmit power and transmit gain of each beam. This beam-level resource allocation scheme is then transmitted to the satellite to configure the payload.

[0112] Step 5: In the user-level resource allocation phase, based on the state s obtained at the target time... t By adopting a user-level resource allocation model, the optimal user channel allocation decision is obtained.

[0113] The user-level resource allocation model is a reinforcement learning network model based on the proximal policy optimization algorithm, including an action network and an evaluation network. The reward r is defined in the reinforcement learning network model. t Define the actual allocated capacity of B beams; define state s. t This includes the transmit power and transmit gain of B beams, the user requested capacity of U users, user location information, and resource reservation information for K orthogonal channels; defining action a. t This is a user channel allocation decision to allocate K orthogonal channels to U users.

[0114] Wherein, state s t The principle for determining the user's location is as follows: if the user is a central user, the user's location information is 0; if the user is a peripheral user, the user's location information is 1.

[0115] state s t The principle for determining the resource reservation status in the satellite information is as follows: based on the busy / idle status of each channel in the satellite information, if the channel is in a "busy" state, the resource reservation information of the channel is 1; if the channel is in an "idle" state, the resource reservation information of the channel is 0.

[0116] Define action a t The output form, i.e., how to allocate K channels to U users, is as follows: If user u occupies channel k, then a u,k =1, otherwise, a u,k =0.

[0117] Define reward r t The actual allocated capacity of beam b is given by B beams. Specifically, the actual allocated capacity of beam b is expressed as follows:

[0118]

[0119] Among them, W k W represents the bandwidth of channel k. k =W b / K;

[0120] E k This represents the useful power of channel k. Where P b,u H represents the transmit power of the beam b where user u is located. sat→u This represents the channel gain from the satellite to user u;

[0121] I k This represents the interference power of channel k. P b',u' H represents the transmit power of beam b' where other users u' occupying channel k are located. sat→u'This represents the channel gain from the satellite to user u'; therefore, I k This represents the interference power caused by other users u' occupying channel k.

[0122] like Figure 3 As shown, the user-level resource allocation model of the present invention is specifically as follows: The action network is established as a fully connected network model, and the input state s... t Using the forward propagation method, output the user resource allocation strategy p(a) t |s t ), for user resource allocation strategy p(a t |s t The user channel allocation decision is obtained through sampling and used as action a. t The evaluation network is also a fully connected network model, with the input state s. t User request capacity, action a t Using the forward propagation method, output the current joint state s. t With action a t The evaluation value.

[0123] After the user-level resource allocation model is determined, it is trained, including the following process:

[0124] Step 501: Set the state s at time t. t Input action network, collect action a t Data; the state s at time t t User request capacity, action a t Input the evaluation network to collect evaluation data; and calculate the reward r at time t. t At each new time t, repeat the above process at least 500 times, storing the collected state, action, reward, and evaluation data as training data in the data cache.

[0125] Step 502: Construct the action network loss function as follows:

[0126]

[0127] Where, θ A For action network model parameters; The ratio of resource allocation strategies for old and new users; This represents the user resource allocation strategy before the action network model parameters are updated; This represents the user resource allocation strategy after the network model parameters are updated; the specific form of the clip(·) function is clip(ρ(θ)). A ),1-ε,1+ε), representing ρ(θ) A The value is reduced to the range [1-ε, 1+ε], where ε is a hyperparameter, which can generally be set to 0.2; The advantage function is calculated using the GAE method in the near-end policy optimization algorithm.

[0128] Step 503: Construct the evaluation network loss function as follows:

[0129]

[0130] Where, θ C To evaluate the network model parameters, To evaluate the evaluation value of the network output, Let γ be the target value to be achieved, where γ is the discount factor, t' is the target time, and r is the target value. t' The reward for time t'.

[0131] Step 504: Input training data and train the action network using the action network loss function, according to... Iteratively update the action network model parameters; train the evaluation network using the evaluation network loss function, according to... Iteratively update the parameters of the evaluation network model.

[0132] Step 505: When the reward is satisfied At that time, the user-level resource allocation model training is complete, where λ user The preset user-level resource allocation precision. For the actual allocated capacity of beam b, The requested capacity for beam b.

[0133] After the user-level resource allocation model is trained, the state data acquired at the target time is input, and the optimal user channel allocation decision is obtained using the trained user-level resource allocation model.

[0134] The trained user-level resource allocation model can also be fine-tuned online. The method is as follows: repeatedly collect state, action, reward, and evaluation value data, and at regular intervals, use the newly collected data to execute the training process described in steps 504 to 505 to achieve online fine-tuning of the user-level resource allocation model parameters in order to output a real-time user-level resource allocation scheme.

[0135] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.

Claims

1. A two-stage, multi-dimensional resource joint allocation method for a satellite communication system, wherein the satellite communication system has a total of There are 3 orthogonal channels, where each beam can use all orthogonal channels within the satellite communication system, but users within the same beam cannot share the same channel. Its characteristics are, Includes the following steps: Step 1: The ground network control center collects satellite information and user information; the user information includes user location coordinates and user requested capacity; the satellite information includes the maximum transmit power of the beam, the maximum transmit gain, and the upper limit of the equivalent isotropic radiated power. Step 2: Based on the user's location coordinates, use the K-means clustering method to cluster the data. Individual users are automatically grouped into Each user cluster uses a single beam to cover users within the same cluster. A user cluster is composed of Coverage is achieved using multiple beams, resulting in the number of beams. , ; Step 3: In the beam-level allocation stage, a self-supervised learning network model is constructed. The collected user information is input into the self-supervised learning network model to obtain the decision factors that determine the beam transmit power and transmit gain. Based on the decision factors, the maximum transmit power of the beam, the maximum transmit gain, and the upper limit of the equivalent isotropic radiated power, the following calculations are performed: The transmit power and transmit gain of each beam; Step 4: In the user-level allocation phase, based on the status obtained at the target time... By adopting a user-level resource allocation model, the optimal user channel allocation decision is obtained; The user-level resource allocation model is a reinforcement learning network model based on the near-end policy optimization algorithm, in which a reward is defined. Define the state for the actual allocated capacity of B beams. include The transmit power and transmit gain of each beam User request capacity and user location information for each user. Resource reservation information for each orthogonal channel, defining actions. To be Each orthogonal channel is assigned to User channel allocation decision for each user; The process of training a self-supervised learning network model includes: Establish loss function The format is: in, For parameters of a self-supervised learning network model; For beam The requested capacity, which is the value of the beam. User request capacity within the coverage area sum; For beam Allocation capacity, , The allocation capacity of each beam constitutes ; The parameters of the self-supervised learning network model are iteratively updated using the gradient descent method. The training of the self-supervised learning network model is completed when the loss function value is less than the expected threshold. beam allocation capacity The calculation process is as follows: Step 301, using Transmit power of each beam With transmit gain Calculated based on link budget Useful power value received by each beam ,as well as Interference power value of each beam ; Step 302: Calculate according to Shannon's formula. The allocated capacity of each beam; where, beam The allocated capacity is , For beam bandwidth, Noise power; Useful power value in step 301 Interference power value The calculation method is as follows: Useful power ; In the above formula, For beam The transmission power, For beam Channel gain, ,in, For beam transmit gain, For beam Receiver gain For beam Path loss; Interference power ; in, , Beams adjacent beams Transmit power and channel gain; for The proportion of edge users per beam; The Edge user ratio per beam The method for determining it is as follows: The soft frequency reuse principle is used to distinguish between edge users and center users. Users located within 3dB of the beam angle are considered center users, while users located within 3dB-1dB of the beam angle are considered edge users, thus obtaining... Edge user ratio per beam .

2. The two-stage multi-dimensional resource joint allocation method for a satellite communication system according to claim 1, characterized in that, The self-supervised learning network model constructed in step three is a fully connected network model, including an input layer, One hidden layer and one output layer; Input layer: Input data is ,in, , , These represent the user's longitude, latitude, and altitude, respectively. The beam that the user accesses. For the user's request capacity, For the number of beams, The input layer has the following dimensions, representing the number of users. The output of the input layer is: The Tanh activation function for the input layer. The weights of the input layer, For the bias of the input layer; Hidden layer: The first hidden layer input is The calculations are performed based on the weights and biases of the network connections in each hidden layer, and the results are output to the next layer. The final output is... This process is represented as: in, Hidden layer 1~ Layer weights, Hidden layer 1~ Layer bias, Hidden layer 1~ The Tanh activation function of the layer; Output layer: The input data for the output layer is The output layer is based on the weights of the network connections. and bias The input data is calculated to obtain the output data as follows: in The output layer uses the sigmoid activation function, and the dimension of the output layer is the number of beams. The output is Decision factors for each beam .

3. The two-stage multi-dimensional resource joint allocation method for a satellite communication system according to claim 1, characterized in that, The The transmit power and transmit gain of each beam are calculated as follows: transmit gain ; Transmit power ; in For beam Maximum transmit gain, This is the upper limit of the equivalent isotropic radiated power, and ensures... Constraints For beam Maximum transmission power, ; The scaling factor for decision factors; This indicates that the calculation is performed in dB.

4. The two-stage multi-dimensional resource joint allocation method for a satellite communication system according to claim 1, characterized in that, Onboard information also includes the busy / idle status of each channel; The state described in step four The principle for determining the user's location is as follows: If the user is a central user, the user's location information is: If the user is a marginal user, the user's location information is: ; state The principle for determining resource reservation status is as follows: based on the busy / idle status of each channel in the on-board information, if a channel is in a "busy" state, the resource reservation information for that channel is... If the channel is in an "idle" state, the resource reservation information for that channel is... ; action The output format is: if the user Channel occupied Then there is ,otherwise, ; , ; award The method for determining the actual allocated capacity of beam b in the image is as follows: in Indicates channel bandwidth, , Indicates channel Useful power, ,in Indicates user Beam The transmission power, Indicates satellite to user Channel gain; Indicates channel Interference power, , Indicates other channels are occupied users Beam The transmission power, Indicates satellite to user Channel gain.

5. The two-stage multi-dimensional resource joint allocation method for a satellite communication system according to claim 4, characterized in that, The reinforcement learning network model based on the proximal policy optimization algorithm includes an action network and an evaluation network. The action network is a fully connected network model, and the input state... Using the forward propagation method, output the user resource allocation strategy. User resource allocation strategy The user channel allocation decision is obtained through sampling and used as an action. The evaluation network is a fully connected network model, and the input state is... User request capacity and actions Using the forward propagation method, output the current joint state. With action The evaluation value.

6. The two-stage multi-dimensional resource joint allocation method for a satellite communication system according to claim 5, characterized in that, The process of training a user-level resource allocation model includes: The state at time t Input action network to collect actions The data will show the state at time t. User request capacity and actions Input the evaluation network, collect evaluation data, and calculate the reward at time t. Update time t, repeat the process at least a preset number of times, and store the collected state, action, reward, and evaluation data as training data in the data cache; Construct the action network loss function, in the form of: ; in, For action network model parameters; The ratio of resource allocation strategies for old and new users; This represents the user resource allocation strategy before the action network model parameters are updated; , which represents the user resource allocation strategy after the network model parameters are updated; The specific form of the function is ,express Laid off Within the range, For hyperparameters, The advantage function is calculated using the GAE method in the near-end policy optimization algorithm; Construct an evaluation network loss function in the following form: ; in, To evaluate the network model parameters, To evaluate the evaluation value of the network output, The target value to be achieved is, where As a discount factor, For the target time, for Momentary rewards; Input training data, train the action network using the action network loss function, and follow... Iteratively update the action network model parameters; train the evaluation network using the evaluation network loss function, according to... Iterative updates to evaluate network model parameters; When the reward is satisfied At that time, the user-level resource allocation model training was completed, among which The preset user-level resource allocation precision. For the actual allocated capacity of beam b, For beam The requested capacity.

Citation Information

Patent Citations

  • Multi-beam satellite communication system resource allocation method considering user association, sub-channel allocation and beam association

    CN115065384A