A Dynamic Beam Hopping and Beam Bandwidth Allocation Method Based on Multi-Agent Reinforcement Learning

By adopting the dynamic beam hopping and beam bandwidth allocation method of multi-agent reinforcement learning in satellite communication, the problem of satellite beam resource allocation in the prior art is solved, and the flexible allocation of resources in three dimensions of time, space and frequency is realized, and business adaptability and performance are improved.

CN114189939BActive Publication Date: 2025-06-03TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111527204.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-06-03
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively allocate satellite beam resources in satellite communications, especially in the three dimensions of time, space and frequency. The convergence speed of traditional algorithms is slow and it is difficult to adapt to time-varying business needs. Machine learning algorithms lack bandwidth-level freedom.

Method used

The dynamic beam jump and beam bandwidth allocation method based on multi-agent reinforcement learning is adopted. By constructing a multi-agent reinforcement learning simulation system model, multiple reinforcement learning agents are used to be responsible for the irradiation direction and bandwidth allocation of the beam respectively, to achieve flexible allocation of resources in three dimensions.

Benefits of technology

Effectively utilize the three degrees of freedom of space-time frequency of satellite beam resources, reduce the complexity of the agent's decision space, improve the throughput rate and delay fairness, and realize real-time beam hopping scheduling and beam bandwidth allocation by time slot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114189939B_ABST
    Figure CN114189939B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic beam hopping and bandwidth allocation method based on multi-agent reinforcement learning. Each agent is responsible for the irradiation direction of a satellite beam or the bandwidth size of the beam, and the agents cooperate with each other to complete the dynamic beam hopping and bandwidth allocation task. This method not only effectively utilizes the three degrees of freedom of time, space, and frequency of the beam resources, but also reduces the complexity of the decision space compared with single-agent reinforcement learning. It has excellent performance in terms of throughput and delay fairness, and has a certain robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of resource allocation in satellite communication, and relates to a method for allocating satellite beam resources in three dimensions of time, space, and frequency. In particular, it relates to a dynamic hopping beam and beam bandwidth allocation method based on multi-agent reinforcement learning. Background Art

[0002] Due to the uneven geographical distribution and time-varying characteristics of ground users, the service sizes of each cell under satellite coverage vary significantly. How to flexibly and effectively allocate satellite resources and match them with non-uniform services is a major problem in the field of satellite communication. As an emerging technology in satellite communication, the hopping beam technology makes full use of the spatio-temporal degrees of freedom of the beam, providing great flexibility for satellite resource allocation. The hopping beam technology divides the satellite coverage area into multiple cells, but only irradiates some of them in one time slot. The cells irradiated by the satellite in different time slots are different, which is called the hopping beam pattern. How to reasonably plan the hopping beam pattern and allocate the bandwidth of each beam is a challenging problem.

[0003] Traditional hopping beam algorithms mainly use heuristic algorithms. These algorithms have a slow convergence speed and cannot well adapt to time-varying service requirements, so it is difficult to achieve real-time hopping beam scheduling. Although some algorithms based on convex optimization have been proposed, they ignore the co-channel interference between beams. Once co-channel interference is considered, the problem will become non-convex, resulting in algorithm failure. With the development of machine learning, applying machine learning to resource allocation has become a hot topic. However, existing machine learning-based hopping beam algorithms generally only allocate resources in the beam time slot dimension and lack degrees of freedom at the bandwidth level. Moreover, existing machine learning algorithms basically adopt the single-agent reinforcement learning method. As the number of beams increases, the decision space of the single agent will increase exponentially, and the scale of the neural network will be greatly improved.

[0004] In summary, there are two problems in the prior art. On the one hand, traditional heuristic hopping beam algorithms have a slow convergence speed and cannot well adapt to time-varying service requirements, making it difficult to achieve real-time hopping beam scheduling. On the other hand, machine learning-based hopping beam algorithms generally only allocate resources in the beam time slot dimension and lack degrees of freedom at the bandwidth level.

[0005] Object of the Invention

[0006] The object of the present invention is to solve the above problems existing in the prior art, and propose a dynamic hopping beam and beam bandwidth allocation method based on multi-agent reinforcement learning, which can not only effectively utilize the three degrees of freedom of time, space, and frequency of beam resources, but also reduce the complexity of the agent decision space compared with single-agent reinforcement learning. Content of the Invention

[0007] The present invention provides a dynamic beam hopping and beam bandwidth allocation method based on multi-agent reinforcement learning for resource allocation of satellite beams in three dimensions of time, space and frequency. The method includes the following steps:

[0008] Step 1, obtain satellite communication system parameters, including the number of beams K, the number of ground cells N, the total satellite transmission power, the satellite altitude, the number of satellite bandwidth resource blocks, and the bandwidth of each satellite broadband resource block;

[0009] Step 2, construct a multi-agent reinforcement learning simulation system model, and perform offline training on the constructed simulation system model;

[0010] Step 3, deploy the trained simulation system model in step 2 to the operation control center or on-board payload;

[0011] Step 4, the operation control center or on-board payload inputs the real-time service queue size of each cell in each time slot into the trained simulation system model deployed thereon, so as to obtain the irradiation position of each beam and the allocated bandwidth size in this time slot.

[0012] Preferably, the multi-agent reinforcement learning simulation system model in step 2 includes a plurality of reinforcement learning agents, a plurality of experience pools, and a satellite beam hopping simulation environment.

[0013] Preferably, in the multi-agent reinforcement learning simulation system model in step 2, one satellite beam corresponds to two agents, which are respectively responsible for the irradiation direction and bandwidth allocation of the satellite beam. Each of the two agents includes two neural networks and an experience pool, wherein the two neural networks are a target network and a Q network.

[0014] Preferably, the working process of the multi-agent reinforcement learning simulation system model in step 2 includes the following sub-steps:

[0015] Step S21, initialize the parameters of 2K neural networks and 2K experience pools; initialize the service arrival rate of each cell; set the environmental initial state as the service size of each ground cell in the first time slot; set the greedy coefficient ε to 1;

[0016] Step S22, start to execute a loop process, which includes M large loops, and each large loop includes T small loops; wherein, the t-th small loop process includes the following sub-processes:

[0017] i. Observe the global state s with each agent t , defined as where represents the total service size of the n-th cell in the t-th time slot;

[0018] ii. Input the global state s t into the Q-network of each agent. The k-th agent makes a decision according to the ε-greedy algorithm This decision refers to the irradiation position of the beam or the number of bandwidth resource blocks;

[0019] iii. Calculate the transmission capacity of each satellite beam according to the decisions of each agent and the link budget, then transmit the traffic of the cell irradiated in the t-th time slot, and update the traffic volume of each cell in the (t + 1)-th time slot to obtain the environmental state variable s t+1 ;

[0020] iv. Calculate the traffic throughput Th t and the inter-cell delay fairness F t in the t-th time slot, and then obtain the reward function rt, which is expressed as shown in the following formula (1):

[0021] r t = βTh t -(1 - β)F t (1),

[0022] where β is a number between 0 and 1, used to balance the traffic throughput Th t and the inter-cell delay fairness F t between the weights;

[0023] v. Each agent stores the experience in the t-th time slot into the corresponding experience pool;

[0024] vi. Each agent randomly extracts M experiences from its corresponding experience pool, calculates the mean square error loss, and uses the Adam algorithm to train the parameters of its own Q-network;

[0025] Step S23: Every C small loops, the agent copies the parameters of its own Q-network to the target network;

[0026] Step S24: Every 1 large loop, reduce the greedy coefficient ε;

[0027] Step S25: Every 1 large loop, reset the traffic arrival rate of each cell and initialize the environmental state.

[0028] Preferably, in the multi-agent reinforcement learning simulation system model described in step 2, the reinforcement learning algorithm used includes but is not limited to the DQN algorithm, Double DQN algorithm, or A3C algorithm.

[0029] Preferably, the agent in step 2 adopts a neural network structure, including but not limited to a fully connected network, a convolutional neural network, and a recurrent neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart of the dynamic hopping beam and beam bandwidth allocation method of the present invention.

[0031] Figure 2 It is a working flowchart of the multi-agent reinforcement learning simulation system model of the present invention.

[0032] Figure 3 It is a satellite hopping beam schematic diagram of Embodiment 1 of the present invention;

[0033] Figure 4 It is a multi-agent reinforcement learning framework of Embodiment 1 of the present invention;

[0034] Figure 5 It is a satellite hopping beam schematic diagram of Embodiment 2 of the present invention;

[0035] Figure 6 It is a multi-agent reinforcement learning framework of Embodiment 2 of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0036] The specific implementation method of the present invention will be further described below in conjunction with the drawings and embodiments. It should be noted that this part is only used to exemplarily illustrate the present invention. The embodiments given are all preferred implementation manners of the present invention, and should not be regarded as a limitation on the protection scope of the present invention. Any variant or alternative technical solution that does not deviate from the gist of the present invention falls within the protection scope of the present invention.

[0037] The present invention proposes a dynamic hopping beam and bandwidth allocation algorithm based on multi-agent reinforcement learning, which is used for resource allocation of satellite beams in three dimensions of time, space, and frequency, and is applicable to scenarios with different numbers of beams and cells.

[0038] Figure 1 It is a flowchart of the dynamic hopping beam and beam bandwidth allocation method of the present invention. As Figure 1 shown, a dynamic hopping beam and beam bandwidth allocation method based on multi-agent reinforcement learning includes the following steps:

[0039] Step 1, obtain satellite communication system parameters, including the number of beams K, the number of ground cells N, the total satellite transmission power, the satellite altitude, the number of satellite bandwidth resource blocks, and the bandwidth of each satellite broadband resource block;

[0040] Step 2, construct a multi-agent reinforcement learning simulation system model, and perform offline training on the constructed simulation system model;

[0041] Step 3: Deploy the trained simulation system model in Step 2 to the operation control center or on-board payload;

[0042] Step 4: The operation control center or on-board payload inputs the real-time service queue size of each cell in each time slot into the trained simulation system model deployed thereon, so as to obtain the irradiation position of each beam and the allocated bandwidth size in this time slot.

[0043] Among them, the multi-agent reinforcement learning simulation system model in Step 2 includes multiple reinforcement learning agents, multiple experience pools, and a satellite hopping beam simulation environment; each satellite beam corresponds to two agents, which are respectively responsible for the irradiation direction and bandwidth allocation of the satellite beam. Each agent in the two agents respectively includes two neural networks and an experience pool, and the two neural networks are a target network and a Q network.

[0044] Figure 2 It is the working flowchart of the multi-agent reinforcement learning simulation system model of the present invention. As shown in the figure, the working process of the multi-agent reinforcement learning simulation system model includes the following sub-steps:

[0045] Step S21: Initialize 2K neural network parameters and 2K experience pools; initialize the service arrival rate of each cell; set the environmental initial state to the service size of each ground cell in the first time slot; set the greedy coefficient ε to 1;

[0046] Step S22: Start to execute a loop process, which includes M large loops, and each large loop contains T small loops; among them, the t-th small loop process includes the following sub-processes:

[0047] i. Each agent observes the global state s t , defined as where represents the total service size of the n-th cell in the t-th time slot;

[0048] ii. Input the global state s t into the Q network of each agent, and the k-th agent makes a decision according to the ε-greedy algorithm This decision refers to the irradiation position of the beam or the number of bandwidth resource blocks;

[0049] iii. According to the decision of each agent and the link budget, calculate the transmission capacity of each satellite beam, then transmit the service of the irradiated cell in the t-th time slot, and update the traffic volume of each cell in the (t + 1)-th time slot to obtain the environmental state variable s t+1 ;

[0050] iv. Calculate the service throughput Th of the t-th time slot according to the service transmission situation of the t-th time slot t and the inter-cell delay fairness F t , and then obtain the reward function r t , which is expressed as follows:

[0051] r t =βTh t -(1-β)F t

[0052] where β is a number between 0 and 1, used to balance the service throughput Th of the t-th time slot t and the inter-cell delay fairness F t between the weights.

[0053] v. Each agent stores the experience of the t-th time slot into the corresponding experience pool;

[0054] vi. Each agent randomly extracts M experiences from its corresponding experience pool, calculates the mean square error loss, and uses the Adam algorithm to train the parameters of its own Q network;

[0055] Step S23: Every C small cycles, the agent copies its own Q network parameters to the target network;

[0056] Step S24: Every 1 large cycle, reduce the greedy coefficient ε;

[0057] Step S25: Every 1 large cycle, reset the service arrival rate of each cell and initialize the environmental state.

[0058] Among them, in specific implementation, the reinforcement learning algorithm adopted by the multi-agent reinforcement learning simulation system model includes but is not limited to the DQN algorithm, the Double DQN algorithm, or the A3C algorithm. The agent adopts a neural network structure, including but not limited to a fully connected network, a convolutional neural network, and a recurrent neural network.

[0059] The following are two specific embodiments of the dynamic beam hopping and beam bandwidth allocation method of the present invention.

[0060]

Embodiment 1

[0061] As Figure 3 shown, assume that the number of satellite beams is 4 and the number of ground cells is 19. The entire multi-agent reinforcement learning framework is as Figure 4 shown. The multi-agent reinforcement learning model training is completed through the following steps:

[0062] a) Initialize eight neural network parameters and eight experience pools, each with a size of 10,000; initialize the traffic arrival rate of each cell, assuming that the traffic arrival rate of each cell is uniformly distributed in the range of 50M to 350M; set the initial state of the environment to the traffic size of each terrestrial cell in the first time slot; set the greedy coefficient ε to 1.

[0063] b) Start the loop process, which includes 10,000 large loops, and each large loop contains 200 small loops. The process of the t-th small loop is as follows:

[0064] i. Each agent observes the global state, defined as where represents the total traffic size of the n-th cell in the t-th time slot.

[0065] ii. Input the global state s t into the Q-network of each agent. The k-th agent makes a decision according to the ε-greedy algorithm This decision refers to the irradiation position of the beam or the number of bandwidth resource blocks allocated.

[0066] iii. According to the decisions of each agent, calculate the transmission capacity of each beam, then transmit the traffic of the cells irradiated in the t-th time slot, and update the traffic volume of each cell in the t + 1-th time slot to obtain the environmental state variable s t+1 .

[0067] iv. According to the traffic transmission situation in the t-th time slot, calculate the traffic throughput Th t and the inter-cell delay fairness F t , and then obtain the reward function

[0068] r t = βTh t -(1 - β)F t , where β is set to 0.5, indicating that the resource allocation focuses on the trade-off between throughput and delay fairness.

[0069] v. Each agent stores the experience in the t-th time slot in the experience pool.

[0070] vi. Each agent randomly extracts 256 experiences from its own experience pool, calculates the mean square error loss, and uses the Adam algorithm to train the parameters of its own Q-network.

[0071] c) Every 2,000 small loops, the agent copies the parameters of its own Q-network to the target network.

[0072] d) Every 1 large loop, the greedy coefficient ε decreases by 0.0001.

[0073] e) After each large loop, re-initialize the service arrival rate of each cell, and the service arrival rate of each cell is uniformly distributed within the range of 50M to 350M.

[0074] After training is completed, deploy the neural network to the operation control center or the on-satellite payload. Input the real-time service size of each cell into the neural network in each time slot, and according to the output of the neural network, the hopping beam pattern and beam bandwidth allocation result in this time slot can be obtained.

[0075]

Embodiment 2

[0076] As Figure 5 shown, assume that the number of satellite beams is 8 and the number of ground cells is 37. The entire multi-agent reinforcement learning framework is as Figure 6 shown. Complete the training of the multi-agent reinforcement learning model through the following steps:

[0077] a) Initialize 16 neural network parameters and 16 experience pools, with the size of each experience pool being 10,000; initialize the service arrival rate of each cell, assuming that the service arrival rate of each cell is uniformly distributed within the range of 50M to 350M; set the greedy coefficient ε to 1.

[0078] b) Start the loop process, which includes 10,000 large loops, and each large loop contains 200 small loops. The process of the t-th small loop is as follows:

[0079] i. Each agent observes the global state, defined as where represents the total service size of the n-th cell in the t-th time slot.

[0080] ii. Input the global state s t into the Q-network of each agent, and the k-th agent makes a decision according to the ε-greedy algorithm This decision refers to the irradiation position of the beam or the number of bandwidth resource blocks allocated.

[0081] iii. According to the decisions of each agent, calculate the transmission capacity of each beam, then transmit the service of the cells irradiated in the t-th time slot, and update the service volume of each cell in the t+1-th time slot to obtain the environmental state variable s t+1 .

[0082] iv. According to the service transmission situation in the t-th time slot, calculate the service throughput Th t and the inter-cell delay fairness F t , and then obtain the reward function

[0083] r t =βTh t-(1-β)F t , where β is set to 1, indicating that the resource allocation only focuses on improving throughput.

[0084] v. Each agent stores the experience at the t-th time slot in the experience pool.

[0085] vi. Each agent randomly samples 256 experiences from its respective experience pool, calculates the mean squared error loss, and uses the Adam algorithm to train the parameters of its Q-network.

[0086] c) Every 2000 mini-cycles, the agent copies the parameters of its Q-network to the target network.

[0087] d) Every 1 large cycle, the greedy coefficient ε decreases by 0.0001.

[0088] e) Every 1 large cycle, the traffic arrival rate of each cell is re-initialized, and the traffic arrival rate of each cell is uniformly distributed in the range of 50M to 350M.

[0089] After the training is completed, the neural network is deployed to the operation control center or the on-board payload. The real-time traffic size of each cell is input into the neural network at each time slot, and according to the output of the neural network, the hopping beam pattern and beam bandwidth allocation result at that time slot can be obtained.

[0090] Compared with the prior art, the dynamic hopping beam and beam bandwidth allocation method of the present invention can not only effectively utilize the three degrees of freedom of time, space, and frequency of beam resources, but also reduces the complexity of the agent decision space compared with single-agent reinforcement learning, has better performance in throughput and delay fairness, can achieve real-time hopping beam scheduling and beam bandwidth allocation for each time slot, and has a certain robustness.

Claims

1. A dynamic beam hopping and beam bandwidth allocation method based on multi-agent reinforcement learning, which is used for resource allocation of satellite beams in three dimensions of time, space and frequency. Characterized in that, This method includes the following steps: Step 1: Obtain satellite communication system parameters, including the number of beams K, the number of ground cells N, the total satellite transmit power, the satellite altitude, the number of satellite bandwidth resource blocks, and the bandwidth of each satellite broadband resource block. Step 2: Construct a multi-agent reinforcement learning simulation system model and perform offline training on the constructed simulation system model. In the multi-agent reinforcement learning simulation system model described in Step 2, one satellite beam corresponds to two agents, which are respectively responsible for the irradiation direction and bandwidth allocation of the satellite beam. Each of the two agents contains two neural networks and an experience pool, where the two neural networks are the target network and the Q network. The multi-agent reinforcement learning simulation system model described in Step 2 has a working process including the following sub-steps: Step S21: Initialize 2K neural network parameters and 2K experience pools; initialize the traffic arrival rate of each cell; set the environmental initial state to the traffic size of each ground cell in the first time slot; set the greedy coefficient ε to 1. Step S22: Start to execute a loop process, which includes M large loops, and each large loop contains T small loops; among them, the t-th small loop process includes the following sub-processes: i. Each agent observes the global state s t , defined as where represents the total traffic volume of the nth cell in the tth time slot; ii. Input the global state s t into the Q-network of each agent, and the k-th agent makes a decision according to the ε-greedy algorithm This decision refers to the irradiation position of the beam or the number of bandwidth resource blocks; iii. Calculate the transmission capacity of each satellite beam according to the decisions of each agent and the link budget, then transmit the traffic of the cell illuminated in the t-th time slot, and update the traffic volume of each cell in the (t + 1)-th time slot to obtain the environmental state variable s t+1 ; iv. Calculate the service throughput Th of the t-th time slot according to the service transmission situation of the t-th time slot t and the inter-cell delay fairness F t , and then obtain the reward function r t , which is expressed as shown in the following formula (1): r t = βTh t -(1 - β)F t (1), where β is a number between 0 and 1, which is used to balance the traffic throughput Th of the t-th time slot t and the inter-cell delay fairness F t between the weights; v. Each agent stores the experience at the t-th time slot into the corresponding The corresponding experience pool. vi. Each agent randomly extracts M experiences from its corresponding experience pool, calculates the mean square error loss, and uses the Adam algorithm to train the parameters of its own Q network. Step S23: Every C small loops, the agent copies the parameters of its own Q network to the target network. Step S24: Every 1 large loop, reduce the greedy coefficient ε. Step S25: Every 1 large loop, reset the traffic arrival rate of each cell and initialize the environmental state. Step 3: Deploy the trained simulation system model in Step 2 to the operation and control center or on-board payload. Step 4: The operation and control center or on-board payload inputs the real-time traffic queue size of each cell in each time slot into the trained simulation system model deployed on it, so as to obtain the irradiation position and the allocated bandwidth size of each beam in this time slot.

2. The dynamic beam hopping and beam bandwidth allocation method according to claim 1, Characterized in that, The multi-agent reinforcement learning simulation system model described in Step 2 includes multiple reinforcement learning agents, multiple experience pools, and a satellite beam hopping simulation environment.

3. The dynamic beam hopping and beam bandwidth allocation method according to claim 1, Characterized in that, The reinforcement learning algorithm adopted by the multi-agent reinforcement learning simulation system model described in Step 2 includes but is not limited to the DQN algorithm, the Double DQN algorithm, or the A3C algorithm.

4. The dynamic beam hopping and beam bandwidth allocation method according to claim 1, Characterized in that, The agent described in Step 2 adopts a neural network structure, including but not limited to a fully connected network, a convolutional neural network, and a recurrent neural network.

Citation Information

Patent Citations

  • Dynamic beam scheduling method based on deep reinforcement learning

    CN108966352A