Multi-beam satellite resource optimization method based on multi-agent reinforcement learning
By optimizing beam scheduling and spectrum allocation of multi-beam satellite systems through multi-agent reinforcement learning and simulated annealing algorithms, the problems of low spectrum resource utilization efficiency and poor adaptability to dynamic user needs in traditional methods are solved, achieving more efficient resource utilization and increased throughput.
Patent Information
- Application Number
- CN202511925310.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional multi-beam satellite systems face challenges in terms of low spectrum resource utilization efficiency and difficulty in adapting to dynamic user needs. Furthermore, spatial interference and spectrum competition between beams increase the complexity of scheduling decisions, limiting the system's resource utilization efficiency and service capabilities.
A multi-agent reinforcement learning approach is adopted, in which a multi-agent model is trained through the QMIX network architecture, and simulated annealing algorithm is used to optimize beam coverage strategy and spectrum allocation. The beam center and 3dB beamwidth are dynamically adjusted to achieve joint resource optimization of beam scheduling and spectrum allocation.
It improves beam space utilization and user coverage accuracy, enhances system spectrum efficiency, has stronger environmental adaptability and system scalability, and significantly improves system throughput and user service capabilities.
Smart Images

Figure CN121508632A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a method for optimizing multi-beam satellite resources based on multi-agent reinforcement learning. Background Technology
[0002] With the rapid growth of global communication demands, satellite communication, especially geostationary orbit (GEO) multi-beam satellite systems, is increasingly becoming one of the key technologies supporting wide-area coverage, long-distance communication, and highly reliable connections. Given the scarcity of spectrum resources and the increasingly complex interference environment, achieving efficient and flexible resource scheduling and management within a multi-beam architecture has become a core challenge for improving the overall system performance.
[0003] Multi-beam satellite systems expand system capacity through space reuse, but traditional static beam scheduling and spectrum allocation strategies rely on fixed cell divisions and lack the ability to perceive the dynamic distribution of users, making it difficult to adapt to rapidly changing service demands. Meanwhile, inter-beam spatial interference and spectrum contention significantly increase the complexity of scheduling decisions, limiting the system's resource utilization efficiency and service capabilities. Summary of the Invention
[0004] The purpose of this invention is to provide a multi-beam satellite resource optimization method based on multi-agent reinforcement learning, so as to alleviate the technical problems of low spectrum resource utilization efficiency and difficulty in adapting to dynamic user needs in existing multi-beam satellite resource optimization methods.
[0005] In a first aspect, the present invention provides a method for optimizing multi-beam satellite resources based on multi-agent reinforcement learning, comprising: acquiring the observation-action history trajectory information of a target beam in the current time slot of a multi-beam satellite communication system; wherein, the target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global state in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global state includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: the beam center position and the 3dB beamwidth; processing the observation-action history trajectory information of the target beam using a target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot; wherein, each beam in the multi-beam satellite communication system is coupled with a target multi-agent reinforcement learning model trained using a QMIX network architecture. Each agent in the learning model is deployed one-to-one, and each agent has an independent beam resource optimization model. The state of an agent is the observation-action history trajectory information of the corresponding beam, and the action of an agent is the beam coverage strategy of the corresponding beam. The rewards of all agents are positively correlated with the total system throughput and user service coverage. The simulated annealing algorithm is used to perform OFDMA subcarrier resource allocation processing on the beam coverage strategy and global state of all beams in the current time slot, to obtain the spectrum allocation result of each beam in the current time slot. The spectrum allocation result includes: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier. The users to be served represent users within the beam coverage area. Based on the beam coverage strategy and spectrum allocation result of all beams in the current time slot, the resource optimization strategy of the multi-beam satellite communication system in the current time slot is determined.
[0006] In an optional implementation, the QMIX network architecture includes a multi-agent reinforcement learning model and a hybrid network model. The method further includes: initializing a global state; repeatedly executing the following steps, and sampling a preset amount of empirical data from the experience replay pool every preset number of rounds, to jointly update the network parameters of the multi-agent reinforcement learning model and the hybrid network model using a mean squared error loss function, until a specified number of rounds is reached, or the mean squared error loss function value is less than a preset threshold; each agent in the multi-agent reinforcement learning model selects an action based on its current state using an ε-greedy policy; using a simulated annealing algorithm to perform OFDMA subcarrier resource allocation processing on the global state and the actions of all agents, obtaining the spectrum allocation result for the beam corresponding to each agent in the current time slot; calculating the reward for all agents based on the global state, the actions of all agents, and the spectrum allocation result for all beams, and determining the next state for each agent; and storing the states, actions, rewards, and next states of all agents in the experience replay pool.
[0007] In an optional implementation, based on the global state, the actions of all agents, and the spectrum allocation results of all beams, the reward for all agents is calculated, including: acquiring the location information of the target terminal user, the multi-beam antenna radiation pattern, free-space transmission loss, and the user receiving antenna gain; wherein, the target terminal user represents any user among all terminal users; based on the location information of the target terminal user and the actions of all agents, a first binary variable is determined to reflect the coverage relationship between the target terminal user and each beam in the multi-beam satellite communication system; based on the spectrum allocation results of all beams, the number of target subcarriers in other beams that have spectral interference with the target terminal user is determined; based on the actions of all agents, a second binary variable is determined to reflect the spatial interference relationship between beams; based on the multi-beam antenna radiation pattern, free-space transmission loss, and user receiving antenna gain, a first channel gain from each beam to the center position of the selected beam is calculated, so as to... And, the second channel gain of the interfering beam that has spatial interference with the target beam to the center position of the target beam; wherein, the center position of the target beam represents the selected center position of the target beam; based on the first binary variable, the number of target subcarriers, the second binary variable, the first channel gain, the second channel gain and the spectrum allocation result, the downlink channel capacity of the target terminal user in the current time slot is calculated; based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the end of the current time slot are determined; the sum of the throughput of all terminal users is taken as the total system throughput of the current time slot; the total number of terminal users whose remaining service demand is 0 after the end of the current time slot is taken as the user service coverage of the current time slot; based on the total system throughput and user service coverage of the current time slot, the reward of all agents in the current time slot is calculated.
[0008] In an optional implementation, based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the end of the current time slot are determined, including: taking the minimum value between the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user as the throughput of the target terminal user in the current time slot; and taking the difference between the service demand of the target terminal user and the throughput in the current time slot as the remaining service demand of the target terminal user after the end of the current time slot.
[0009] In an optional implementation, the formula for the downlink channel capacity of the target terminal user in the current time slot is: ;in, Indicates the target end user Downlink channel capacity in the current time slot, Indicates the target end user With beam The first binary variable that represents the overlay relationship between them. This represents the bandwidth of a single subcarrier. Indicates beam Assigned to standard end users The number of subcarriers, Indicates beam To the center position of the selected beam The first channel gain, , Indicates beam To the center of the beam Transmit antenna gain in the direction of direction Indicates free space transmission loss. Indicates the user's receiving antenna gain. This indicates the transmit power of each subcarrier. Represents the noise power spectral density. Indicates the beam used to reflect With beam The second binary variable representing the spatial interference relationship between them. Indicates beam Interference beams with spatial interference relationship to beam Selected center position The second channel gain, Indicates other beams and the target end user Number of target subcarriers subject to spectral interference. This indicates the total number of beams in a multi-beam satellite communication system.
[0010] In an optional implementation, the reward formula for all agents is: ;in, express Rewards for all agents in the time slot, express Total system throughput of time slots , express Time slot target end user throughput, This indicates the total number of end users in a multi-beam satellite communication system. express User service coverage of time slots This indicates the preset weight parameters. and All represent the homogenized reference values.
[0011] In an optional implementation, the hybrid network model is a supernetwork driven by a global state, and the agent's local Q-value undergoes attention enhancement processing before being input into the supernetwork; the formula for the global Q-value output by the hybrid network is: ;in, This represents the global state of time slot t. This represents the set of actions of all agents in time slot t. Represents the network parameters of the hybrid network model. , , This represents the local Q-value of agent k. This indicates the total number of beams in a multi-beam satellite communication system. This indicates the preset attention enhancement processing function. and Both indicate that the hypernetwork is based on global state. Determined weights, and Both indicate that the hypernetwork is based on global state. Determined bias.
[0012] Secondly, the present invention provides a multi-beam satellite resource optimization device based on multi-agent reinforcement learning, comprising: an acquisition module for acquiring observation-action history trajectory information of a target beam in the current time slot in a multi-beam satellite communication system; wherein, the target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global state in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global state includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: the beam center position and the 3dB beamwidth; and a processing module for processing the observation-action history trajectory information of the target beam using a target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot; wherein, each beam in the multi-beam satellite communication system is coupled with a target multi-agent reinforcement learning model trained using a QMIX network architecture. Each agent in the learning model is deployed one-to-one, and each agent has an independent beam resource optimization model. The state of an agent is the observation-action history trajectory information of the corresponding beam, and the action of an agent is the beam coverage strategy of the corresponding beam. The rewards of all agents are positively correlated with the total system throughput and user service coverage. The allocation module is used to perform OFDMA subcarrier resource allocation processing on the beam coverage strategy and global state of all beams in the current time slot using the simulated annealing algorithm to obtain the spectrum allocation result of each beam in the current time slot. The spectrum allocation result includes: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier. The users to be served represent users within the beam coverage area. The determination module is used to determine the resource optimization strategy of the multi-beam satellite communication system in the current time slot based on the beam coverage strategy and spectrum allocation result of all beams in the current time slot.
[0013] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the multi-beam satellite resource optimization method based on multi-agent reinforcement learning as described in any of the foregoing embodiments.
[0014] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the multi-beam satellite resource optimization method based on multi-agent reinforcement learning as described in any of the foregoing embodiments.
[0015] The multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in this invention is essentially a user-centric beam scheduling mechanism. It combines multi-agent collaborative learning with a fine-grained spectrum resource scheduling strategy to achieve joint resource optimization of beam scheduling and spectrum allocation. In the beam scheduling part, by dynamically adjusting the beam center and 3dB beamwidth, beam space utilization and user coverage accuracy are improved. In the spectrum allocation part, simulated annealing algorithm is used for subcarrier resource allocation and interference control, improving system spectrum efficiency. Compared with traditional schemes based on fixed area partitioning and static scheduling, the method of this invention has stronger environmental adaptability and system scalability, and can continuously maintain optimized performance in scenarios with dynamically changing user distribution. Furthermore, by using the QMIX network architecture to train the multi-agent reinforcement learning model, efficient collaboration and policy sharing among beams can be achieved, significantly improving system throughput and user service capabilities while keeping system complexity under control. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 A flowchart of a multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in an embodiment of the present invention; Figure 2 This is an application scenario diagram of a multi-beam satellite resource optimization method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a QMIX network architecture provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the variation of antenna gain with off-axis angle, provided by an embodiment of the present invention. Figure 5 A functional block diagram of a multi-beam satellite resource optimization device based on multi-agent reinforcement learning provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0020] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0021] Example 1 Figure 1 A flowchart of a multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in this embodiment of the invention is shown below. Figure 1 As shown, the method specifically includes the following steps: Step S102: Obtain the observation-action history trajectory information of the target beam in the current time slot in the multi-beam satellite communication system.
[0022] The target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global status in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global status includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: the beam center position and the 3dB beamwidth.
[0023] Specifically, the multi-beam satellite communication system described in the embodiments of the present invention, such as a geostationary satellite relay link multi-beam satellite system, is... Figure 2 As shown, its end users are distributed across the Earth's surface. Beamforming technology is required to ensure that the satellite can continuously illuminate end users within a designated area. Multi-beam satellites can provide... Several beams, and have coverage throughout the entire coverage area. Individual end users, settings , This represents the set of end users. It exists in the system space. One available beam coverage center.
[0024] To determine the resource optimization strategy for a multi-beam satellite communication system, this invention proposes a centralized training and distributed execution architecture for multi-agent beam scheduling and spectrum optimization strategies. Each agent corresponds to one beam, and after joint training of multiple agents, the policy network of all agents converges. Therefore, in the actual deployment phase, all agents can perform online inference using a fixed strategy; that is, each agent can independently decide its chosen action based solely on its local state.
[0025] In this embodiment of the invention, the state of the agent is the observation-action history trajectory information of its corresponding beam. Therefore, in the actual deployment phase, in order to determine the resource optimization strategy of the system in the current time slot, it is necessary to obtain the observation-action history trajectory information of each beam in the system in the current time slot. , , , , ,in, This represents the observation-action history trajectory information of beam k in time slot t. This represents the local observation state of beam k in time slot t. This represents the beam coverage strategy of beam k in time slot t-1. This represents the global state in time slot t. In other words, in this embodiment of the invention, all agents share the global state. Indicates the terminal user in time slot t The business demand, Indicates the terminal user in time slot t The queue is waiting for the first The traffic volume per time slot, where TTL represents the lifespan of data traffic.
[0026] As can be seen from the expression for the global state, it reflects the traffic demands of each terminal user in the multi-beam satellite system. Specifically, the global state consists of the storage traffic (i.e., service demand) of terminal users within each time slot, reflecting the current load and resource requirements of the system. The system updates its state information by observing the service demand of each user within a time slot, that is, it dynamically adjusts as user demand changes.
[0027] Step S104: The observation-action history trajectory information of the target beam is processed using the target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot.
[0028] In this multi-beam satellite communication system, each beam corresponds one-to-one with each agent in the target multi-agent reinforcement learning model trained using the QMIX network architecture, and each agent has an independent beam resource optimization model. The state of an agent is the observation-action history trajectory information of the corresponding beam, and the action of an agent is the beam coverage strategy of the corresponding beam. The rewards of all agents are positively correlated with the total system throughput and user service coverage.
[0029] In this embodiment of the invention, the specific method of beam scheduling is to precisely select the beam center position and 3dB beamwidth of each beam to maximize system throughput while increasing the number of users served. The multi-agent reinforcement learning model contains... Each intelligent agent is responsible for The beam coverage strategy for each beam involves each agent selecting an appropriate 3dB beamwidth and beam center position based on its own state to optimize the overall system performance of the multi-beam communication system. The set of actions of all agents in time slot t is represented as: ,in, This represents the action of agent k in time slot t.
[0030] The objective of this invention is to maximize total system throughput and user service coverage. User service coverage represents the number of end users whose service needs are fully met. Therefore, all agents share a unified reward mechanism to promote beam coordination and ensure overall system optimization. Furthermore, this invention sets the rewards for all agents to be positively correlated with total system throughput and user service coverage. Clearly, this reward mechanism can guide agents to adjust beam parameters according to system status, improving overall performance and resource utilization efficiency.
[0031] In multi-agent reinforcement learning, achieving efficient collaborative decision-making among multiple agents is crucial for improving the overall system performance. Therefore, this invention introduces QMIX as the core decision-making learning framework to enhance the efficiency and stability of joint optimization in multi-beam satellite systems. As described above, the satellite... Each agent makes independent decisions regarding the coverage center and range of each beam. Traditional multi-agent reinforcement learning methods (such as DQN) train a Q-value network for each agent and allow them to make decisions independently. However, in complex multi-agent environments, this independent decision-making may not adequately account for the influence of other agents. Therefore, QMIX introduces a hybrid network to combine the Q-values of each agent, resulting in a global Q-value function. QMIX uses a hybrid network that generates the global Q-value based on the local Q-values of each agent, ensuring that the global Q-value is a monotonically increasing function of the local Q-values.
[0032] Step S106: Using the simulated annealing algorithm, OFDMA subcarrier resource allocation processing is performed on the beam coverage strategy of all beams in the current time slot and the global state under the current time slot to obtain the spectrum allocation result of each beam under the current time slot.
[0033] The spectrum allocation results include: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier; the user to be served refers to the user within the beam coverage area.
[0034] To utilize spectrum resources more efficiently, this embodiment of the invention introduces a simulated annealing algorithm for OFDMA subcarrier resource allocation. The total bandwidth is divided into multiple subcarrier blocks, and subcarriers are allocated to users under each beam as needed. Simultaneously, the orthogonality of subcarriers within a beam is guaranteed, and spectrum reuse between beams is supported. Specifically, after obtaining the global state in the current time slot and the beam coverage strategy of all beams in the current time slot, these two are used as inputs to the simulated annealing algorithm. Under the constraints of the OFDMA mechanism, the simulated annealing algorithm can output the number of subcarriers allocated to users within the coverage area of each beam and the frequency band position corresponding to each subcarrier.
[0035] The specific process is as follows: 1. Input preparation: After completing step S104, the system has obtained the beam coverage strategy (beam center position and 3dB beamwidth) of all beams in the current time slot and the global status (service demand of all end users).
[0036] 2. Resource Modeling: The total available bandwidth of the system is divided into multiple subcarrier blocks, each with a fixed bandwidth and frequency band location number; subcarrier blocks within a beam must remain orthogonal and cannot be repeatedly assigned to different users within the same beam.
[0037] 3. Optimization Objective: To maximize the total system throughput and improve user service coverage, a joint optimization model for subcarrier number and frequency band allocation is established. To facilitate the determination of the optimal solution, the weighted sum of the total system throughput and user service coverage is used as the objective function of the joint optimization model.
[0038] 4. Simulated annealing solution: (1) Randomly generate the current spectrum allocation scheme (number of subcarriers + frequency band position) and use it as the initial solution.
[0039] (2) Randomly adjust the allocation scheme in the neighborhood space (such as increasing / decreasing the number of subcarriers of a user, changing the frequency band position) to obtain the updated spectrum allocation scheme.
[0040] (3) Determine whether to accept the updated spectrum allocation scheme based on the magnitude of the first optimization objective function value corresponding to the updated spectrum allocation scheme and the second optimization objective function value corresponding to the spectrum allocation scheme before the update, so as to escape the local optimum; if the first optimization objective function value is greater than the second optimization objective function value, then accept the update; otherwise, reject the update.
[0041] (4) Gradually reduce the temperature until convergence or the iteration limit is reached, and output the optimal or near-optimal spectrum allocation scheme.
[0042] 5. Output: The output shows the number of subcarriers allocated to each user in the current time slot for each beam and the frequency band position they occupy, which is then used by the system to execute spectrum allocation commands.
[0043] Through the above methods, the present invention can not only accurately control the amount of resources allocated to each user, but also flexibly plan the occupied frequency band positions, realize intra-beam orthogonality and efficient inter-beam multiplexing, thereby maximizing spectrum utilization efficiency and overall system performance.
[0044] Step S108: Based on the beam coverage strategy and spectrum allocation results of all beams in the current time slot, determine the resource optimization strategy of the multi-beam satellite communication system in the current time slot.
[0045] That is, in this embodiment of the invention, the resource optimization strategy of the multi-beam satellite communication system includes: beam coverage strategy and spectrum allocation results for all beams.
[0046] The multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in this invention is essentially a user-centric beam scheduling mechanism. It combines multi-agent collaborative learning with a fine-grained spectrum resource scheduling strategy to achieve joint resource optimization of beam scheduling and spectrum allocation. In the beam scheduling part, the beam center and 3dB beamwidth are dynamically adjusted to improve beam space utilization and user coverage accuracy. In the spectrum allocation part, simulated annealing is used for subcarrier resource allocation and interference control, improving system spectrum efficiency. Compared with traditional schemes based on fixed area partitioning and static scheduling, this invention's method has stronger environmental adaptability and system scalability, maintaining optimized performance even in scenarios with dynamically changing user distribution. Furthermore, by training the multi-agent reinforcement learning model using the QMIX network architecture, efficient beam collaboration and policy sharing can be achieved, significantly improving system throughput and user service capabilities while keeping system complexity under control.
[0047] In one alternative implementation, such as Figure 3 As shown, the QMIX network architecture includes: a multi-agent reinforcement learning model and a hybrid network model; the method of this invention also includes the following: Initialize the global state; repeat steps S201 to S204 below, and sample a preset number of experience data from the experience replay pool every preset round, so as to jointly update the network parameters of the multi-agent reinforcement learning model and the network parameters of the hybrid network model using the mean squared error loss function, until a specified round is reached, or the mean squared error loss function value is less than a preset threshold.
[0048] Step S201: Each agent in the multi-agent reinforcement learning model selects an action based on its current state using an ε-greedy policy.
[0049] Step S202: The simulated annealing algorithm is used to perform OFDMA subcarrier resource allocation processing on the global state and the actions of all agents to obtain the spectrum allocation result of the beam corresponding to each agent in the current time slot.
[0050] Step S203: Based on the global state, the actions of all agents, and the spectrum allocation results of all beams, calculate the rewards of all agents and determine the next state of each agent.
[0051] Step S204: Store the states, actions, rewards, and next states of all agents into the experience replay pool.
[0052] Specifically, as described above, the multi-agent beam scheduling and spectrum optimization strategy proposed in this embodiment of the invention adopts a centralized training and distributed execution architecture. During the training phase, each agent selects an action based on its current state through an ε-greedy strategy, and determines the spectrum allocation result of each beam based on the simulated annealing algorithm, thereby enabling the system to execute a joint scheduling strategy and calculate the rewards for all agents.
[0053] The known states of the agent include: the global state in the current time slot and the action in the previous time slot. Because, after obtaining the action in the current time slot, to obtain the agent's next (time slot) state, the following formula needs to be used to calculate the end user's state. Business demand in the next time slot: ;in, Indicates the terminal user in time slot t+1 The business demand, Indicates the terminal user in time slot t The business demand, Indicates the terminal user in time slot t The throughput (i.e., the amount of traffic processed in time slot t). Indicates the terminal user in time slot t Discarded traffic (i.e., traffic that is waiting for TTL slots). Indicates the terminal user in time slot t+1 The volume of business that has arrived.
[0054] Next, the states, actions, rewards, and next states of all agents are recorded and stored in the experience replay pool. Every certain number of steps, small batches of data are sampled from the replay pool, and the mean squared error loss function is used to jointly update the hybrid network model and the multi-agent reinforcement learning model until the iteration termination condition is met: reaching a specified number of iterations, or the mean squared error loss function value is less than a preset threshold.
[0055] In one alternative implementation, the hybrid network model is a supernetwork driven by a global state, and the agent's local Q-value is enhanced with attention before being input into the supernetwork.
[0056] Specifically, to further improve the collaborative efficiency and policy expression capabilities among multi-agent networks, this invention introduces an attention mechanism into the QMIX hybrid network. This mechanism models the policy dependencies between agents, thereby highlighting the interaction relationships between key beams and suppressing interference from invalid information. Specifically, the agent's local Q-value undergoes attention enhancement processing before being input into the supernetwork.
[0057] In this embodiment of the invention, each beam is modeled as an intelligent agent, and in each time slot, the system has a total of Each beam entity maintains a local Q-function. ,in, This represents the observation-action history trajectory information of beam k in time slot t, that is, the state of agent k in time slot t. This represents the beam coverage strategy of beam k in time slot t, that is, the action of agent k in time slot t. This represents the network parameters of agent k in time slot t. The agent network determines its local policy selection by evaluating historical trajectory information, thereby achieving distributed beam selection.
[0058] To enhance the collaborative capabilities among agents, this embodiment of the invention introduces an attention mechanism to enhance the local Q-value of each agent. The enhanced Q-value after fusing global policy information is expressed as: ;in, , representing the set of original action values of all intelligent agents. This represents the local Q-value of agent k; , , , Each represents the mapping matrix of the i-th attention head. This represents the dimension of each attention head; all attention heads have the same dimension; and the total number of attention heads. It is an adjustable hyperparameter. This indicates the output mapping matrix.
[0059] based on As can be seen from the expression, this embodiment of the invention employs a multi-head attention mechanism based on action value vectors to model the policy dependencies between agents. Specifically, the set of action values for all agents... First, attention vectors corresponding to multiple attention heads are obtained through linear mapping. Then, the outputs of all attention heads are concatenated and mapped to the final enhanced representation.
[0060] Subsequently, all enhanced local Q-values are fused through a hybrid network to calculate a global Q-value, which serves as a monitoring signal for the optimization process. In this embodiment of the invention, the formula for the global Q-value output by the hybrid network is: ;in, This represents the global state of time slot t. This represents the set of actions of all agents in time slot t. Represents the network parameters of the hybrid network model. , , This represents the local Q-value of agent k. This indicates the total number of beams in a multi-beam satellite communication system. This indicates the preset attention enhancement processing function. and Both indicate that the hypernetwork is based on global state. Determined weights, and Both indicate that the hypernetwork is based on global state. A defined bias enables nonlinear fusion of local Q values.
[0061] During training, the loss function is defined as the mean squared error (MSE) between the target Q-value and the predicted global Q-value, and its objective function is as follows: ;in, This indicates the batch size sampled from the experience replay pool. This represents the target Q value, and the formula for calculating the target Q value is: ,in, This represents the reward for all agents in time slot t. This represents the reward discount factor. This represents the global state of time slot t+1. This represents the joint action chosen by all agents in time slot t+1 to maximize the objective Q-value. Network parameters representing the delay target network, Periodically extract network parameters from hybrid networks The process is replicated in the middle to stabilize the training process.
[0062] Finally, the network parameters of the hybrid network are updated using gradient descent, with the following update formula: .in, The learning rate and gradient are calculated relative to the network parameters of the hybrid network. This update process follows the standard backpropagation mechanism.
[0063] In an optional implementation, step S203 above, which calculates the rewards for all agents based on the global state, the actions of all agents, and the spectrum allocation results of all beams, specifically includes the following steps: Step S20301: Obtain the location information of the target terminal user, the multi-beam antenna radiation pattern, the free space transmission loss, and the user receiving antenna gain; wherein, the target terminal user refers to any user among all terminal users.
[0064] Specifically, in this embodiment of the invention, the radiation pattern of the multi-beam antenna refers to ITU-RS.672-4, and its specific form is as follows: ;in, Indicates beam To the center of the beam Transmit antenna gain in the direction of direction =2.88, =6.32, =-25dB, Indicates the off-axis angle. This represents half of the 3dB beamwidth. This indicates the maximum gain of the satellite transmitting antenna. , Figure 4 This is a schematic diagram illustrating the variation of antenna gain with off-axis angle provided in an embodiment of the present invention. Figure 4 It can be seen that increasing the beamwidth by 3dB will reduce the maximum gain of the transmitting antenna, thus reducing the beam capacity, but at the same time, a larger coverage area is obtained, which allows more end users to be served. This shows that there is a negative correlation between beam coverage area and beam capacity.
[0065] The user receiver antenna gain is expressed as The formula for calculating free space transmission loss is as follows: , Indicates the distance between the satellite and the ground. Indicates the signal frequency.
[0066] Step S20302: Based on the location information of the target terminal user and the actions of all intelligent agents, determine a first binary variable to reflect the coverage relationship between the target terminal user and each beam in the multi-beam satellite communication system.
[0067] The actions of each agent represent the beam coverage strategy of its corresponding beam: the beam center position and the 3dB beamwidth, which determines the coverage range of each beam. Therefore, based on the location information of the target terminal user, it can be determined whether the target terminal user is covered by beam k, thus obtaining the first binary variable. 0 indicates the target end user Not covered by beam k, 1 indicates the target end user It is covered by beam k.
[0068] Step S20303: Based on the spectrum allocation results of all beams, determine the number of target subcarriers in other beams that have spectrum interference with the target terminal user.
[0069] Specifically, the first step is to extract the target end users. The target subcarrier is obtained by first identifying the set of subcarriers, then traversing other beams, marking overlapping subcarriers, and finally counting the number of interfering subcarriers (i.e., overlapping subcarriers).
[0070] Step S20304: Based on the actions of all agents, determine a second binary variable to reflect the spatial interference relationship between beams.
[0071] Based on the actions of each intelligent agent, the beam center position of each beam can be determined. In this embodiment of the invention, the distance between the beam center positions is used to determine whether spatial interference exists between beams. If the distance between the center positions of two beams is less than a preset distance threshold, then spatial interference is determined to exist between the two beams. The second binary variable used to reflect the spatial interference relationship between beams is... 0 indicates beam With beam There is no spatial interference between them, and 1 indicates the beam. With beam There is spatial interference.
[0072] Step S20305: Based on the multi-beam antenna radiation pattern, free-space transmission loss, and user receiving antenna gain, calculate the first channel gain of each beam to the selected beam center position, and the second channel gain of the interfering beam that has spatial interference with the target beam to the target beam center position; wherein, the target beam center position represents the selected beam center position of the target beam.
[0073] Known beam To the center of the beam The gain of the transmitting antenna in the direction is Free space transmission loss and user receiving antenna gain Therefore, during downlink transmission, the beam To the center position of the selected beam The first channel gain is: With beam Interference beams with spatial interference relationship to beam Central position The second channel gain is expressed as Calculation method reference , that is, , Refer to the multi-beam antenna radiation pattern for specific values.
[0074] Step S20306: Based on the first binary variable, the number of target subcarriers, the second binary variable, the first channel gain, the second channel gain, and the spectrum allocation result, calculate the downlink channel capacity of the target terminal user in the current time slot.
[0075] In one alternative implementation, the formula for the downlink channel capacity of the target terminal user in the current time slot is: ;in, Indicates the target end user Downlink channel capacity in the current time slot, Indicates the target end user With beam The first binary variable that represents the overlay relationship between them. This represents the bandwidth of a single subcarrier. Indicates beam Assigned to standard end users The number of subcarriers, Indicates beam To the center position of the selected beam The first channel gain, , Indicates beam To the center of the beam Transmit antenna gain in the direction of direction Indicates free space transmission loss. Indicates the user's receiving antenna gain. This indicates the transmit power of each subcarrier. Represents the noise power spectral density. Indicates the beam used to reflect With beam The second binary variable representing the spatial interference relationship between them. Indicates beam Interference beams with spatial interference relationship to beam Selected center position The second channel gain, Indicates other beams and the target end user Number of target subcarriers subject to spectral interference. This indicates the total number of beams in a multi-beam satellite communication system.
[0076] Step S20307: Based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, determine the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the current time slot ends.
[0077] Knowing the service demand of the target terminal user in the current time slot And calculate the downlink channel capacity of the target terminal user in the current time slot. Then, the actual throughput that the target end user can obtain in the current time slot can be determined. Then, based on its throughput, the remaining service demand of the end user after the current time slot ends can be calculated.
[0078] In one optional embodiment, step S20307, based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, determines the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the current time slot ends, specifically including the following: The minimum of the downlink channel capacity and the service demand of the target terminal user in the current time slot is taken as the throughput of the target terminal user in the current time slot. That is, Taking a smaller throughput value means that if the channel capacity is less than the service demand, it is limited by the channel capacity; if the service demand is less than the channel capacity, it is limited by the actual service demand of the user.
[0079] The difference between the target end user's service demand and the throughput in the current time slot is taken as the target end user's remaining service demand after the current time slot ends.
[0080] Step S20308: The sum of the throughput of all terminal users is taken as the total system throughput of the current time slot. The total system throughput of the current time slot is expressed as: .
[0081] Step S20309: The total number of terminal users whose remaining service demand is 0 after the current time slot ends will be used as the user service coverage rate of the current time slot.
[0082] In other words, to measure the overall capability of the system to serve users, user service coverage is defined as the total number of users whose business needs are fully met within a time slot.
[0083] Step S20310: Calculate the rewards for all agents in the current time slot based on the total system throughput and user service coverage of the current time slot.
[0084] In one alternative implementation, the reward formula for all agents is: ;in, express Rewards for all agents in the time slot, express Total system throughput of time slots , express Time slot target end user throughput, This indicates the total number of end users in a multi-beam satellite communication system. express User service coverage of time slots This indicates the preset weight parameters. and All represent the homogenized reference values.
[0085] This invention also provides a parameter setting example in a simulation scenario to evaluate a geostationary orbit (GEO) multi-beam satellite communication system. The system is configured with six beams, each capable of serving 360 users within a target area. The satellite transmission frequency band is set to 20 GHz (Ka band), and Orthogonal Frequency Division Multiple Access (OFDMA) is used as the resource allocation mechanism, dividing the total bandwidth of 500 MHz into multiple subcarrier units, each with a bandwidth of 120 kHz. During the simulation, the user service request rate follows a uniform distribution from 3 Mbps to 18 Mbps, and the system schedules requests at a time granularity of 2 ms. At each time step, the agent dynamically adjusts the beam center position and 3 dB beamwidth based on the current environmental state, allocates spectrum resources to users using a simulated annealing algorithm, and then calculates the total system throughput and user coverage as reward feedback.
[0086] Ultimately, the trained agent can independently execute distributed scheduling strategies in actual system deployments, enhance inter-beam coordination capabilities through attention enhancement mechanisms, and obtain globally optimal strategies through centralized training. Simulation results show that, under dynamic business demand scenarios, the user-center joint scheduling strategy proposed in this embodiment outperforms existing greedy scheduling, genetic algorithms, and traditional deep reinforcement learning methods in terms of system throughput, spectrum utilization, and number of user services.
[0087] In summary, this invention addresses the problems of traditional beam scheduling methods in multi-beam satellite communication systems, such as strong staticity, low spectrum resource utilization efficiency, and difficulty in adapting to dynamic user needs. It proposes a multi-agent reinforcement learning-based method for optimizing multi-beam satellite resources. This method is designed for scenarios with dynamically changing user distribution, comprehensively considering key factors such as inter-beam interference, spectrum allocation, and service coverage to achieve coordinated optimization of beam scheduling and spectrum resources. Specifically, it proposes the following three innovations: A. A user-centric beam scheduling mechanism is proposed, which dynamically adjusts the center position of the beam and the 3dB beamwidth to achieve on-demand coverage for users in different areas, effectively improving service accuracy and spatial resource utilization.
[0088] B. Construct a multi-agent reinforcement learning architecture that combines attention mechanisms, model each beam as an independent agent, and capture the cooperative relationship between beams through a hybrid network with attention enhancement to improve the consistency and convergence efficiency of overall decision-making.
[0089] C. Introduce a spectrum allocation algorithm based on simulated annealing, divide the system bandwidth into multiple fine-grained subcarrier blocks, and dynamically allocate them according to user needs to achieve intra-beam spectrum orthogonality and inter-beam frequency reuse, thereby improving spectrum utilization efficiency.
[0090] Example 2 This invention also provides a multi-beam satellite resource optimization device based on multi-agent reinforcement learning. This device is mainly used to execute the multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in Embodiment 1 above. The device provided in this invention will be described in detail below.
[0091] Figure 5 A functional block diagram of a multi-beam satellite resource optimization device based on multi-agent reinforcement learning provided in an embodiment of the present invention is shown below. Figure 5 As shown, the device mainly includes: an acquisition module 10, a processing module 20, an allocation module 30, and a determination module 40, wherein: The acquisition module 10 is used to acquire the observation-action history trajectory information of the target beam in the current time slot in the multi-beam satellite communication system; wherein, the target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global state in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global state includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: the beam center position and the 3dB beamwidth.
[0092] Processing module 20 is used to process the observation-action history trajectory information of the target beam using the target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot. In this multi-beam satellite communication system, each beam is deployed one-to-one with each agent in the target multi-agent reinforcement learning model trained using the QMIX network architecture, and each agent has an independent beam resource optimization model. The state of the agent is the observation-action history trajectory information of the corresponding beam, the action of the agent is the beam coverage strategy of the corresponding beam, and the reward of all agents is positively correlated with the total system throughput and user service coverage.
[0093] The allocation module 30 is used to perform OFDMA subcarrier resource allocation processing on the beam coverage strategy and global state of all beams in the current time slot using the simulated annealing algorithm, so as to obtain the spectrum allocation result of each beam in the current time slot; wherein, the spectrum allocation result includes: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier; the user to be served refers to the user within the beam coverage area.
[0094] The determination module 40 is used to determine the resource optimization strategy of the multi-beam satellite communication system in the current time slot based on the beam coverage strategy and spectrum allocation results of all beams in the current time slot.
[0095] The multi-beam satellite resource optimization device based on multi-agent reinforcement learning provided in this invention essentially employs a user-centric beam scheduling mechanism, combining multi-agent collaborative learning with a fine-grained spectrum resource scheduling strategy to achieve joint resource optimization of beam scheduling and spectrum allocation. In the beam scheduling section, dynamic adjustment of the beam center and 3dB beamwidth improves beam space utilization and user coverage accuracy. In the spectrum allocation section, simulated annealing is used for subcarrier resource allocation and interference control, improving system spectrum efficiency. Compared with traditional schemes based on fixed area partitioning and static scheduling, this invention offers stronger environmental adaptability and system scalability, maintaining optimized performance even in scenarios with dynamically changing user distribution. Furthermore, by training the multi-agent reinforcement learning model using the QMIX network architecture, efficient beam collaboration and policy sharing can be achieved, significantly improving system throughput and user service capabilities while keeping system complexity under control.
[0096] Optionally, the QMIX network architecture includes: a multi-agent reinforcement learning model and a hybrid network model; the device also includes: The initialization module is used to initialize the global state.
[0097] The repeated execution module is used to repeatedly call the following selection unit, allocation unit, calculation unit and storage unit, and every preset round, it samples a preset number of experience data from the experience replay pool to jointly update the network parameters of the multi-agent reinforcement learning model and the network parameters of the hybrid network model using the mean squared error loss function, until a specified round is reached, or the mean squared error loss function value is less than a preset threshold.
[0098] Selection Unit: In a multi-agent reinforcement learning model, each agent selects an action based on its current state using an ε-greedy policy.
[0099] The allocation unit is used to perform OFDMA subcarrier resource allocation processing on the global state and the actions of all agents using the simulated annealing algorithm, so as to obtain the spectrum allocation result of the beam corresponding to each agent in the current time slot.
[0100] The computing unit is used to calculate the rewards of all agents and determine the next state of each agent based on the global state, the actions of all agents, and the spectrum allocation results of all beams.
[0101] The storage unit is used to store the states, actions, rewards, and next states of all agents into the experience replay pool.
[0102] Optionally, the computing unit includes: The acquisition sub-unit is used to acquire the location information of the target terminal user, the multi-beam antenna radiation pattern, the free space transmission loss, and the user receiving antenna gain; wherein, the target terminal user refers to any user among all terminal users.
[0103] The first determining subunit is used to determine a first binary variable reflecting the coverage relationship between the target terminal user and each beam in the multi-beam satellite communication system, based on the location information of the target terminal user and the actions of all intelligent agents.
[0104] The second determining subunit is used to determine the number of target subcarriers in other beams that have spectral interference with the target terminal user, based on the spectrum allocation results of all beams.
[0105] The third determining subunit is used to determine a second binary variable that reflects the spatial interference relationship between beams based on the actions of all agents.
[0106] The first calculation subunit is used to calculate, based on the multi-beam antenna radiation pattern, free-space transmission loss, and user receiving antenna gain, the first channel gain of each beam to the selected beam center position, and the second channel gain of interfering beams that have spatial interference with the target beam to the target beam center position; wherein, the target beam center position represents the selected beam center position of the target beam.
[0107] The second calculation subunit is used to calculate the downlink channel capacity of the target terminal user in the current time slot based on the first binary variable, the number of target subcarriers, the second binary variable, the first channel gain, the second channel gain, and the spectrum allocation result.
[0108] The fourth determining subunit is used to determine the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the current time slot ends, based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state.
[0109] The fifth determining subunit is used to sum the throughput of all terminal users as the total system throughput of the current time slot.
[0110] The sixth sub-unit is used to determine the total number of terminal users whose remaining service demand is 0 after the current time slot ends, which is taken as the user service coverage rate of the current time slot.
[0111] The third computational subunit is used to calculate the rewards for all agents in the current time slot based on the total system throughput and user service coverage of the current time slot.
[0112] Optionally, the fourth determining subunit is specifically used for: The minimum of the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user are taken as the throughput of the target terminal user in the current time slot.
[0113] The difference between the target end user's service demand and the throughput in the current time slot is taken as the target end user's remaining service demand after the current time slot ends.
[0114] Optionally, the formula for the downlink channel capacity of the target terminal user in the current time slot is: ;in, Indicates the target end user Downlink channel capacity in the current time slot, Indicates the target end user With beam The first binary variable that represents the overlay relationship between them. This represents the bandwidth of a single subcarrier. Indicates beam Assigned to standard end users The number of subcarriers, Indicates beam To the center position of the selected beam The first channel gain, , Indicates beam To the center of the beam Transmit antenna gain in the direction of direction Indicates free space transmission loss. Indicates the user's receiving antenna gain. This indicates the transmit power of each subcarrier. Represents the noise power spectral density. Indicates the beam used to reflect With beam The second binary variable representing the spatial interference relationship between them. Indicates beam Interference beams with spatial interference relationship to beam Selected center position The second channel gain, Indicates other beams and the target end user Number of target subcarriers subject to spectral interference. This indicates the total number of beams in a multi-beam satellite communication system.
[0115] Optionally, the reward formula for all agents is: ;in, express Rewards for all agents in the time slot, express Total system throughput of time slots , express Time slot target end user throughput, This indicates the total number of end users in a multi-beam satellite communication system. express User service coverage of time slots This indicates the preset weight parameters. and All represent the homogenized reference values.
[0116] Optionally, the hybrid network model is a supernetwork driven by the global state, and the agent's local Q-value is enhanced with attention before being input into the supernetwork.
[0117] The formula for the global Q-value output by the hybrid network is: ;in, This represents the global state of time slot t. This represents the set of actions of all agents in time slot t. Represents the network parameters of the hybrid network model. , , This represents the local Q-value of agent k. This indicates the total number of beams in a multi-beam satellite communication system. This indicates the preset attention enhancement processing function. and Both indicate that the hypernetwork is based on global state. Determined weights, and Both indicate that the hypernetwork is based on global state. Determined bias.
[0118] Example 3 See Figure 6 This invention provides an electronic device, which includes a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected via the bus 62. The processor 60 is used to execute executable modules, such as computer programs, stored in the memory 61.
[0119] The memory 61 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0120] Bus 62 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0121] The memory 61 is used to store programs. After receiving an execution instruction, the processor 60 executes the program. The method executed by the apparatus defined by the process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.
[0122] Processor 60 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 60 or by instructions in software form. Processor 60 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 61. Processor 60 reads the information in memory 61 and, in conjunction with its hardware, completes the steps of the above method.
[0123] The computer program product of the multi-beam satellite resource optimization method based on multi-agent reinforcement learning provided in this embodiment of the invention includes a computer-readable storage medium storing non-volatile program code executable by a processor. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0124] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0125] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0126] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0127] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0128] Furthermore, terms such as "horizontal," "vertical," and "sag" do not imply that components must be absolutely horizontal or suspended, but rather that they can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0129] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-beam satellite resource optimization method based on multi-agent reinforcement learning, characterized in that, include: Acquire the observation-action history trajectory information of a target beam in the current time slot of a multi-beam satellite communication system; wherein, the target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global state in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global state includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: beam center position and 3dB beamwidth; The observation-action history trajectory information of the target beam is processed using a target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot. In this multi-beam satellite communication system, each beam corresponds one-to-one with each agent in the target multi-agent reinforcement learning model trained using the QMIX network architecture, and each agent has an independent beam resource optimization model. The state of an agent is the observation-action history trajectory information of the corresponding beam, and the action of an agent is the beam coverage strategy of the corresponding beam. The rewards for all agents are positively correlated with the total system throughput and user service coverage. The simulated annealing algorithm is used to perform OFDMA subcarrier resource allocation processing on the beam coverage strategy and global state of all beams in the current time slot, resulting in the spectrum allocation result for each beam in the current time slot. The spectrum allocation result includes: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier. The user to be served refers to the user within the beam coverage area. Based on the beam coverage strategy and spectrum allocation results of all beams in the current time slot, the resource optimization strategy of the multi-beam satellite communication system in the current time slot is determined.
2. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 1, characterized in that, The QMIX network architecture includes: a multi-agent reinforcement learning model and a hybrid network model; the method also includes: Initialize the global state; Repeat the following steps, and every preset number of rounds, sample a preset number of experience data from the experience replay pool to jointly update the network parameters of the multi-agent reinforcement learning model and the network parameters of the hybrid network model using the mean squared error loss function, until a specified number of rounds is reached, or the mean squared error loss function value is less than a preset threshold. In the multi-agent reinforcement learning model, each agent selects an action based on its current state using an ε-greedy policy. The simulated annealing algorithm is used to perform OFDMA subcarrier resource allocation processing on the global state and the actions of all agents to obtain the spectrum allocation result of the beam corresponding to each agent in the current time slot; Based on the global state, the actions of all agents, and the spectrum allocation results of all beams, calculate the rewards for all agents and determine the next state for each agent. Store the state, actions, rewards, and next state of all agents in the experience replay pool.
3. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 2, characterized in that, Based on the global state, the actions of all agents, and the spectrum allocation results of all beams, the rewards for all agents are calculated, including: The location information, multi-beam antenna radiation pattern, free-space transmission loss, and user receiving antenna gain of the target terminal user are obtained; wherein, the target terminal user refers to any user among all the terminal users; Based on the location information of the target terminal user and the actions of all the intelligent agents, a first binary variable is determined to reflect the coverage relationship between the target terminal user and each beam in the multi-beam satellite communication system. Based on the spectrum allocation results of all beams, determine the number of target subcarriers in other beams that have spectrum interference with the target terminal user; Based on the actions of all the agents, a second binary variable is determined to reflect the spatial interference relationship between beams; Based on the multi-beam antenna radiation pattern, the free-space transmission loss, and the user receiving antenna gain, calculate the first channel gain of each beam to the selected beam center position, and the second channel gain of the interfering beam that has spatial interference with the target beam to the target beam center position; wherein, the target beam center position represents the selected beam center position of the target beam. Based on the first binary variable, the number of target subcarriers, the second binary variable, the first channel gain, the second channel gain, and the spectrum allocation result, the downlink channel capacity of the target terminal user in the current time slot is calculated; Based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the current time slot ends are determined. The sum of the throughput of all end users is taken as the total system throughput of the current time slot; The total number of end users whose remaining service demand is 0 after the current time slot ends will be used as the user service coverage rate of the current time slot. Calculate the rewards for all agents in the current time slot based on the total system throughput and user service coverage of the current time slot.
4. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 3, characterized in that, Based on the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user in the global state, the throughput of the target terminal user in the current time slot and the remaining service demand of the target terminal user after the current time slot ends are determined, including: The minimum value between the downlink channel capacity of the target terminal user in the current time slot and the service demand of the target terminal user is taken as the throughput of the target terminal user in the current time slot. The difference between the service demand of the target terminal user and the throughput in the current time slot is taken as the remaining service demand of the target terminal user after the current time slot ends.
5. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 3, characterized in that, The formula for calculating the downlink channel capacity of the target terminal user in the current time slot is: ; in, Indicates the target end user Downlink channel capacity in the current time slot, Indicates the target end user With beam The first binary variable that represents the overlay relationship between them. This represents the bandwidth of a single subcarrier. Indicates beam Assigned to standard end users The number of subcarriers, Indicates beam To the center position of the selected beam The first channel gain, , Indicates beam To the center of the beam Transmit antenna gain in the direction of direction This represents the free space transmission loss. This indicates the gain of the user's receiving antenna. This indicates the transmit power of each subcarrier. Represents the noise power spectral density. Indicates the beam used to reflect With beam The second binary variable representing the spatial interference relationship between them. Indicates beam Interference beams with spatial interference relationship to beam Selected center position The second channel gain, Indicates other beams and the target end user Number of target subcarriers subject to spectral interference. This indicates the total number of beams in the multi-beam satellite communication system.
6. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 3, characterized in that, The formula for the rewards of all agents is: ; in, express Rewards for all agents in the time slot, express Total system throughput of time slots , express Time slot target end user throughput, This represents the total number of terminal users in the multi-beam satellite communication system. express User service coverage of time slots This indicates the preset weight parameters. and All represent the homogenized reference values.
7. The multi-beam satellite resource optimization method based on multi-agent reinforcement learning according to claim 2, characterized in that, The hybrid network model is a super network driven by the global state, and the agent's local Q-value undergoes attention enhancement processing before being input into the super network; The formula for the global Q-value output by the hybrid network is: ;in, This represents the global state of time slot t. This represents the set of actions of all agents in time slot t. This represents the network parameters of the hybrid network model. , , This represents the local Q-value of agent k. This indicates the total number of beams in the multi-beam satellite communication system. This indicates the preset attention enhancement processing function. and Both indicate that the hypernetwork is based on global state. Determined weights, and Both indicate that the hypernetwork is based on global state. Determined bias.
8. A multi-beam satellite resource optimization device based on multi-agent reinforcement learning, characterized in that, include: The acquisition module is used to acquire the observation-action history trajectory information of a target beam in the current time slot in a multi-beam satellite communication system; wherein, the target beam represents any beam in the multi-beam satellite communication system; the observation-action history trajectory information includes: the global state in the current time slot and the beam coverage strategy of the target beam in the previous time slot; the global state includes: the set of service demands of all terminal users in the multi-beam satellite communication system; the beam coverage strategy includes: beam center position and 3dB beamwidth; The processing module is used to process the observation-action history trajectory information of the target beam using the target beam resource optimization model to obtain the beam coverage strategy of the target beam in the current time slot. Each beam in the multi-beam satellite communication system is deployed one-to-one with each agent in the target multi-agent reinforcement learning model trained using the QMIX network architecture, and each agent has an independent beam resource optimization model. The state of the agent is the observation-action history trajectory information of the corresponding beam, and the action of the agent is the beam coverage strategy of the corresponding beam. The rewards of all agents are positively correlated with the total system throughput and user service coverage. The allocation module is used to perform OFDMA subcarrier resource allocation processing on the beam coverage strategy and global state of all beams in the current time slot using the simulated annealing algorithm, so as to obtain the spectrum allocation result of each beam in the current time slot; wherein, the spectrum allocation result includes: the number of subcarriers allocated to each user to be served and the frequency band position corresponding to each subcarrier; the user to be served refers to the user within the beam coverage area; The determination module is used to determine the resource optimization strategy of the multi-beam satellite communication system in the current time slot based on the beam coverage strategy and spectrum allocation results of all beams in the current time slot.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the multi-beam satellite resource optimization method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the multi-beam satellite resource optimization method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.