Frequency spectrum access method and system for large-scale hybrid constellation
By employing a locally interactive Markov game model and a deep Q-network in a hybrid constellation of GEO and LEO satellites within the LEO satellite network, the dynamic nature of spectrum access strategies in the LEO satellite network is addressed, maximizing ground user satisfaction and improving the cumulative performance of spectrum access.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies cannot effectively solve the problem of dynamic spectrum access strategies caused by rapid location changes in LEO satellite networks, and traditional methods are difficult to solve by introducing cumulative effects across consecutive time slots.
A large-scale hybrid constellation spectrum access system is constructed by using a hybrid constellation based on multiple GEO and LEO satellites and combining the Local Interactive Markov Game Model (LIMGM) and the Deep Q Network (DQN) to achieve dynamic allocation of downlinks and maximize terrestrial user satisfaction.
It enables dynamic spectrum access decisions within continuous time slots, maximizes the perceived satisfaction of ground users, solves the problems of dynamism and cumulative effects in satellite networks, and improves the cumulative performance of spectrum access.
Smart Images

Figure CN121664259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communications, and more specifically to a spectrum access method and system for a large hybrid constellation. Background Technology
[0002] The high-speed motion of low Earth orbit (LEO) satellites makes antenna resource optimization in LEO satellite networks a complex decision-making problem involving multiple time slots. The rapid changes in the relative positions between LEO satellites and ground users necessitate real-time adjustments to the downlink spectrum access strategy. Currently, there is no method to handle this dynamic nature. While dividing system time into continuous static time slots and treating the network topology within each slot as static simplifies the optimization problem within a single time slot, it introduces cumulative effects across multiple time slots, which are also difficult to address. Summary of the Invention
[0003] This invention provides a spectrum access method and system for large-scale hybrid constellations, which can solve the technical problems of the prior art.
[0004] To achieve the above objectives, in one aspect, embodiments of the present invention provide a spectrum access method for large hybrid constellations, comprising:
[0005] Step 1: Provide downlinks to ground users within the coverage area using a hybrid constellation consisting of multiple GEO satellites and multiple LEO satellites, where the LEO satellites use a frequency within the frequency set of the GEO satellites; and construct a dynamic allocation problem for the downlink of each ground user based on the satisfaction of the ground users.
[0006] Step 2: Construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlinks based on the downlinks of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium.
[0007] Step 3: The Local Interaction Markov Game Model (LIMGM) is mapped to the Local Influence Graph Model (LIMG), and combined with the Deep Q Network (DQN) to obtain a graph-based Deep Q Network. The graph-based Deep Q Network is then trained to obtain the trained graph-based Deep Q Network.
[0008] Step 4: Using the trained graph-based deep Q-network GDQN, solve the problem of dynamically allocating downlinks for each ground user based on their satisfaction with the ground user, and obtain a dynamic allocation scheme for downlinks for each ground user.
[0009] On the other hand, embodiments of the present invention provide a spectrum access system for a large hybrid constellation, comprising:
[0010] The basic model building unit is used to build a system model to provide downlinks to ground users in the coverage area through a hybrid constellation of multiple GEO satellites and multiple LEO satellites. The LEO satellites use a frequency within the frequency set of GEO satellites. The dynamic allocation problem of downlinks for each ground user is constructed based on the satisfaction of ground users.
[0011] The first model conversion unit is used to construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlink based on the downlink of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium.
[0012] The second model conversion unit is used to map the Local Interaction Markov Game Model (LIMGM) to the Local Influence Graph Model (LIMG), and combine it with the Deep Q Network (DQN) to obtain a graph-based Deep Q Network. The graph-based Deep Q Network is then trained to obtain the trained graph-based Deep Q Network.
[0013] The solution unit is used to solve the problem of dynamic allocation of downlinks for each ground user based on the satisfaction of ground users by using the trained graph-based deep Q-network GDQN, and obtain the dynamic allocation scheme of downlinks for each ground user.
[0014] The above technical solution has the following beneficial effects: it can solve the dynamic allocation of downlinks with continuous time slots for ground users, and can maximize the perceived satisfaction of each ground user. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of a spectrum access method for a large hybrid constellation according to an embodiment of the present invention;
[0017] Figure 2 This is a structural diagram of a large-scale hybrid constellation spectrum access system according to an embodiment of the present invention;
[0018] Figure 3 This is a spectrum sharing scenario in a hybrid constellation according to an embodiment of the present invention;
[0019] Figure 4 This is the downlink interference model of the hybrid constellation in this embodiment of the invention;
[0020] Figure 5 This is a comparison of the convergence of different schemes in the embodiments of the present invention;
[0021] Figure 6 This is a hybrid GDQN framework according to an embodiment of the present invention;
[0022] Figure 7 This refers to the distribution of LU and GS in the hybrid simulation scenario of this invention embodiment;
[0023] Figure 8 This is the relationship between the hybrid spectrum usage balance and the number of proxies in this embodiment of the invention;
[0024] Figure 9 This is the relationship between the hybrid average perceived user satisfaction and the number of agents in an embodiment of the present invention;
[0025] Figure 10 This describes the relationship between the hybrid network fairness index and the number of agents in an embodiment of the present invention.
[0026] Figure 11 The hybrid exploration rate and in this embodiment of the invention Relationship;
[0027] Figure 12 The hybrid exploration rate and in this embodiment of the invention Relationship;
[0028] Figure 13 This is a hybrid average user perceived satisfaction and... (The sentence is incomplete and requires more context to translate accurately.) The relationship. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Due to the highly dynamic nature of satellite-terrestrial network topology, this problem constitutes a complex continuous time-slot spectrum access decision-making process. Traditional distributed algorithms are not suitable for dynamic decisions in continuous time slots, while existing multi-agent deep reinforcement learning (MADRL) methods, although capable of handling such problems, lack theoretical convergence guarantees.
[0031] To overcome these limitations, this invention provides a method for asymptotically optimal multi-agent deep reinforcement learning spectrum access for large-scale LEO constellations based on Markov games. First, the problem is formulated as a Locally Interactive Markov Game (LIMGM) model, and it is proven to be an exact potential game (EPG) with at least one pure policy Nash equilibrium (NE). Second, a graph-based Deep Q-Network (GDQN) framework is used to derive the equilibrium solution of the LIMGM. Finally, using stochastic approximation theory, the convergence and asymptotic optimality of the graph-based Deep Q-Network (GDQN) algorithm are theoretically proven, thus filling a key theoretical gap in MADRL applications for satellite networks. Simulation results demonstrate that the method of this invention has significant effects on optimizing the cumulative performance of continuous time-slot spectrum access decisions.
[0032] like Figure 1 As shown, in conjunction with embodiments of the present invention, a spectrum access method for a large hybrid constellation is provided, comprising:
[0033] Step 1: Provide downlinks to ground users within the coverage area using a hybrid constellation consisting of multiple GEO satellites and multiple LEO satellites, where the LEO satellites use a frequency within the frequency set of the GEO satellites; and construct a dynamic allocation problem for the downlink of each ground user based on the satisfaction of the ground users.
[0034] Step 2: Construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlinks based on the downlinks of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium.
[0035] Step 3: The Local Interaction Markov Game Model (LIMGM) is mapped to the Local Influence Graph Model (LIMG), and combined with the Deep Q Network (DQN) to obtain a graph-based Deep Q Network. The graph-based Deep Q Network is then trained to obtain the trained graph-based Deep Q Network.
[0036] Step 4: Using the trained graph-based deep Q-network GDQN, solve the problem of dynamically allocating downlinks for each ground user based on their satisfaction with the ground user, and obtain a dynamic allocation scheme for downlinks for each ground user.
[0037] like Figure 2 As shown, in conjunction with embodiments of the present invention, a large-scale hybrid constellation spectrum access system is provided, comprising:
[0038] The basic model building unit 21 is used to build a system model to provide downlinks to ground users in the covered area through a hybrid constellation including multiple GEO satellites and multiple LEO satellites. The LEO satellites use a frequency within the frequency set of GEO satellites. The dynamic allocation problem of downlinks for each ground user is constructed based on the satisfaction of ground users.
[0039] The first model conversion unit 22 is used to construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlink based on the downlink of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium.
[0040] The second model conversion unit 23 is used to correspond the Local Interaction Markov Game Model (LIMGM) to the Local Influence Graph Model (LIMG), and combine it with the Deep Q Network (DQN) to obtain a graph-based deep Q network. The graph-based deep Q network is then trained to obtain the trained graph-based deep Q network.
[0041] The solution unit 24 is used to solve the problem of dynamic allocation of downlinks for each ground user based on the satisfaction of ground users by using the trained graph-based deep Q network GDQN, and obtain a dynamic allocation scheme for downlinks for each ground user.
[0042] Preferably, step 1 includes step 1.1:
[0043] Downlinks are provided to terrestrial users within the coverage area via a hybrid constellation consisting of three GEO satellites and a group of LEO satellites configured in a Walker constellation. Each LEO satellite can dynamically adjust the beam direction and frequency of the downlink to serve terrestrial users within the coverage area, wherein:
[0044] The receiving antenna of the GEO ground station GS scans data packets within the available frequency range and identifies the current downlink frequency of the GEO satellite by decoding the flag bits of the data packets;
[0045] At the beginning of each time slot t, the LEO satellite senses the downlink frequency of the GEO satellite and determines the frequency of one or more downlinks within the frequency set of the GEO satellite used by the LEO satellite.
[0046] Each LEO satellite includes multiple beams, with each beam corresponding to a frequency, and each beam is used as a transmitting antenna;
[0047] Within each time slot t, the downlink frequency of the GEO satellite is switched once, and this frequency switching process follows a Markov model.
[0048] Each ground user accesses the downlink of the LEO satellite and receives the beam of the downlink through a receiving antenna;
[0049] The system model is constructed through the above steps.
[0050] The aforementioned additional technologies are also specific embodiments of the basic model building unit 21.
[0051] Preferably, step 1 further includes step 1.2:
[0052] The gain of each transmit antenna of the LEO satellite is determined based on the maximum gain of the LEO satellite's transmit antenna, the transmitter off-axis angle, the far sidelobe level, and the co-channel interference between two ground users accessing the same downlink.
[0053] The gain of the ground user's receiving antenna is determined based on the maximum gain of the ground user's receiving antenna and the co-channel interference between two ground users accessing the same downlink.
[0054] An antenna model is constructed by measuring the gain of the transmitting antenna and the gain of the receiving antenna.
[0055] The aforementioned additional technologies are also specific embodiments of the basic model building unit 21.
[0056] Preferably, step 1 further includes step 1.3:
[0057] Within time slot t, an interference model for the hybrid constellation to provide downlink access to ground users is determined, the interference model including interference from GEO satellites to ground users. Co-channel interference, LEO satellites to ground users The aggregation interference from LEO satellites and the aggregation interference from LEO satellites to the GEO ground station GS, among which:
[0058] GEO satellite for ground users Co-channel interference refers to: when and When using the same frequency within time slot t right Co-channel interference It depends on the antenna transmit power of the GEO satellite and the ground users of the GEO satellite. The off-axis angle, This indicates the downlink connection between the ground station GS and the GEO satellite. Indicates ground users Access to the downlink of LEO satellites;
[0059] LEO satellite for ground users Co-channel interference refers to the interference that occurs when multiple LEO satellites use the same frequency to serve ground users in adjacent areas. This interference occurs between the LEO satellites on the same frequency within time slot t. Aggregation interference It depends on the antenna transmission power of the LEO satellite and the LEO satellite and its ground users. The antenna gain between the LEO satellite's transmitting antenna and the ground user's antenna gain. The gain of the receiving antenna;
[0060] LEO satellite aggregation interference to GEO ground station GS refers to the aggregation interference caused by the downlink of the LEO satellite to the GEO ground station GS when the LEO satellite and GEO satellite share the same frequency. This aggregation interference occurs within time slot t. It depends on the antenna transmit power of the LEO satellite and the antenna gain between the LEO satellite and the ground station GS, which includes the gain of the LEO satellite's transmit antenna and the gain of the ground station GS's receive antenna.
[0061] The aforementioned additional technologies are also specific embodiments of the basic model building unit 21.
[0062] Preferably, step 1 further includes step 1.4:
[0063] Within time slot t, ground users are connected via GEO satellite. Co-channel interference, LEO satellites to ground users Aggregation interference, ground users The downlink channel noise of the access, as well as the antenna transmit power of the LEO satellite and ground users. Access downlink Channel gain, calculate ground user The signal-to-noise ratio received by the receiving antenna;
[0064] Based on ground users The signal-to-noise ratio received by the receiving antenna and the downlink The frequency bandwidth is used to calculate the downlink frequency bandwidth. communication capacity ;
[0065] Ground users Perceived satisfaction is used as an indicator to balance communication capacity. With traffic demand threshold The gap between them.
[0066] The aforementioned additional technologies are also specific embodiments of the basic model building unit 21.
[0067] Preferably, step 1 further includes step 1.5:
[0068] ground users The dynamic allocation problem of downlink is described as follows: within time slot t, based on the allocation to ground users... downlink frequency Optimize to include ground users To maximize perceived satisfaction; and at the same time satisfy the set constraints, which are: the aggregated interference of LEO satellites to GEO ground station GS is not less than the maximum interference threshold that does not affect the downlink communication of GEO satellites.
[0069] The aforementioned additional technologies are also specific embodiments of the basic model building unit 21.
[0070] Preferably, step 2 includes:
[0071] Based on the downlink of ground users, a Locally Interactive Markov Game (LIMGM) model is constructed to address the dynamic allocation problem of downlinks. The LIMGM model consists of seven tuples. ,in:
[0072] This represents the set of game participants, where game participants refer to ground users. Downlink access to LEO satellites ;
[0073] Represents the set of state spaces. S represents the Cartesian product of the states of all game participants. n Ground users State space, ground users State space refers to ground users The perceived satisfaction of ground users and their neighbors, as well as the ground station GS and ground users The frequency of their respective downlinks;
[0074] This represents the set of actions of the game participants, where each action refers to the downlink. The frequency of the downlink The frequencies are selected from the set of available downlink frequencies;
[0075] This represents the set of neighbors of a game participant. A game participant's neighbors are those who, if the ground user... and ground users The distance between them is no greater than the interference distance threshold. Then determine the ground user Ground users Neighbors;
[0076] It is the state transition probability function, expressed as Where s and s' represent the current and new states of the downlink frequency for ground users accessing LEO satellites, respectively. This represents the actions of all game participants at time slot t;
[0077] It is the reward function for the participants in the game;
[0078] It is a discount factor used to measure the weight of immediate rewards and future rewards. .
[0079] The aforementioned additional technology is also a specific embodiment of the first model conversion unit 22.
[0080] Preferably, step 3 includes:
[0081] The edges in the Local Influence Graph (LIMG) model correspond to the neighbor relationships in the Local Interaction Markov Game (LIMGM) model, where the neighbor relationships represent the interference relationships between game participants.
[0082] The state space of each player in the game includes the player's state space. Local state information Game participants Interacting neighbor game participants Local state information Composition, game participants The state space is represented as:
[0083] (31)
[0084] in, Indicates the participants in the game The set of participants in the neighbor game. Indicates the participants in the game Neighbor game participants, , Indicates the participants in the game Local state information, Indicates the participants in the neighbor game. Local state information;
[0085] DQN is used to approximate each game participant. Action value function The participants in the game Current state and actions As input, through game participants reward function The output is an action. Expected rewards ;
[0086] The above steps yield a deep Q-network based on a graph model;
[0087] During the training of a graph-based deep Q-network, each game participant... use Greedy strategies rely on probability. Choose a random action, with probability Select the action with the highest value based on the current action value function;
[0088] As training progresses, The loss function is gradually decayed to reduce the exploration frequency until it is lower than the preset loss value, thus obtaining the trained graph model-based deep Q network. The convergence and asymptotic optimality of the deep Q network are proved using stochastic approximation theory.
[0089] The aforementioned additional technology is also a specific embodiment of the second model conversion unit 23.
[0090] Preferably, step 4 includes:
[0091] Each player in the game First, observe your current local state information. and the local state information of its interacting neighboring intelligent agents Based on a pre-trained graph-based deep Q-network GDQN, game participants use - A greedy strategy selects actions that cause the game participants to... Select the action with the highest value based on the current action value function:
[0092] (34)
[0093] Game participants Perform the selected action and observe the environmental response, which includes the game participants. Instant rewards and the next state And the game participants obtained through dynamic monitoring Neighbor game participants Actions and rewards;
[0094] And update the game participants based on the environmental response. Local impact diagram To reflect the latest interactive information;
[0095] Until the average perceived satisfaction of all game participants is maximized, the downlink of each corresponding ground user is used as a time slot t dynamic allocation scheme.
[0096] The aforementioned additional technology is also a specific embodiment of the solving unit 24.
[0097] The technical solutions of the present invention will be described in detail below with reference to specific application examples. For technical details not described in the implementation process, please refer to the relevant descriptions above.
[0098] I. System Model and Problem Formalization
[0099] like Figure 3 As shown, the hybrid constellation comprises three geostationary orbit (GEO) satellites and a group of low Earth orbit (LEO) satellites configured using the Walker constellation. The GEO satellites are in geostationary orbit at an altitude of approximately 35,786 km, while the LEO satellites operate in low Earth orbit at an altitude of approximately 550 km. Adjacent LEO satellites are interconnected via inter-satellite links, and the GEO satellites connect to the LEO constellation via ground gateways. Each LEO satellite is equipped with a multi-beam phased array antenna, capable of dynamically adjusting beam direction and frequency to serve ground users within its coverage area. In contrast, the GEO satellites use high-gain parabolic antennas with narrower beamwidths, optimized for long-distance communication.
[0100] Table 1 Parameter Description
[0101] 1. System Model
[0102] like Figure 3 As shown, the area jointly covered by the GEO and LEO satellite constellations includes one GEO ground station (GS) and multiple randomly distributed ground user groups (LUs). The set of ground user groups (LUs) is defined as... The set of LEO satellites is represented as , This represents the k-th LEO satellite. Each LEO satellite can generate a maximum of [number missing]. Each satellite has one beam, acting as a transmitter, capable of tracking and serving one ground user unit (LU). Each LU is equipped with a receiver to receive downlink signals. The transmit power of the GEO satellite beams (antennas) and the LEO satellite beams (antennas) are fixed at [values to be filled in]. and In this embodiment of the invention, a phased array antenna is used, and each beam has no physical antenna. It can be understood that the phased array antenna is a virtual antenna power for each beam.
[0103] To ensure communication security, the downlink frequency (i.e., the frequency of the carrier used in the downlink) between GEO satellites and GEO ground stations (GS) adopts a dynamic access mechanism. The carrier is the frequency signal used in the downlink between GEO and LEO satellites; it is a radio signal at a specific frequency used in satellite communication and is the basic unit of spectrum resources. The set of available downlink frequencies is represented as... Each carrier (beam) occupies a bandwidth of .time Divided into equal-length time slots t, i.e. Within each time slot t, the downlink frequency of the GEO satellite switches once. The frequency switching process follows a Markov model, and each downlink selects a frequency for communication. There is a one-to-one relationship between the downlink and the frequency. The receiver at the GEO ground station GS scans for data packets within the available frequency range and identifies the current downlink frequency of the GEO satellite by decoding the flag bits of the data packets.
[0104] In a hybrid constellation, LEO satellites share the frequency set of GEO satellites for their downlinks to maximize spectrum utilization. At the beginning of each time slot t, the LEO satellite first senses the downlink frequencies of the GEO satellites and then determines the frequency it will use for its downlink, i.e., determining a specific frequency within the GEO satellite frequency set to be adopted by the LEO satellite. By constructing a dynamic decision-making method for LEO satellite downlink frequency selection, the aim is to minimize inter-downlink interference and improve the average perception satisfaction of the ground user base (LUs).
[0105] 2. Antenna Model Protocol
[0106] Each beam of the LEO satellite is an independent transmitting antenna (transmitter), and its radiation pattern conforms to the ITU-RS.1528 standard. The gain of the LEO satellite's transmitting antenna (transmitter) is:
[0107] (1)
[0108] in, This represents the off-axis angle of the LEO satellite's transmitting antenna in time slot t, i.e. and The angle between them This represents the ground user in time slot t. Downlink access Link to downlink Co-channel interference (within time slot t) right Co-channel interference), with the superscript u indicating the u-th ground user. The subscript n represents the nth ground user. , Indicates ground users The downlink of the access; the subscript n indicates ; It is a 3 dB beamwidth Half of This indicates the point where the intersection of the main beam and the near-side lobe mask is lower than the peak gain. Indicates the level of the far sidelobe. .
[0109] Maximum gain of the transmitting antenna Represented as:
[0110] (2)
[0111] The parameter Z is represented as:
[0112] (3)
[0113] Each ground user LU is equipped with a receiving antenna (receiver) whose radiation pattern conforms to the ITU-R S.465-6 standard. The ground user receives the signal, and the gain of the receiving antenna is:
[0114] (4)
[0115] Among them, the superscript in formula (4) Indicates the first ground users subscript Indicates the first ground users ; Where D is the diameter of the receiving antenna aperture for the ground user, and λ represents the carrier wavelength of the LEO satellite. The maximum gain of the receiving antenna... Represented as:
[0116] (5).
[0117] 3. Interference Model
[0118] To calculate the signal-to-noise ratio (SNR) received by the GEO ground station GS and the ground user group LUs in the system. Figure 4 The downlink interference model is illustrated. Within time slot t, the co-channel interference experienced by the GEO ground station GS and the ground user group LUs includes interference from the GEO satellite to the ground users. Co-channel interference, LEO satellites to ground users The aggregation interference of LEO satellites and the aggregation interference of LEO satellites to the GEO ground station GS.
[0119] 1) GEO satellites to ground users Co-channel interference: when ( (This indicates the downlink of the ground station GS accessing the GEO satellite) and ( Indicates ground users When the downlink accessing the LEO satellite uses the same frequency within time slot t, right Co-channel interference It depends on the antenna transmit power of the GEO satellite and the GEO satellite antenna's position relative to the LEO satellite. ground users The off-axis angle.
[0120] (6)
[0121] in, This indicates the beam of the GEO satellite to ground users. Co-channel interference links, In Indicates the first ground users g represents a GEO satellite;
[0122] This indicates the beam antenna transmit power of the GEO satellite.
[0123] This indicates that within time slot t, GEO satellites... Antenna interference gain:
[0124] (7)
[0125] (8)
[0126] In equations (7) and (8), the subscripts are... Indicates the first ground users The subscript 'g' indicates a GEO satellite; This represents a binary indicator function used to determine the relationship between ground station GS and the time slot t. ground users Whether to use the same frequency, This indicates the downlink frequency of the ground station GS accessing the GEO satellite. Indicates ground users The frequency of the downlink access to LEO satellites; This indicates the gain of the transmitting antenna (transmitter) of a GEO satellite. This indicates the gain of the receiving antenna when a ground user receives downlink data from a GEO satellite. This indicates the off-axis angle of the transmitting antenna of a GEO satellite. Among them, express Euclidean distance.
[0127] 2) LEO satellites to ground users Co-channel interference: When multiple LEO satellites use the same frequency to serve adjacent ground user groups (LUs), co-channel interference will occur between the multiple LEO satellites. Within time slot t, LEO satellites interfere with ground users... Aggregation interference It depends on the antenna transmission power of the LEO satellite and the nth ground user of the LEO satellite. The antenna gain between (including the LEO satellite's transmitting antenna and the nth ground user) (receiving antenna).
[0128] (9)
[0129] In formula (9), This represents the number of ground users for LEO satellites, with the subscript u indicating the u-th ground user for the LEO satellite. The subscript n indicates the nth ground user of the LEO satellite. , The mark in Indicates LEO satellite; Indicates the LEO satellite antenna transmit power, subscript Indicates LEO satellite; Indicates ground users The downlink access to ground users Interference gain of the accessed downlink; Define something similar to equation (8) A binary indicator function; where:
[0130] (10)
[0131] in, This indicates the transmit antenna gain of the LEO satellite, where the subscript... This refers to a LEO satellite; the superscript tran indicates transmission. (Parameters) Indicates the first ground users To the ground users The off-axis angle of the link relative to the LEO satellite's antenna; Indicates the first ground users Receive antenna gain, subscript : indicates the first ground users The superscript "rece" indicates receiving. Indicates the first ground users The signal relative to the first ground users Off-axis angle of the receiving antenna; Indicates within time slot t right Co-channel interference Euclidean distance.
[0132] 3) Aggregation interference from LEO satellites to GEO ground station GS: When LEO and GEO satellites share the same frequency, the downlink signals from the LEO satellites will cause aggregation interference to the GEO ground station GS. Aggregation interference from LEO satellites to ground station GS occurs within time slot t. It depends on the transmit power of the LEO satellite's antenna and the antenna gain between the LEO satellite and the ground station GS (the transmit antenna of the LEO satellite and the receive antenna of the ground station GS).
[0133] (11)
[0134] in, Subheadings Indicates LEO satellite; Indicates the LEO satellite antenna transmit power, subscript Indicates LEO satellite; Indicates ground users The interference gain of the access downlink to the downlink of the GEO satellite accessed by the ground station GS; Define something similar to equation (8) A binary indicator function; where:
[0135] (12)
[0136] in, Indicates the transmit antenna gain of the LEO satellite, subscript This refers to a LEO satellite; the superscript tran indicates transmission. (Parameters) The uth ground user The off-axis angle of the LEO satellite antenna relative to the GS direction of the ground station; This indicates the receiving antenna gain of the ground station GS. The subscript GS indicates the ground station, and the superscript Receive indicates Receive. The uth ground user The off-axis angle of the signal relative to the ground station's GS receiving antenna; This indicates that within time slot t, starting from the u-th ground user downlink Co-channel interference to ground station GS, where the subscript u indicates the u-th ground user served by the LEO satellite. ; This represents the u-th ground user served by the LEO satellite. Interference to GEO ground station GS during access Euclidean distance.
[0137] Therefore, the first time within time slot t will be... ground users Received signal-to-noise ratio (SNR) Represented as:
[0138] (13)
[0139] in, This indicates the antenna transmit power of the LEO satellite; This indicates the ground user within time slot t. Access downlink Channel gain, This represents the downlink accessed by ground users within time slot t. Channel noise;
[0140] Downlink within time slot t ( The link communication capacity of ) can be expressed as:
[0141] (14)
[0142] in, This indicates the frequency bandwidth of the downlink.
[0143] In order to measure With traffic demand threshold The gap between them led to the introduction of perceived satisfaction. As an indicator, this metric captures users' personalized needs and preferences. Higher perceived satisfaction indicates better optimization results. ground users Perceived satisfaction Represented as:
[0144] (15)
[0145] The higher the value, the more significant the th... ground users The higher the satisfaction with communication services, the better. The lower the value, the less effectively the requirement is met. Parameter This represents the slope of the adjusted demand utility curve, reflecting... Traffic demand threshold Sensitivity.
[0146] 4. Problem Description
[0147] ground users The goal of optimizing perceived satisfaction is to find the optimal allocation scheme of downlink frequencies for LEO satellites to maximize the average perceived satisfaction of the ground user group (LUs) within time T. This involves assigning ground users... The goal of optimizing perceived satisfaction is described as a problem. :
[0148] (16)
[0149] The constraints are:
[0150] (17)
[0151] in, This indicates the ground user within time slot t. The frequency of the downlink access to LEO satellites, This represents the maximum interference threshold that ensures GEO satellite downlink communication is unaffected. This indicates the number of ground users within the coverage area. Indicates ground users The downlink frequency for accessing LEO satellites is Perceived satisfaction.
[0152] Unlike the static environmental characteristics of terrestrial wireless systems, the satellite system in the real-time example of this invention involves a sequential decision-making problem for the downlink frequency of LEO satellites. "Sequential" refers to a decision-making process that proceeds continuously in chronological order. LEO satellites need to make spectrum access decisions in consecutive time slots (t=1,2,...,T). The decision in each time slot affects the decision space of subsequent time slots; this continuous decision-making across time is called "sequential decision-making." This problem requires consideration of the cumulative effect of variable decisions within consecutive time slots.
[0153] II. Markov Games with Local Interactions Among Multiple Agents
[0154] To reduce communication and computing costs, the problem Modeled as a Locally Interactive Markov Game (LIMGM), the MADRL framework based on LIMGM captures the temporal continuity of the optimization problem. It is proven that LIMGM is an exact potential game (EPG), and its Nash equilibrium (NE) can be achieved through the local information interaction and cooperation of the ground user group (LUs). Theoretical convergence guarantees are provided within the exact potential game (EPG) framework.
[0155] 1. Locally Interactive Markov Game Model (LIMGM)
[0156] Definition 1: The Locally Interactive Markov Game (LIMGM) model is defined as a seven-tuple. ,in:
[0157] This represents the set of game participants, where game participants refer to ground users. Downlink access to LEO satellites , This means all downlinks A set of.
[0158] Represents the set of state spaces. , representing the Cartesian product of the states of all game participants, where game participants refer to ground users. Downlink of the accessed LEO satellite, S n Ground users State space, ground users State space refers to ground users The perceived satisfaction of ground users and their neighbors, as well as the ground station GS and ground users The downlink frequency. The joint state space of time slot t is defined as... ,pass express A certain value, , Indicates the participants in the game In the time slot The historical spatial state, i.e., ground users Historical satisfaction with downlink access to LEO satellites ; Indicates the participants in the game In the time slot The historical state space of all neighboring ground users. (Index) Index representing the participants in the neighbor game. Participant in the game The set of neighbors; This indicates the sensing result of the frequency of the ground station's GS access to the GEO satellite link for time slot t, with the subscript... Indicates GEO satellite; Indicates time slot t pair Sensing results for frequencies accessing LEO satellite links, subscript This refers to the LEO satellite.
[0159] This represents the space of action sets for game participants, where each participant's action refers to the frequency of their downlink choices. Represents all participants in the game. The Cartesian product of the action set. ,and Indicates downlink The set of available frequencies, This represents the set of available frequencies for the downlink.
[0160] Let represent the set of neighbors of the game participants, where ,and Let represent the set of neighbors of game participant n. If and The distance between them is no greater than the interference distance threshold. ,but It is considered The neighbors.
[0161] It is the state transition probability function, expressed as Where s and s' represent the current and new states of frequency selection for all ground users' LU access to the LEO satellite, respectively. This represents the joint action of all game participants at time slot t. The current state refers to the frequency currently used for the downlink of the ground user, and the new state refers to the frequency the link will use in the future.
[0162] It is the set of reward functions (also known as utility functions) of game participants. Specifically, the reward function of game participant n in time slot t is defined as... .
[0163] It is a discount factor used to measure the weight of immediate rewards and future rewards, defined as follows: .
[0164] The Locally Interactive Markov Game (LIMGM) model enhances the flexibility and adaptability of the model while reducing communication overhead and computational resource requirements.
[0165] 2. Local Collaboration Methods in Reward Design
[0166] By introducing local altruistic behavior among neighboring game participants, global optimization is achieved through local optimization. Each game participant not only focuses on its own interests but also considers the interests of its neighbors, thereby promoting more efficient resource allocation and global optimization.
[0167] Reward function of game participant n in time slot t The definition is as follows:
[0168] (18)
[0169] in, Indicates the participants in the game Actions in time slot t, Including the game participants The set of actions of all participants in the game. Indicates ground users Perceived satisfaction in time slot t (i.e., time interval t). Indicates the participants in the game Neighbors Actions in time slot t, Including the game participants Neighbors The set of actions of all external game participants express The sum of neighbors' perceived satisfaction, subscript Including the game participants All other game participants, i.e. ; Subscript Including the game participants All other game participants, i.e. .parameter It is a weighting factor used to balance ground users The perceived satisfaction level of its neighboring ground users.
[0170] To evaluate the long-term performance of game participants, the expected average long-run return (EALR) is defined as follows:
[0171] (19)
[0172] in, Indicates the participants in the game The sequence of actions within time T This indicates that within the same time slot t, excluding game participants... The set of joint actions of all external game participants.
[0173] To reformulate the optimization problem Γ as a game model :
[0174] (20)
[0175] To describe the problem With game theory models The relationship between them, to prove It is a precise potential game (EPG).
[0176] Definition 2: Action set A game is a pure strategy (NE) if and only if no player can increase its utility by unilaterally changing its current action:
[0177] (twenty one)
[0178] The formal expression is:
[0179] (twenty two)
[0180] in, Represents the actions of the game participants Action sequence, Actions of other game participants A set of action sequences; Indicates the participants in the game The optimal sequence of actions over the entire time period T. Including the game participants Besides, the set of optimal action sequences for all other game participants, This represents the optimal sequence of actions for the first player in the game over the entire time period T. Let represent the optimal sequence of actions for the Nth player in the entire time period T.
[0181] Definition 3: A game is an exact potential game (EPG) if there exists a potential function. This makes it possible for the participants in the game. When the action is from Become When the following conditions are met
[0182] (twenty three)
[0183] in, Indicates the participants in the game An alternative action sequence (different from the current action sequence). Represents the potential function. Represents the set of real numbers. This represents the Cartesian product operator. This indicates a mapping relationship.
[0184] Theorem 1: Game Theory Model It is an exact potential game (EPG) and has at least one pure strategy (NE).
[0185] Proof: Potential function The structure is as follows:
[0186] (twenty four)
[0187] When any participant in a game unilaterally changes their sequence of actions from Become When the individual's expected average long-term return (EALR) changes, it can be calculated as follows:
[0188] (25)
[0189] in, Indicates the participants in the game An alternative action sequence (different from the current action sequence). Indicates the participants in the game Neighbor An alternative action sequence (different from the current action sequence);
[0190] On the other hand, the change of the potential function is:
[0191] (26)
[0192] in, Indicates the use of new actions Post-game participants Perceived satisfaction; This indicates when game participants Changing actions affects neighbors At that time, the neighbor Perceived satisfaction.
[0193] in, ( Indicates the index of game participants. In addition to the game participants The set of all game participants other than its neighbors). Due to the game participants Changes in the action sequence of a player will not affect the decisions of players outside its neighbors. Therefore, it can be deduced that:
[0194] (27)
[0195] in, Indicates the participants in the game The reward function (also known as the utility function).
[0196] Therefore, combining the derivation of the above equations: equation (25), equation (26), and equation (27), we obtain:
[0197] (28)
[0198] therefore, It is an EPG. It has at least one pure strategy Night Elf.
[0199] A strategy refers to the frequency selection sequence for the downlink of each LEO satellite. A pure strategy refers to a deterministic frequency selection (not a probability distribution). A Nash equilibrium state is a stable state in which no agent can increase its utility by unilaterally changing its frequency selection.
[0200] The purpose of pure strategy NE is: (1) to ensure that the system has a stable solution: to ensure that the spectrum allocation scheme will not oscillate indefinitely; (2) to provide a convergence objective: to provide a clear optimization objective for the GDQN algorithm, where the EPG property makes it possible to prove the convergence of the algorithm using stochastic approximation theory; (3) to guarantee theoretical optimality: to prove that there exists a theoretically optimal spectrum allocation scheme.
[0201] The potential function of the EPG can be used as an optimization objective, and the Nash equilibrium can be found by maximizing the potential function.
[0202] For the sequence decision problem of downlink spectrum access in LEO satellite systems, a game-theoretic utility function with continuous time slot accumulation effect is adopted. Compared with the traditional Markov game model, LIMGM has a significant feature: the utility function of each player is jointly determined by the actions of the player itself and its neighboring players. This structural feature effectively captures the local interference effect, similar to the co-channel interference patterns observed in satellite communication systems.
[0203] Theorem 2: Problem The optimal solution constitutes the game. Pure strategy NE.
[0204] Proof: Assume The question raised The optimal solution, which means the game potential function The local or global maximum value has been reached.
[0205] (29)
[0206] in, Indicates within time slot t The carrier frequency of the access downlink, in the EPG, is represented by the potential function. The global or local maximum value corresponds to the pure policy Nash equilibrium (NE).
[0207] (30)
[0208] III. Distributed Multi-Agent Deep Reinforcement Learning Scheme Based on I3DQN Framework
[0209] question Modeled as LIMGM, and the game theory model It is an exact potential EPG, and the problem optimal solution Corresponding to game theory model Pure policy Negative Algebra (NE) is a problem. To solve NE, the multi-agent deep reinforcement learning algorithm GDQN is employed, which can obtain the optimal policy in a distributed and adaptive manner. .
[0210] A multi-agent deep learning network scheme based on GDQN is adopted. The GDQN framework combines the Local Influence Graph Model (LIMG) with Deep Q-Network (DQN, which combines deep neural networks with Q-learning algorithms to solve decision problems in high-dimensional state spaces). Edges in LIMG correspond to neighbor relationships in LIMGM, representing the interference relationships between agents (the agents refer to the downlink established between LEO satellites and ground users; agents are the game participants). Its core idea is to effectively model the interference relationships in LEO satellite beam frequency resource allocation through local interactions between neighboring agents. This allows agents to more accurately estimate the value of their actions. Finally, the convergence and asymptotic optimality of the graph-based Deep Q-Network (GDQN) algorithm are rigorously proven using stochastic approximation theory.
[0211] The graph-based deep Q-network (GDQN) framework effectively captures the time dependency in the downlink frequency sequential decision-making process of LEO satellites, enabling LIMGM to have an equilibrium solution while maintaining low communication and computational overhead.
[0212] untie This refers to the frequency selection strategy for each agent (downlink). Specifically, it is represented from the downlink frequency set as The selected frequency. Implemented for each ground user. Determine the optimal downlink frequency within time slot t. At the same time, it can maximize average perceived satisfaction. .
[0213] 1. Training Phase
[0214] During the training phase of the GDQN scheme, LIMG is constructed based on the neighbor relationships defined in LIMGM. This model, LIMGM, focuses on local interactions, thereby reducing communication and computational costs. The state of each agent (game participant) is determined by its local state information. Local state information of its interacting neighboring intelligent agents Composition, formally represented as:
[0215] (31)
[0216] In the local state information, "local" is relative to "global," containing only information about the agent itself and its neighbors; it does not require obtaining global information about the entire network, reducing communication overhead; small Meaning: Indicates the first The index of the intelligent agent. Corresponding to the ground user. The downlink to which the connection is made. Represents intelligent agents The set of neighboring intelligent agents, intelligent agents Distance less than the interference threshold Other agents, neighbor representation, represent a set of agents that may cause mutual interference. Represents intelligent agents The neighbor agent index, , Represents intelligent agents Local state information, Representing neighboring intelligent agents Local state information.
[0217] in, Represents intelligent agents Local state information. It comprises the following components:
[0218] (1) Historical perceived satisfaction information: Intelligent agent (correspond The perceived satisfaction level in the previous time slot reflects the user's past service quality.
[0219] (2) Frequency-sensing information: The ground station's GS access to the current downlink frequency sensing results of the GEO satellite. This information provides ground users with the frequency information currently used for LEO satellite downlink access. This information helps the agent understand the current spectrum occupancy.
[0220] (3) Link quality related information: current channel gain, noise level, and other key parameters that affect communication quality.
[0221] Representing neighboring intelligent agents Local state information for each neighbor ,Include:
[0222] (1) Neighbors' historical perceived satisfaction: :Neighbor The perceived satisfaction level in the previous time slot helps the intelligent agent. Understand the service situation of your neighbors;
[0223] (2) Neighbors' frequency of use: Neighbors The currently selected frequency is used to assess potential co-channel interference;
[0224] (3) Relative position information: relative to the agent Distance (within the interference distance threshold) (Inner), used to estimate interference intensity.
[0225] The GDQN scheme uses DQN to approximate the action-value function of each agent. The network input includes the agent's current state. and actions The output is the expected reward for the selected action. To enhance the model's robustness and generalization ability, an experience replay buffer is used to store the agent's historical experience. Training is performed by randomly sampling from this buffer, breaking temporal correlation and improving learning stability.
[0226] During training, each agent uses A greedy strategy is used to balance exploration and exploitation. Specifically, the agent uses probability... Choose a random action, with probability The action with the highest value is selected based on the current action value function. As training progresses, Gradual decay is used to reduce exploration frequency and increase utilization. To further stabilize training, GDQN employs a target network. Its parameters are periodically obtained from the main network. Update to calculate the target Q-value. The loss function is defined as:
[0227] (32)
[0228] in, Participant in the game The instant reward value obtained at a certain moment. It is a discount factor. Represents intelligent agents The current state, Represents intelligent agents The current action (frequency selection). This is the next state of the agent. Indicates the candidate action for the next state.
[0229] Game participants The reward function is expressed as:
[0230] (33)
[0231] 2. Execution Phase
[0232] GDQN framework such as Figure 6 As shown, during the execution phase of the GDQN scheme, each agent first observes its current local state information. and the local state information of its interacting neighboring intelligent agents Based on the trained DQN network, the agent uses... - A greedy strategy is used to select actions. In actual deployments, this is typically... Set to a smaller value to minimize the exploration frequency and ensure that the agent primarily selects the highest-value action based on the current action-value function:
[0233] (34)
[0234] The agent performs the selected action and observes the environmental response, which includes an immediate reward. (Calculated according to formula (33) to reflect the quality of the current frequency selection) and the next state After each action, the agent updates its local influence map. This reflects the latest interaction information. The update is achieved by dynamically monitoring the actions and rewards of neighboring agents.
[0235] Among them, the local impact map It is an intelligent agent A graph structure centered on the agent is used to represent local disturbance relationships. The basic elements of a graph are: (1) Nodes, including: Central node: Agent itself (corresponding link) ); Neighbor nodes: all that satisfy intelligent agents (Other links within the interference range). (2) Edges, including those connecting agents. with his neighbors The weight of the edge reflects the interference intensity: calculated based on distance and antenna gain. (3) Attribute information, including the following stored for each node: the frequency currently used, perceived satisfaction, and link status; and the following stored for each edge: interference coefficient (calculated according to formula (10)).
[0236] The dynamic update mechanism requires real-time updates to the local impact map because: LEO satellites are constantly moving, and neighbor relationships will change; neighbor frequency selection will change; and interference intensity changes with distance.
[0237] Interaction information is key data exchanged between intelligent agents, including: (1) Frequency usage information: Current frequency selection: ,Neighbor Selected downlink frequency, frequency occupancy status: which frequencies are currently in use, co-channel interference indication: (Formula (8)), indicating whether there is co-channel interference; (2) Performance feedback information: perceived satisfaction: The neighbor's current service quality, with instant rewards: The reward value obtained by the neighbor; Link quality: Signal-to-noise ratio .
[0238] 3. Topology-related information, including: relative positions: distances between agents. Antenna pointing: Angle information that affects interference .
[0239] During the execution phase, the agent can continue learning by adding new experiences to the experience replay buffer and performing incremental training periodically to adapt to environmental changes and emerging interaction effects. This approach enables the GDQN scheme to maintain high adaptability and learning capacity in dynamic environments, thereby improving overall system performance and stability.
[0240]
[0241]
[0242] 3. Convergence Analysis
[0243] Despite the introduction of neural networks and experience replay mechanisms, the core update process of DQN gradually approaches the optimal Q-function. The training process of DQN can be viewed as an iterative sequence of Q-learning, where each Q-update follows the rules of Q-learning. Subsequently, the convergence and asymptotic optimality of the graph-based deep Q-network (GDQN) algorithm are proved using stochastic approximation theory.
[0244] The purposes of having convergence include: (1) ensuring that the algorithm can converge stably to a solution, rather than diverging or oscillating; (2) ensuring that a feasible spectrum allocation scheme can be found within a finite time; and (3) providing theoretical guarantees to make the algorithm reliable in actual deployment.
[0245] The objectives of asymptotic optimality include: (1) ensuring that the solution the algorithm eventually converges to is the optimal solution, not the suboptimal solution; (2) ensuring that the spectrum allocation scheme found can maximize user satisfaction; and (3) providing a theoretical upper bound guarantee for performance.
[0246] Theorem 3: Convergence of the GDQN Algorithm: Let the game... It is an EPG and meets the following conditions:
[0247] (35)
[0248] Then the Q-function of each DQN network In the number of iterations It converges with probability 1.
[0249] in, State-action pairs represent agents. In state Take action below , This indicates that the reward function is bounded. : No. In the next iteration, the agent Instant rewards : The upper bound constant of the reward function; : No. The learning rate for each iteration; This refers to the sum of learning rates diverging, which ensures that the algorithm can fully explore the state space; This indicates that the sum of squared learning rates converges, which is used to ensure that the variance of the algorithm is bounded and to guarantee convergence.
[0250] Proof: Step 1: Proof
[0251] In EPG, the reward function can be rewritten as:
[0252] (36)
[0253] Time difference (TD) error is defined as:
[0254] (37)
[0255] in, This represents the target Q-network, used to calculate the target Q-value; parameters are periodically extracted from the main network. Copying improves training stability; Indicates the first The next state in the next iteration. Indicates the first The candidate actions for the next iteration, marked with an apostrophe, are variables used for the max operation, not the actual actions to be performed, but rather all possible actions. Indicates all possible actions Maximum value operation.
[0256] Substituting the reward function (35) into the time difference TD error, we get:
[0257] (38)
[0258] Take the conditional expectation:
[0259] (39)
[0260] Because the model satisfies the Markov property:
[0261] (40)
[0262] therefore:
[0263] (41)
[0264] in, Indicates: the The delta algebra, consisting of all historical information up to the last iteration, includes information up to time step [time]. It contains all the information about states, actions, rewards, etc. In stochastic process theory, this is called a filter. This represents the conditional expectation based on historical information.
[0265] Step 2: Proof
[0266] in, It is a bounded constant: , Represents the potential function Bounded constant. The bounds of the Q function, which in English guarantee that the variance of the TD error is bounded, are the key conditions for proving convergence.
[0267] The square of the time difference (TD) error is:
[0268]
[0269] Substituting into the reward function (34), we get:
[0270] (42)
[0271] Expanding the expected value (37), we get:
[0272] (43)
[0273] Where D and M are the bounded constants of the potential function and the Q function, respectively.
[0274] Step 3: Prove convergence using stochastic approximation theory: According to the Robbins-Siegmund theorem, if the learning rate condition and the bounded variance condition are satisfied, the Q-function update converges to the optimal Q-function with probability 1. Therefore:
[0275] (44)
[0276] Equation (42) indicates that the infinite series of the product of the learning rate and the TD error converges to 0, with a probability of 1, meaning almost certain convergence. Mathematical expression: This ensures that the Q-function update will eventually stabilize and no longer change significantly.
[0277] This means that the Q-function update can be viewed as stochastic gradient descent under these conditions. exist It converges with probability 1.
[0278] Theorem 4: [Asymptotic optimality of the GDQN algorithm]: If the same conditions as Theorem 3 are satisfied, then It converges to the optimal Q-function with probability 1. .
[0279] in, The asterisk (*) in the figure represents the optimal Q-function: indicating that in state... Next, take action Then, follow the optimal strategy to obtain the maximum expected cumulative reward.
[0280] Differences from the ordinary Q function: (1) : The currently learned estimate of the Q-function (which may not be optimal), (2) The theoretically optimal Q-function value (the actual optimal value).
[0281] Significance in the LEO satellite scenario, optimal Q function This represents: (1) Optimal frequency selection strategy: can guide the selection of the optimal frequency in any state; (2) Long-term reward maximization: not only considers immediate rewards, but also the cumulative rewards of all future time slots; (3) Nash equilibrium solution: in the game theory framework, the value function corresponding to the Nash equilibrium state. * The symbol is a standard notation in the field of reinforcement learning, used to distinguish between "current estimate" and "theoretical optimal value", and is a key concept for evaluating the convergence and optimality of the algorithm.
[0282] Proof: Define the error of the Q function as:
[0283] (45)
[0284] in, Represents intelligent agents The optimal Q-function is the theoretically optimal action-value function, representing the state at which the action-value is obtained. Take action below Maximum expected cumulative reward: ,error It measures the difference between the current Q-function and the optimal Q-function.
[0285] According to the Q function update rules:
[0286] (46)
[0287] In the Q function update formula Meaning:
[0288] in, The meaning is learning rate, the first The learning step size in each iteration controls the magnitude of each update and determines the weight of new information. Convergence condition must be met.
[0289] Represents the updated Q function value (the first one). (nth iteration) Indicates the current Q function value (the first one). (nth iteration) : No. Temporal difference (TD) error of the next iteration.
[0290] Substituting the error into the Q function and redefining it, we get:
[0291] (47)
[0292] in, : No. Error after the nth iteration It is the error decay coefficient, which controls the retention ratio of historical errors. Formula (40) describes the evolution of error with the number of iterations.
[0293] If the conditions of Theorem 3 are satisfied, the error exist It converges to 0 with probability 1. Therefore:
[0294] (48)
[0295] in, A limit is a complete mathematical statement that includes the following elements: The limit expression: Indicates the number of iterations As it approaches infinity; the convergent object represents the learned Q-function. Convergence objective: Optimal Q-function Convergence property: "With probability 1" indicates almost certain convergence. Its mathematical meaning is... In practical terms, this theorem guarantees that, under certain conditions, after a sufficient number of iterations, the Q-function learned by the GDQN algorithm will converge to the optimal Q-function with probability 1, thus ensuring that the optimal spectrum allocation strategy is found.
[0296] The convergence and optimality of distributed learning algorithms have been extensively studied in numerous research works. However, proving the convergence and optimality of the DQN algorithm remains a challenging problem. Theorems 3 and 4, based on the convergence condition of the Q-learning algorithm, first prove that the variance of the difference sequence of the DQN algorithm within the EPG framework is bounded. Subsequently, the convergence and asymptotic optimality of the GDQN algorithm are proved using stochastic approximation theory.
[0297] IV. Numerical Results and Discussion
[0298] 1. Scene Setup
[0299] The convergence, advantages, and parameter sensitivity of the GDQN framework in a mega-hybrid satellite constellation were analyzed through simulation. The simulated constellation consists of three GEO satellites and one LEO Walker constellation, with design parameters based on commercial satellites such as Starlink. Detailed parameters are shown in Tables 2 and 3.
[0300] The simulation was conducted on a computer equipped with a 2.4 GHz Intel Core i5-1135G7 CPU and 16 GB of RAM. The simulation lasted 30 minutes, divided into 180 time slots, each slot being 10 seconds long. In this constellation, it is assumed that LEO satellites can mitigate interference caused by spectrum sharing and improve user satisfaction by optimizing beam spectrum access strategies.
[0301] Table 2 Parameters of Mixed Constellations
[0302] Table 3 GDQN Model Parameters
[0303] The experimental setup is based on the following assumptions:
[0304] LU Distribution: LUs are randomly distributed within the beams of GEO satellites. Specifically, GS is located near the equator, and LUs are randomly distributed within a radius of 200 kilometers around GS, such as... Figure 7 As shown.
[0305] Beam service limitation: Each LEO satellite beam can only serve a maximum of one LU at a time.
[0306] Static parameter assumptions: The satellite's orbital position, LU position, and flow requirements remain unchanged within each 10-second time slot.
[0307] The experiment aims to evaluate the performance of the GDQN framework in a mega-hybrid constellation and to verify its effectiveness and stability in spectrum sharing management.
[0308] 2. Convergence
[0309] Convergence is a key metric for evaluating algorithm performance. This experiment quantifies convergence speed by recording the reward value for each iteration. The experimental setup includes 50 LUs randomly distributed across the aforementioned region. Experimental results are as follows: Figure 5 As shown.
[0310] Centralized Control Model Training Scheme (CCMTS): CCMTS exhibits the fastest convergence speed, with the algorithm rapidly reaching high reward values and converging to the optimal reward value in the early stages. However, it incurs a significant information exchange burden and high computational cost, making it less suitable for large and complex networks.
[0311] GDQN: GDQN exhibits a moderate convergence speed, with the reward value steadily increasing throughout the iterations. Although its convergence reward value is slightly lower than CCMTS, its low information transmission and computation costs make it well-suited for dynamic satellite network environments where a balance between convergence speed and performance is required.
[0312] Non-cooperative Model Training Scheme (NCMTS): NCMTS has the slowest convergence speed. Although the algorithm gradually increases its reward value, it converges to the minimum reward value. Since NCMTS operates independently and has no information exchange, its information transmission and computation costs are the lowest, making it suitable for small networks with limited inter-node interaction.
[0313] 3. Advantages
[0314] Traditional spectrum management methods rely on centralized management and optimization, which is inefficient in large, dynamic networks. These methods typically require significant communication and computational overhead, leading to prolonged response times and increased system latency. They are also susceptible to single points of failure. In contrast, the GDQN framework introduces a locally cooperative LIMG mechanism, where each agent interacts only with its neighbors. This approach reduces communication and computational overhead while maintaining flexibility and robustness.
[0315] 1) Balanced use of spectrum
[0316] The spectrum is calculated using the balance method as follows:
[0317] Where N is the number of LUs, yes Actual spectrum resources used yes Total available spectrum resources.
[0318] The spectrum comparison between GDQN and traditional methods uses balanced simulation results as follows: Figure 6 As shown, although CCMTS achieved optimal spectrum usage balance in its early stages, its communication and computational overhead increased dramatically with the number of users, leading to a rapid decline in spectrum usage balance. NCMTS reduced these overheads, but its spectrum usage balance remained low due to a lack of effective information exchange in independent decision-making. In contrast, GDQN not only significantly reduced communication and computational overhead but also maintained a higher spectrum usage balance across networks of different sizes through local information exchange and collaborative decision-making based on deep reinforcement learning. This demonstrates that GDQN possesses a more efficient and flexible spectrum resource allocation capability in large-scale dynamic LEO satellite networks, achieving better optimal resource allocation.
[0319] 2) User perceived satisfaction
[0320] according to Figure 9 The simulation results shown compare the impact of the three algorithms on perceived user satisfaction. As the number of users increases, the average perceived user satisfaction decreases for all three algorithms, which is consistent with the reality that interference becomes more severe with an increased number of users under limited spectrum resources.
[0321] CCMTS initially achieved the highest average user satisfaction, but its network transmission and computation costs increased dramatically with the number of users. NCMTS reduced these costs, but due to a lack of effective information exchange, spectrum conflicts and interference occurred frequently. This independent decision-making resulted in the lowest average user satisfaction. In contrast, GDQN maintained a moderate level of average user satisfaction. Through local information exchange and collaborative decision-making, GDQN effectively reduced co-channel interference and computational costs, and was able to adapt to the dynamic characteristics of LEO satellite network topologies. This demonstrates that GDQN has more efficient and flexible spectrum resource management capabilities, better addressing the challenges of large-scale dynamic networks.
[0322] 3) Network fairness
[0323] Most existing papers focus on overall optimization and often neglect user fairness
[42] . Inspired by
[23] , the Network Experience Quality (QoE) fairness index is introduced to better demonstrate the effectiveness of the proposed algorithm. Considering the current alliance structure, the QoE fairness index is expressed as follows:
[0324]
[0325] Where N is the number of LUs, yes User perceived satisfaction. It can be concluded that... The closer it is to 1, the higher the fairness among users.
[0326] according to Figure 10 The simulation results shown indicate that, similar to average user perceived satisfaction, the CCMTS network initially exhibits the best fairness index. However, as the number of users increases, its communication and computational overhead rises sharply, making it unsuitable for large-scale satellite networks. NCMTS, due to a lack of information exchange and coordination, suffers from severe spectrum conflicts and interference, resulting in unbalanced resource allocation and the lowest network fairness index. In contrast, GDQN maintains a stable level of network fairness across different user scales by combining local information exchange and collaborative decision-making. This makes GDQN suitable for large-scale, highly dynamic LEO satellite networks.
[0327] 4) Parameter sensitivity analysis
[0328] By adjusting the parameters and We analyzed the impact of exploration rate decay on algorithm performance. For example... Figure 11 and Figure 12 As shown, the smaller A higher value helps in finding more local optima in complex environments because it increases the chances of exploration in the early stages. Conversely, a larger value... While this accelerates the convergence of the algorithm, it may prematurely abandon the exploration, causing the opportunity to miss the chance to find a potentially better strategy.
[0329] Larger Maintaining a high exploration rate in the later stages of training helps in discovering new strategies, but may slow down the convergence speed. Smaller values... The value is reduced in the later stages of exploration, thereby improving the stability and final performance of the algorithm.
[0330] By adjusting We analyzed the impact of local altruistic rewards on cooperative and competitive behavior in multi-agent environments. For example... Figure 13 As shown, the smaller Values make agents more competitive and focused on personal gain. This can lead to a decline in overall performance due to a lack of information exchange and cooperation among agents. Conversely, larger values... Values promote cooperation among neighboring agents through information exchange and collaboration, reducing interference from competitors and improving average user satisfaction. However, excessive cooperation may also reduce the learning efficiency of individual agents, as they may not be able to fully explore their optimal strategies.
[0331] Appropriate A value (e.g., 0.5) can strike a balance between cooperation and competition. This balance enhances collaboration and individual learning, ultimately leading to optimal overall performance in a multi-agent environment.
[0332] The GDQN framework is used for dynamic spectrum optimization in mega-hybrid LEO satellite constellations under spectrum sharing. By integrating LIMGM with a deep Q-network, the framework effectively adapts to the dynamic LEO environment, reducing communication and computational overhead while maintaining scalability. Theoretical analysis confirms that the framework converges to a pure policy Nash equilibrium and its asymptotic optimality. Simulation results show that GDQN improves user satisfaction and spectrum efficiency compared to traditional methods, balancing fairness and performance. Key contributions include formulating the problem as an EPG, developing the scalable GDQN framework, and providing rigorous theoretical guarantees.
[0333] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process may be rearranged without departing from the scope of this disclosure. The appended method claims provide elements of various steps in an exemplary order and are not intended to limit the scope to the specific order or hierarchy described.
[0334] In the above detailed description, various features are combined together in a single embodiment to simplify this disclosure. This approach to disclosure should not be construed as reflecting an intention that embodiments of the claimed subject matter require more features than are explicitly stated in each claim. Rather, as reflected in the appended claims, the invention is presented with fewer features than all of the features of the single disclosed embodiment. Therefore, the appended claims are hereby explicitly incorporated into the detailed description, wherein each claim stands alone as a preferred embodiment of the invention.
[0335] The disclosed embodiments have been described above to enable any person skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0336] The foregoing description includes examples of one or more embodiments. It is certainly impossible to describe all possible combinations of components or methods in order to describe the above embodiments, but those skilled in the art will recognize that further combinations and arrangements of the various embodiments are possible. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. Furthermore, the term "comprising" as used in the specification or claims is interpreted in a manner similar to the term "including," as interpreted when used as a conjunction in the claims. Additionally, the use of any term "or" in the specification of the claims is intended to mean "non-exclusive or."
[0337] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A spectrum access method for a large hybrid constellation, characterized in that, include: Step 1: Provide downlinks to ground users within the coverage area using a hybrid constellation consisting of multiple GEO satellites and multiple LEO satellites, where the LEO satellites use a frequency within the frequency set of the GEO satellites; and construct a dynamic allocation problem for the downlink of each ground user based on the satisfaction of the ground users. Step 2: Construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlinks based on the downlinks of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium. Step 3: The Local Interaction Markov Game Model (LIMGM) is mapped to the Local Influence Graph Model (LIMG), and combined with the Deep Q Network (DQN) to obtain a graph-based Deep Q Network. The graph-based Deep Q Network is then trained to obtain the trained graph-based Deep Q Network. Step 4: Using the trained graph-based deep Q-network GDQN, solve the problem of dynamically allocating downlinks for each ground user based on their satisfaction with the ground user, and obtain a dynamic allocation scheme for downlinks for each ground user.
2. The spectrum access method for a large LEO constellation according to claim 1, characterized in that, Step 1 includes step 1.1: Downlinks are provided to terrestrial users within the coverage area via a hybrid constellation consisting of three GEO satellites and a group of LEO satellites configured in a Walker constellation. Each LEO satellite can dynamically adjust the beam direction and frequency of the downlink to serve terrestrial users within the coverage area, wherein: The receiving antenna of the GEO ground station GS scans data packets within the available frequency range and identifies the current downlink frequency of the GEO satellite by decoding the flag bits of the data packets; At the beginning of each time slot t, the LEO satellite senses the downlink frequency of the GEO satellite and determines the frequency of one or more downlinks within the frequency set of the GEO satellite used by the LEO satellite. Each LEO satellite includes multiple beams, with each beam corresponding to a frequency, and each beam is used as a transmitting antenna; Within each time slot t, the downlink frequency of the GEO satellite is switched once, and this frequency switching process follows a Markov model. Each ground user accesses the downlink of the LEO satellite and receives the beam of the downlink through a receiving antenna; The system model is constructed through the above steps.
3. The spectrum access method for a large LEO constellation according to claim 2, characterized in that, Step 1 also includes step 1.2: The gain of each transmit antenna of the LEO satellite is determined based on the maximum gain of the LEO satellite's transmit antenna, the transmitter off-axis angle, the far sidelobe level, and the co-channel interference between two ground users accessing the same downlink. The gain of the ground user's receiving antenna is determined based on the maximum gain of the ground user's receiving antenna and the co-channel interference between two ground users accessing the same downlink. An antenna model is constructed by measuring the gain of the transmitting antenna and the gain of the receiving antenna.
4. The spectrum access method for a large LEO constellation according to claim 3, characterized in that, Step 1 also includes step 1.3: Within time slot t, an interference model for the hybrid constellation to provide downlink access to ground users is determined, the interference model including interference from GEO satellites to ground users. Co-channel interference, LEO satellites to ground users The aggregation interference from LEO satellites and the aggregation interference from LEO satellites to the GEO ground station GS, among which: GEO satellite for ground users Co-channel interference refers to: when and When using the same frequency within time slot t right Co-channel interference It depends on the antenna transmit power of the GEO satellite and the ground users of the GEO satellite. The off-axis angle, This indicates the downlink connection between the ground station GS and the GEO satellite. Indicates ground users Access to the downlink of LEO satellites; LEO satellite for ground users Co-channel interference refers to the interference that occurs when multiple LEO satellites use the same frequency to serve ground users in adjacent areas. This interference occurs between the LEO satellites on the same frequency within time slot t. Aggregation interference It depends on the antenna transmission power of the LEO satellite and the LEO satellite and its ground users. The antenna gain between the LEO satellite's transmitting antenna and the ground user's antenna gain. The gain of the receiving antenna; LEO satellite aggregation interference to GEO ground station GS refers to the aggregation interference caused by the downlink of the LEO satellite to the GEO ground station GS when the LEO satellite and GEO satellite share the same frequency. This aggregation interference occurs within time slot t. It depends on the antenna transmit power of the LEO satellite and the antenna gain between the LEO satellite and the ground station GS, which includes the gain of the LEO satellite's transmit antenna and the gain of the ground station GS's receive antenna.
5. The spectrum access method for a large LEO constellation according to claim 4, characterized in that, Step 1 also includes step 1.4: Within time slot t, ground users are connected via GEO satellite. Co-channel interference, LEO satellites to ground users Aggregation interference, ground users The downlink channel noise of the access, as well as the antenna transmit power of the LEO satellite and ground users. Access downlink Channel gain, calculate ground user The signal-to-noise ratio received by the receiving antenna; Based on ground users The signal-to-noise ratio received by the receiving antenna and the downlink The frequency bandwidth is used to calculate the downlink frequency bandwidth. communication capacity ; Ground users Perceived satisfaction is used as an indicator to balance communication capacity. With traffic demand threshold The gap between them.
6. The spectrum access method for a large LEO constellation according to claim 5, characterized in that, Step 1 also includes step 1.5: ground users The dynamic allocation problem of downlink is described as follows: within time slot t, based on the allocation to ground users... downlink frequency Optimize to include ground users To maximize perceived satisfaction; and at the same time satisfy the set constraints, which are: the aggregated interference of LEO satellites to GEO ground station GS is not less than the maximum interference threshold that does not affect the downlink communication of GEO satellites.
7. The spectrum access method for a large LEO constellation according to claim 6, characterized in that, Step 2 includes: Based on the downlink of ground users, a Locally Interactive Markov Game (LIMGM) model is constructed to address the dynamic allocation problem of downlinks. The LIMGM model consists of seven tuples. ,in: This represents the set of game participants, where game participants refer to ground users. Downlink access to LEO satellites ; Represents the set of state spaces. S represents the Cartesian product of the states of all game participants. n Ground users State space, ground users State space refers to ground users The perceived satisfaction of ground users and their neighbors, as well as the ground station GS and ground users The frequency of their respective downlinks; This represents the set of actions of the game participants, where each action refers to the downlink. The frequency of the downlink The frequencies are selected from the set of available downlink frequencies; This represents the set of neighbors of a game participant. A game participant's neighbors are those who, if the ground user... and ground users The distance between them is no greater than the interference distance threshold. Then determine the ground user Ground users Neighbors; It is the state transition probability function, expressed as Where s and s' represent the current and new states of the downlink frequency for ground users accessing LEO satellites, respectively. This represents the actions of all game participants at time slot t; It is the reward function for the participants in the game; It is a discount factor used to measure the weight of immediate rewards and future rewards. .
8. The spectrum access method for a large LEO constellation according to claim 7, characterized in that, Step 3 includes: The edges in the Local Influence Graph (LIMG) model correspond to the neighbor relationships in the Local Interaction Markov Game (LIMGM) model, where the neighbor relationships represent the interference relationships between game participants. The state space of each player in the game includes the player's state space. Local state information Game participants Interacting neighbor game participants Local state information Composition, game participants The state space is represented as: (31) in, Indicates the participants in the game The set of participants in the neighbor game. Indicates the participants in the game Neighbor game participants, , Indicates the participants in the game Local state information, Indicates the participants in the neighbor game. Local state information; DQN is used to approximate each game participant. Action value function The participants in the game Current state and actions As input, through game participants reward function The output is an action. Expected rewards ; The above steps yield a deep Q-network based on a graph model; During the training of a graph-based deep Q-network, each game participant... use Greedy strategies rely on probability. Choose a random action, with probability Select the action with the highest value based on the current action value function; As training progresses, The loss function is gradually decayed to reduce the exploration frequency until it is lower than the preset loss value, thus obtaining the trained graph model-based deep Q network. The convergence and asymptotic optimality of the deep Q network are proved using stochastic approximation theory.
9. The spectrum access method for a large LEO constellation according to claim 8, characterized in that, Step 4 includes: Each player in the game First, observe your current local state information. and the local state information of its interacting neighboring intelligent agents Based on a pre-trained graph-based deep Q-network GDQN, game participants use - A greedy strategy selects actions that cause the game participants to... Select the action with the highest value based on the current action value function: (34) Game participants Perform the selected action and observe the environmental response, which includes the game participants. Instant rewards and the next state And the game participants obtained through dynamic monitoring Neighbor game participants Actions and rewards; And update the game participants based on the environmental response. Local impact diagram To reflect the latest interactive information; Until the average perceived satisfaction of all game participants is maximized, the downlink of each corresponding ground user is used as a time slot t dynamic allocation scheme.
10. A spectrum access system for a large-scale hybrid constellation, characterized in that, include: The basic model building unit is used to build a system model to provide downlinks to ground users in the coverage area through a hybrid constellation of multiple GEO satellites and multiple LEO satellites. The LEO satellites use a frequency within the frequency set of GEO satellites. The dynamic allocation problem of downlinks for each ground user is constructed based on the satisfaction of ground users. The first model conversion unit is used to construct a Locally Interactive Markov Game (LIMGM) model for the dynamic allocation problem of downlink based on the downlink of ground users. The LIMGM model is an exact potential game (EPG), and the LIMGM model has at least one pure policy Nash equilibrium. The second model conversion unit is used to map the Local Interaction Markov Game Model (LIMGM) to the Local Influence Graph Model (LIMG), and combine it with the Deep Q Network (DQN) to obtain a graph-based Deep Q Network. The graph-based Deep Q Network is then trained to obtain the trained graph-based Deep Q Network. The solution unit is used to solve the problem of dynamic allocation of downlinks for each ground user based on the satisfaction of ground users by using the trained graph-based deep Q-network GDQN, and obtain the dynamic allocation scheme of downlinks for each ground user.