A multi-beam interference avoidance method and system based on multi-agent reinforcement learning
By employing a multi-agent reinforcement learning method, the problem of beam interference in multi-beam systems was solved, achieving system stability and fairness, and improving throughput and communication quality.
Patent Information
- Application Number
- CN202411235093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-09-04
AI Technical Summary
Beam interference in multi-beam systems severely impacts system performance, leading to signal attenuation, increased bit error rate, and decreased communication quality. Existing technologies struggle to effectively suppress inter-beam interference and ensure system stability and fairness.
A multi-agent reinforcement learning-based approach is adopted. This involves establishing an initial system model graph, performing multi-beam interference analysis, constructing an M/G/1 vacation queuing model, context learning, and Markov decision process transformation. Finally, a multi-agent deep deterministic policy gradient algorithm is used to optimize beam scheduling and achieve interference avoidance.
It effectively avoids interference between multiple beams, improves system throughput and communication quality, enhances system stability and adaptability, and achieves more efficient beam scheduling and interference management.
Smart Images

Figure CN119233427B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication technology, specifically relating to a multi-beam interference avoidance method and system based on multi-agent reinforcement learning. Background Technology
[0002] With the ever-increasing global demand for high-speed internet and data transmission, satellite communication technology is also constantly evolving. Multi-beam technology allows each beam to precisely cover a specific small area on the Earth's surface, meeting the communication needs of high-density users while also enabling more precise resource allocation. However, in multi-beam systems, beam interference severely impacts system performance. Inter-beam interference refers to the signal attenuation, increased bit error rate, and degraded communication quality caused by mutual interference between multiple beams present simultaneously. Therefore, effectively suppressing inter-beam interference and improving the stability and reliability of communication systems has become an important research direction. Based on this, this invention designs a multi-beam interference avoidance method and system based on multi-agent reinforcement learning, which reduces inter-beam interference while ensuring fairness and improving system throughput. Summary of the Invention
[0003] To address the aforementioned problems, this invention proposes a multi-beam interference avoidance method and system based on multi-agent reinforcement learning. This invention analyzes the data packet traffic within a cell using a queuing theory model to determine the optimal beam scheduling and allocation scheme. This invention solves the problems existing in the prior art, achieving interference avoidance in current low-Earth orbit satellite multi-beam systems, and improving intra-beam throughput while ensuring fairness between cells.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] A multi-beam interference avoidance method based on multi-agent reinforcement learning includes the following steps:
[0006] S1. Establish an initial system model diagram based on the relationship between the satellite-generated multibeams and ground cells;
[0007] S2. Analyze the interference between multiple beams;
[0008] S3. Generate a service arrival model based on ground status information;
[0009] S4. Construct the M / G / 1 vacation queuing model and input the results of the business arrival model in S3 into the M / G / 1 vacation queuing model;
[0010] S5. Perform contextual learning for multi-beam beams;
[0011] S6. Transform the beam interference avoidance problem into a multi-agent deep deterministic policy gradient (MADDPG) algorithm learning problem, and define the transformation of the Markov decision process (POMDP) problem.
[0012] S7. Based on the transformation definition of the problem and the MADDPG algorithm, perform reinforcement learning of the MADDPG algorithm, and plan the optimal strategy for satellite beam scheduling based on the maximum probability action value of multiple satellite beam outputs.
[0013] As a preferred option, in step S1, an initial system model diagram is established based on the multibeams generated by the satellite and the relationship between ground cells, including:
[0014] The system employs a single-satellite multi-beam model, comprising a low-Earth orbit satellite, a gateway station, and ground cells. The antennas utilize a planar uniform phased array, capable of randomly generating K beams, denoted as K = {k | k = 1, 2, ..., K}, to cover N ground cells, denoted as N = {n | n = 1, 2, ..., N}, where K... <N。
[0015] As a preferred option, step S2 involves performing inter-beam interference analysis, specifically including:
[0016] S21. As can be seen from step S1, when the same frequency beams communicate via a link, the side lobes or main lobes of each point beam overlap spatially with the main lobes of other beams. Therefore, in addition to the required signal, interference signals from the main lobes or side lobes of adjacent beams will be received, which will affect the signal quality and data transmission rate of the receiver.
[0017] S22. Based on S21, the carrier-to-interference ratio (C / I) of the terminal received signal is used to quantify inter-beam interference:
[0018]
[0019] Among them, G t (θ0) is the desired beam transmit antenna gain, G t,I (θ i () is the antenna gain for transmission to other interfering beams.
[0020] As a preferred option, step S3 involves generating a service arrival model based on ground status information, specifically including:
[0021] A service arrival model was established based on the geographical and topographical differences of ground cells. Service volume is calculated by weighting three factors: terrain, development, and time. Geographical terrain is categorized into four types: ocean, land, desert, and mountain, with different coefficients assigned to each. Development status is categorized into developing and developed. The time factor is assigned a value considering the 24-hour work-rest pattern of human life. The calculation considers the weighted relationship of the three factors. At time t, the device deployment density of grid i is defined as:
[0022]
[0023] Among them, S i ρ represents the area of raster i, m represents the number of geographic environment types contained within the raster, and ρ represents the area of raster i. j S represents the device deployment density corresponding to geographical environment type j, calculated by weighting based on terrain and development conditions. i,j This represents the area occupied by geographic environment type j within grid i. Traffic flow within the system follows a Poisson distribution with arrival rate λ.
[0024] As a preferred option, in step S4, an M / G / 1 vacation queuing model is constructed, and the results output from the business arrival model in S3 are input into the queuing model. Specifically, this includes:
[0025] S41. The satellite can provide N queues to store the traffic arriving at each cell. The arrival traffic in time slot t is represented as:
[0026] L t ={l t,n |n∈N}
[0027] Among them, l t,n Let n be the arrival flow rate of cell n in time slot t, which follows the arrival rate λ. t,n The Poisson distribution. The total flow stored in the t-slot queue is represented as D. t ={d t,n |n∈N}, where d t,n Let t be the total traffic stored in queue n in time slot t.
[0028] S42. Each cell is relatively independent. Within a cell, data packets arrive via a Poisson process with an average arrival rate of λ. Upon arrival and entry into the cell, packets queue and wait for processing by the beam, following a first-come, first-served principle. When the beam is active, the service time for each data packet is a random variable, following a general distribution with an average service rate of μ. on The mean of the service time is E[S] and the variance is σ. S 2 Furthermore, the service rate is 0 when the beam is off, and data packets still arrive at the cell and are queued even when the beam is off.
[0029] S43. Based on step S42, the data packet queuing and waiting service process is designed as an M / G / 1 model, with beam on being regarded as the working state and beam off being the vacation state.
[0030] S44. Based on step S43, when the beam is turned on, the average waiting time (including queuing time and service time) can be given by the Pollaczek-Khinchine (PK) formula:
[0031]
[0032] Among them, E[S 2 [ ] is the second moment of the service time. It refers to the system utilization rate.
[0033] During beamout, assume the average beamout time is T. off During the shutdown period, the average cumulative latency of each data packet is:
[0034] τ off =T off
[0035] S45. Based on step S44, let the probability of beam activation be p. on The probability of closing is p off p on and p off This can be obtained through statistical analysis of historical data from the cell. Taking into account both the on and off states of the beam, the average waiting time can be expressed as:
[0036] τ t,n =p on ·τ on +p off ·τ off
[0037] As a preferred option, in step S4, beam activation is considered the working state, and beam deactivation is the vacation state; the information data obtained from analyzing the queuing model is the average waiting time within the cell.
[0038] As a preferred option, step S5, the context learning of the multi-beams, specifically includes:
[0039] S51. Use the operating states of all beams to form a context for interference avoidance; when a beam serves a cell, it is said to be in an active state; otherwise, it is said to be in an idle state; the context vector of the operating states of all beams at observation time t is collected and defined as follows:
[0040] c(t) = {c1(t), c2(t), ..., c K(t)}
[0041] S52. The agent observes the state of beams that interfere with its own transmission and applies the resulting context. Each agent masks the state of non-interfering beams to form its own version of the context. Based on the underlying deployment layout, the non-interfering beams for each beam can be easily identified, and corresponding mask settings can be applied. Each machine learning (ML) agent masks the state of non-interfering beams in the context to form its own version of the context; the mask vector representing the non-interfering beams of beam bi can be defined as follows:
[0042] M i ={m 1|i ,m 2|i ,…,m K|i}
[0043] Where, m k|i This indicates whether the activation of beam k will not interfere with beam i, that is, when beam k can interfere with beam i, m k|i It becomes 1 if it is set to 1, otherwise it becomes 0.
[0044] S53, c i (t) represents the observation context of beam i after all non-interfering beam states are masked to 0:
[0045]
[0046] Where c(t)={c1(t),c2(t),…,c K (t)} represents all the context observed by the agent initially.
[0047] As a preferred approach, in step S6, the beam interference avoidance problem is transformed into a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm learning problem, and a Markov Decision Process (POMDP) problem is defined, specifically including:
[0048] S61. Define the global state s, including the average waiting time τ of the ground cell. i Data packets q generated by the beam transmission process of the sanitary system b And the data traffic D within the community:
[0049] s={τ i ,q b ,D}
[0050] S62. Define the context c of the t-slot beam observation and the basic observation feature s as the global state o:
[0051]
[0052] o = {s, c}
[0053] S63. The agent should make decisions based on the state to improve long-term gains. The agent dynamically adjusts the beam coverage of each time slot unit, i.e., turns the beam on or off, intelligently adjusting the selection based on the state in each time slot. Therefore, the action performed at time t is defined as:
[0054] a = {x1, x1, ..., x1} k |x i ∈{0,1}}
[0055] S64. To maximize the defined long-term optimization objective, data throughput and latency fairness are used as immediate rewards:
[0056] r = αTP total -(1-α)TD
[0057] Among them, TP total TD represents the latency difference.
[0058] As a preferred option, in step S7, based on the transformed definition of the problem and the MADDPG algorithm, reinforcement learning of the MADDPG algorithm is performed, and the optimal strategy for satellite beam scheduling is planned based on the maximum probability action value of multiple satellite beam outputs, including:
[0059] S71. Randomly initialize the actor network and critic network for all beams and configure the experience playback buffer.
[0060] S72. Utilize a deterministic policy network to collect observation data of the current environment. This data includes the beam context c of time slot t and basic observation features s, where s includes the average latency of the ground cell, the number of packets processed by the current beam, the number of packets dropped within the cell, and the data traffic within the cell.
[0061] o = {s, c}
[0062] S73. Based on the environmental observation data obtained in step S72, determine the target beam generation action for the low-orbit satellite and execute the action.
[0063] S74. After completing the beam decision action, record the obtained reward and the environmental observation data of the next state, and store the current environmental observation data, the target satellite's execution action, the obtained reward and the new environmental observation data in the experience playback buffer.
[0064] S75. Randomly sample a batch of experiences from the experience replay buffer for training, update the parameters of the actor network and the critic network, aiming to maximize the Q-value of the critic network, minimize the error of the Q-value, and use the updated Q-network to generate the optimal policy path.
[0065] S76. Repeat steps S72-S75 until the algorithm converges (e.g., the beam scheduling strategy tends to stabilize during training and no longer changes significantly), thereby determining the optimal beam scheduling strategy.
[0066] This invention also discloses a multi-beam interference avoidance system based on multi-agent reinforcement learning, used to execute the above-described method, comprising the following modules:
[0067] Initial Multibeam System Model Diagram Construction Module: Establishes an initial system model diagram based on the multibeams generated by the satellite and the relationship between ground cells;
[0068] Information and data acquisition module: Based on the characteristics of multi-beams, analyze the interference between multi-beams and quantify the interference;
[0069] Service arrival model module: Generates service arrival models based on ground status information;
[0070] Queuing Model Module: Constructs M / G / 1 queuing models, inputs service arrival information into the queuing model, and generates an independent queuing model for each cell;
[0071] Context learning module: Multi-beam context learning, generating a unique context for each beam based on the underlying beam information;
[0072] Transformation Module: Transforms the beam interference avoidance problem into a multi-agent deep deterministic policy gradient algorithm learning problem and performs a transformation definition of the Markov decision process problem;
[0073] Reinforcement Learning and Beam Scheduling Module: Based on the transformation definition of the problem and the multi-agent deep deterministic policy gradient algorithm, reinforcement learning of the multi-agent deep deterministic policy gradient algorithm is performed, and the optimal strategy for satellite beam scheduling is planned based on the maximum probability action value of the output of multiple satellite beams.
[0074] Compared with the prior art, the present invention has the following technical effects:
[0075] (1) The present invention adopts a context learning model, in which the target beam learns its own context environment, masks non-interfering beams through the underlying beam information, and finally generates specific context state information for each target beam. This model can effectively capture and utilize the context information of the environment while saving state space, thereby optimizing the beam scheduling strategy and improving the overall performance and adaptability of the system.
[0076] (2) This invention selects the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm in reinforcement learning. By simultaneously inputting the observed environment state information and context information into a partial Markov decision formula, it can better adapt to changes and complexities in the environment. At the same time, this algorithm can handle complex decision problems in multi-agent environments, improve the overall system performance by coordinating the behavior of multiple agents, and thus achieve more efficient beam scheduling and interference management. Attached Figure Description
[0077] Figure 1 This is a flowchart of a preferred embodiment of the present invention, which describes a multi-beam interference avoidance method based on multi-agent reinforcement learning.
[0078] Figure 2 This is a model diagram based on context learning according to a preferred embodiment of the present invention;
[0079] Figure 3 This is a model diagram of a preferred embodiment of the present invention based on a multi-agent deep deterministic policy gradient algorithm for learning.
[0080] Figure 4 This is a block diagram of a multi-beam interference avoidance system based on multi-agent reinforcement learning, which is a preferred embodiment of the present invention. Detailed Implementation
[0081] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0082] like Figure 1As shown, this embodiment proposes a multi-beam interference avoidance method based on multi-agent reinforcement learning. This method combines a multi-agent deep deterministic policy gradient algorithm model and a context learning model, while also using a queuing model. First, an initial system model diagram is established based on the relationship between the satellite-generated multi-beams and ground cells. Interference analysis between multi-beams is performed, and a service arrival model is generated by combining ground state information. Simultaneously, the model output is input into the queuing model to construct an M / G / 1 queuing model. The beam interference avoidance problem is transformed into a multi-agent deep deterministic policy gradient (MADDPG) algorithm learning problem. A Markov decision process (POMDP) problem is defined, and context learning for the multi-beams is performed. Based on the transformed definition of the problem and the MADDPG algorithm, reinforcement learning of the MADDPG algorithm is performed. The optimal strategy for satellite beam scheduling is planned based on the maximum probability action value output by multiple satellite beams. The specific steps of this embodiment are as follows:
[0083] S1. Establish an initial system model diagram based on the relationship between the satellite-generated multibeams and ground cells;
[0084] S2. Analyze the interference between multiple beams;
[0085] S3. Generate a service arrival model based on ground status information;
[0086] S4. Construct the M / G / 1 vacation queuing model and input the results of the business arrival model in S3 into the queuing model;
[0087] S5. Perform contextual learning for multi-beam beams;
[0088] S6. Transform the beam interference avoidance problem into a multi-agent deep deterministic policy gradient (MADDPG) algorithm learning problem, and define the transformation of the Markov decision process (POMDP) problem.
[0089] S7. Based on the transformation definition of the problem and the MADDPG algorithm, perform reinforcement learning of the MADDPG algorithm, and plan the optimal strategy for satellite beam scheduling based on the maximum probability action value of multiple satellite beam outputs.
[0090] The steps of this embodiment will be described in detail below.
[0091] In step S1, an initial system model diagram is established based on the multibeams generated by the satellite and the relationship between ground cells, including:
[0092] The system employs a single-satellite multi-beam model, comprising a low-Earth orbit satellite, a gateway station, and ground cells. The antennas utilize a planar uniform phased array, capable of randomly generating K beams, denoted as K = {k | k = 1, 2, ..., K}, to cover N ground cells, denoted as N = {n | n = 1, 2, ..., N}, where K... <N。
[0093] Step S2 involves performing inter-beam interference analysis, specifically including:
[0094] S21. As can be seen from step S1, when the same frequency beams communicate via a link, the side lobes or main lobes of each point beam overlap spatially with the main lobes of other beams. Therefore, in addition to the required signal, interference signals from the main lobes or side lobes of adjacent beams will be received, which will affect the signal quality and data transmission rate of the receiver.
[0095] S22. Based on S21, the carrier-to-interference ratio (C / I) of the terminal received signal is used to quantify inter-beam interference:
[0096]
[0097] Among them, G t (θ0) is the desired beam transmit antenna gain, G t,I (θ i () is the antenna gain for transmission to other interfering beams.
[0098] In step S3, a service arrival model is generated based on ground status information, specifically including:
[0099] A service arrival model was established based on the geographical and topographical differences of the ground cells. Service volume was calculated by weighting three factors: terrain, development, and time. At time t, the device deployment density of grid i was defined as:
[0100]
[0101] Among them, S i ρ represents the area of raster i, m represents the number of geographic environment types contained within the raster, and ρ represents the area of raster i. j S represents the device deployment density corresponding to geographical environment type j, calculated by weighting based on terrain and development conditions. i,j This represents the area occupied by geographic environment type j within grid i. Traffic flow within the system follows a Poisson distribution with arrival rate λ.
[0102] In step S4, the M / G / 1 vacation queuing model is constructed by inputting the results of the business arrival model in S3 into the queuing model. Specifically, this includes:
[0103] S41. The satellite can provide N queues to store the traffic arriving at each cell. The arrival traffic in time slot t is represented as:
[0104] L t ={l t,n |n∈N}
[0105] Among them, l t,n Let n be the arrival flow rate of cell n in time slot t, which follows the arrival rate λ. t,n The Poisson distribution. The total flow stored in the t-slot queue is represented as D. t ={d t,n |n∈N}, where d t,n Let t be the total traffic stored in queue n in time slot t.
[0106] S42. Each cell is relatively independent. Within a cell, data packets arrive via a Poisson process with an average arrival rate of λ. Upon arrival and entry into the cell, packets queue and wait for processing by the beam, following a first-come, first-served principle. When the beam is active, the service time for each data packet is a random variable, following a general distribution with an average service rate of μ. on The mean of the service time is E[S] and the variance is σ. S 2 Furthermore, the service rate is 0 when the beam is off, and data packets still arrive at the cell and are queued even when the beam is off.
[0107] S43. Based on step S42, the data packet queuing and waiting service process is designed as an M / G / 1 model, with beam on being regarded as the working state and beam off being the vacation state.
[0108] S44. Based on step S43, when the beam is turned on, the average waiting time (including queuing time and service time) can be given by the Pollaczek-Khinchine (PK) formula:
[0109]
[0110] Among them, E[S 2 [ ] is the second moment of the service time. It refers to the system utilization rate.
[0111] During beamout, assume the average beamout time is T. off During the shutdown period, the average cumulative latency of each data packet is:
[0112] τ off =T off
[0113] S45. Based on step S44, let the probability of beam activation be p. on The probability of closing is p off p on and poff This can be obtained through statistical analysis of historical data from the cell. Taking into account both the on and off states of the beam, the average waiting time can be expressed as:
[0114] τ t,n =p on ·τ on +p off ·τ off
[0115] This embodiment presents a multi-beam interference avoidance method based on multi-agent reinforcement learning, using a context learning model, the model diagram of which is shown below. Figure 2 As shown, by using this model to learn and analyze the information of each beam, interference beams can be identified in advance, reducing computational complexity. The specific steps include the following:
[0116] S51. Use the operating status of all beams to form a context for interference avoidance. Specifically:
[0117] S511. A beam can be in an active or idle state. When a beam is serving a cell, it is said to be in an active state; otherwise, it is said to be in an idle state.
[0118] S512. Assign a separate agent to each beam, enabling it to independently observe the context and interact with the environment. Collect the context vector of the operating state of all beams at observation time t, defined as follows:
[0119] c(t) = {c1(t), c2(t), ..., c K (t)}
[0120] S52. Each ML agent masks the state of the non-interfering beam in the context to form its own version of the context.
[0121] Based on the underlying deployment layout, the non-interfering beams of each beam can be easily identified, and corresponding mask settings can be applied. The mask vector representing the non-interfering beam of beam i can be defined as follows:
[0122] M i ={m 1|i ,m 2|i ,…,m K|i}
[0123] Where, m k|i This indicates whether the activation of beam k will not interfere with beam i, that is, when beam k can interfere with beam i, m k|i It becomes 1 if it is set to 1, otherwise it becomes 0.
[0124] S53. Based on all beam state information obtained in step S51 and the mask information based on non-interference beams obtained in S52, the context information of each agent is obtained. (The last part, "c," appears to be a typo and can be left as is.) i (t) represents the observation context of beam i after all non-interfering beam states are masked to 0. i (t) can be represented as follows.
[0125]
[0126] This embodiment presents a multi-beam interference avoidance method based on multi-agent reinforcement learning, which is based on a multi-agent deep deterministic policy gradient algorithm model. The model diagram is shown below. Figure 3 As shown, in this network, during the training phase, the Q-network receives environmental information composed of observations from all beams generated by the satellite, state information generated through context learning, and actions of all agents as input. It then centrally calculates the action-value function for each beam. Each beam learns an individual Q-value as feedback for its actions. The Q-network is trained based on the estimated and actual Q-values, and the satellite updates its policy based on the feedback from the Q-network. Once each beam is sufficiently trained, it can take appropriate actions based on its state, without needing the states or actions of other beams. During the execution phase, the beam's local observations are used as input to output its action. Specifically, the steps include:
[0127] S61. Define the environmental basic observation characteristics s, including the average waiting time τ of the ground cell. i The data packets q generated by the satellite are transmitted and processed by the beam. b And the data traffic D within the community:
[0128] s={τ i ,q b ,D}
[0129] S62. Define the context c of the t-slot beam observation and the basic observation feature s as the global state o:
[0130]
[0131] o = {s, c}
[0132] S63. Define the action space as beam on or off:
[0133] a = {x1, x2, ..., x} k |x i ∈{0,1}}
[0134] Where, x i =0 indicates beam off, x i =1 indicates that the beam is turned on.
[0135] S64. Define the reward function based on the characteristics of traffic and waiting latency within the cell:
[0136] r = αTP total -(1-α)TD
[0137] Among them, TP total TD represents the latency difference.
[0138] Step S7 is as follows:
[0139] S71. Randomly initialize the actor network and critic network for all beams and configure the experience playback buffer.
[0140] S72. Utilize a deterministic policy network to collect observation data of the current environment. This data includes the beam context c of time slot t and basic observation features s, where s includes the average latency of the ground cell, the number of packets processed by the current beam, the number of packets dropped within the cell, and the data traffic within the cell.
[0141] S73. Based on the environmental observation data obtained in step S72, determine the target beam generation action for the low-orbit satellite and execute the action.
[0142] S74. After completing the beam decision action, record the obtained reward and the environmental observation data of the next state, and store the current environmental observation data, the target satellite's execution action, the obtained reward and the new environmental observation data in the experience playback buffer.
[0143] S75. Randomly sample a batch of experiences from the experience replay buffer for training, update the parameters of the actor network and the critic network, aiming to maximize the Q-value of the critic network, minimize the error of the Q-value, and use the updated Q-network to generate the optimal policy path.
[0144] S76. Repeat steps S71-S75 until the algorithm converges, thereby determining the optimal beam scheduling strategy.
[0145] This invention solves the problem of interference between multiple beams in existing low-orbit satellite multi-beam systems, making it more reasonable and flexible, and effectively avoiding interference, thereby improving the overall system performance.
[0146] like Figure 4 As shown in the figure, this embodiment discloses a multi-beam interference avoidance system based on multi-agent reinforcement learning, used to execute the method of the above embodiment, which includes the following modules:
[0147] Initial Multibeam System Model Diagram Construction Module: Establishes an initial system model diagram based on the multibeams generated by the satellite and the relationship between ground cells;
[0148] Information and data acquisition module: Based on the multi-beam characteristics of the system, analyze the interference between the multi-beams and quantify the interference;
[0149] Service arrival model module: Generates service arrival models based on ground status information;
[0150] Queuing Model Module: Constructs M / G / 1 vacation queuing models, inputs the results of the business arrival model into the queuing model, and generates an independent queuing model for each cell;
[0151] Context learning module: Multi-beam context learning, generating a unique context for each beam based on the underlying beam information;
[0152] Transformation Module: Transforms the beam interference avoidance problem into a multi-agent deep deterministic policy gradient algorithm learning problem and performs a transformation definition of the Markov decision process problem;
[0153] Reinforcement Learning and Beam Scheduling Module: Based on the transformation definition of the problem and the multi-agent deep deterministic policy gradient algorithm, reinforcement learning of the multi-agent deep deterministic policy gradient algorithm is performed, and the optimal strategy for satellite beam scheduling is planned based on the maximum probability action value of the output of multiple satellite beams.
[0154] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.
Claims
1. A multi-beam interference avoidance method based on multi-agent reinforcement learning, characterized by: Includes the following steps: S1. Establish an initial system model diagram based on the relationship between the satellite-generated multibeams and ground cells; S2. Analyze the interference between multiple beams; S3. Generate a service arrival model based on ground status information; S4. Construct the M / G / 1 vacation queuing model and input the results of the business arrival model in S3 into the M / G / 1 vacation queuing model; S5. Perform contextual learning for multi-beam beams; S6. Transform the beam interference avoidance problem into a multi-agent deep deterministic policy gradient algorithm learning problem, and define the transformation into a Markov decision process problem. S7. Based on the transformation definition of the problem and the multi-agent deep deterministic policy gradient algorithm, perform reinforcement learning of the multi-agent deep deterministic policy gradient algorithm, and plan the optimal strategy for satellite beam scheduling based on the maximum probability action value output by multiple satellite beams.
2. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 1, characterized in that, In step S1, a single-satellite multi-beam model is adopted, including low-Earth orbit satellites, gateway stations, and ground cells; the antenna selection is set to a planar uniform phased array antenna array, which can be randomly generated. K A beam, denoted as K={k│k=1,2,...,K}, covers the ground. N There are N cells, denoted as N={n│n=1,2,…,N}, where K < N .
3. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 1, characterized in that, In step S2, the carrier-to-interference ratio (C / I) of the terminal received signal is used to quantize the inter-beam interference: Among them, G t (θ0) is the desired beam transmit antenna gain, G t,I (θ i () is the antenna gain for transmission to other interfering beams.
4. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 1, characterized in that, In step S3, considering the geographical and topographical differences of the ground cells, a service arrival model is established. This model comprehensively reflects the service volume under different conditions by weighting the calculations based on terrain features, development status, and time factors; at time... t Define grid i The equipment deployment density is: Among them, S i Represents grid i area, m ρ represents the number of geographic environment types contained within a raster. j Indicates geographical environment type j The corresponding device deployment density, S i,j Represents grid i Inland geographical environment types j The area occupied; the flow rate follows a Poisson distribution with an arrival rate of λ.
5. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 4, characterized in that, Step S4 involves constructing the M / G / 1 vacation queuing model, inputting the results from the business arrival model in S3 into the queuing model, specifically including: S41, satellite provided N One queue is used to store the traffic arriving at each cell; time slots t The arrival flow is expressed as: L t ={l t,n |n∈N} in, l t,n Let n be the arrival flow rate of cell n in time slot t, which follows the arrival rate. λ t,n The Poisson distribution; the total flow stored in the t-slot queue is represented by D. t ={d t,n |n∈N}, where d t,n Let n be the total throughput stored in queue n during time slot t. S42. Each cell is relatively independent. Within a cell, data packets arrive via a Poisson process with an average arrival rate of λ. Upon arrival and entry into the cell, packets queue and wait for processing by the beam, following a first-come, first-served principle. When the beam is active, the service time for each data packet is a random variable, following a general distribution with an average service rate of μ. on The mean of the service time is E[S] and the variance is... When the beam is off, the service rate is 0, but data packets still arrive at the cell and are queued even when the beam is off. S43. Design the data packet queuing service process as an M / G / 1 vacation queuing model, and regard the beam opening as the working state and the beam closing as the vacation state. S44. When the beam is on, the average waiting time is given by the PK formula: Among them, E[S 2 ] is the second moment of the service time. It is the system utilization rate; During beam shutdown, let the average duration of beam shutdown be... T off During the shutdown period, the average cumulative latency of each data packet is: S45. Let the probability of beam activation be... p on The probability of closing is p off , p on and p off The average waiting time is obtained through statistical analysis of historical data from the cell; taking into account both the on and off states of the beam, the average waiting time is expressed as: 。 6. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 5, characterized in that, Step S5 specifically includes: S51. Use the operating states of all beams to form a context for interference avoidance; when a beam serves a cell, it is said to be in an active state; otherwise, it is said to be in an idle state; the context vector of the operating states of all beams at observation time t is collected and defined as follows: c (t)={ c 1(t), c 2(t),…, c K (t) } in, c K (t) To collect observation time t Time Beam K Context vector information; S52. Each machine learning agent masks the state of the non-interfering beam in the context to form its own version of the context; representing the beam. i The mask vector for the non-interference beam is defined as follows: Mi={m 1│i, m 2│i,…, m K│i } Where, m k│i This indicates whether the activation of beam k will not interfere with the beam. i When the beam k Capable of interfering with beams i m k│i Change to 1 if it is set to 1, otherwise change to 0. S53. Based on the operating status information of all beams obtained in step S51 and the mask vector information based on non-interference beams obtained in step S52, the context information of each agent is obtained.
7. The multi-beam interference avoidance method based on multi-agent reinforcement learning according to claim 6, characterized in that, Step S6 specifically includes: S61. Define environmental basic observation characteristics s, including average latency of ground cells and data packets processed by satellite-generated beam transmission. q b And the data traffic D within the community: S62. Define the context c of the t-slot beam observation and the basic observation feature s as the global state o: S63. Define the action space as beam on or off: When x i =0 indicates beam off, x i =1 indicates that the beam is turned on; S64. Define the reward function based on the characteristics of traffic and waiting latency within the cell: in, These are the weighting coefficients. TP total Indicates throughput. TD This indicates the time delay difference.
8. A multi-beam interference avoidance method based on multi-agent reinforcement learning according to any one of claims 1-7, characterized in that, Step S7 specifically includes: S71. Randomly initialize the actor network and critic network for all beams and configure the experience playback buffer; S72. Utilize a deterministic policy network to collect observational data of the current environment, including time slots. t beam context c and basic observational features s ,in, s This includes the average waiting time of the ground cell, the number of data packets currently being processed by the beam, the number of data packets dropped within the cell, and the data traffic within the cell; S73. Based on the environmental observation data obtained in step S72, determine the target beam generation action for the low-orbit satellite and execute the action; S74. After executing the beam decision action, record the obtained reward and the environmental observation data of the next state, and store the current environmental observation data, the target satellite's execution action, the obtained reward and the new environmental observation data in the experience playback buffer; S75. Randomly sample a batch of experiences from the experience replay buffer for training, updating the parameters of the actor network and the critic network, aiming to maximize the critic network's performance. Q Value, minimize Q The error of the value, and utilize the updated Q The network generates the optimal strategy path; S76. Repeat steps S71-S75 until convergence is achieved, thereby determining the optimal beam scheduling strategy.
9. A multi-beam interference avoidance system based on multi-agent reinforcement learning, used to perform the method as described in any one of claims 1-8, characterized in that, Includes the following modules: Initial Multibeam System Model Diagram Construction Module: Establishes an initial system model diagram based on the multibeams generated by the satellite and the relationship between ground cells; Information and data acquisition module: Based on the characteristics of multi-beams, analyze the interference between multi-beams and quantify the interference; Service arrival model module: Generates service arrival models based on ground status information; Queuing Model Module: Constructs the M / G / 1 holiday queuing model, inputs the results of the service arrival model into the M / G / 1 holiday queuing model, and generates an independent queuing model for each cell; Context learning module: Multi-beam context learning, generating a unique context for each beam based on the underlying beam information; Transformation Module: Transforms the beam interference avoidance problem into a multi-agent deep deterministic policy gradient algorithm learning problem and performs a transformation definition of the Markov decision process problem; Reinforcement Learning and Beam Scheduling Module: Based on the transformation definition of the problem and the multi-agent deep deterministic policy gradient algorithm, reinforcement learning of the multi-agent deep deterministic policy gradient algorithm is performed, and the optimal strategy for satellite beam scheduling is planned based on the maximum probability action value of the output of multiple satellite beams.
Citation Information
Patent Citations
Dynamic resource allocation method for beam hopping satellite system based on deep reinforcement learning
CN114499629A
Flow prediction satellite path selection method and system based on reinforcement learning
CN116781139A