A multi-satellite hop beam scheduling method for random access

By combining a multi-agent deep deterministic policy gradient and conditional generation diffusion model with a primal-dual optimization multi-satellite hop beam scheduling method, the problems of low user access efficiency and QoS constraints in LEO satellite constellations are solved, achieving efficient user access and resource utilization.

CN121462066BActive Publication Date: 2026-05-22BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Traditional fixed-beam or static polling schemes are difficult to handle uplink communication scenarios in LEO satellite constellations where user access behavior is highly random and resource requirements change dynamically, resulting in low access efficiency and poor resource utilization. Furthermore, existing deep reinforcement learning methods lack effective modeling of security and QoS constraints.

Method used

A multi-agent deep deterministic policy gradient (MADDPG) framework is adopted, combined with conditional generative diffusion model (GDM) and primal-dual optimization method, to construct a multi-star hop beam scheduling method. User access is optimized through dynamic beam coverage strategy to meet strict QoS constraints and security requirements.

Benefits of technology

It significantly improves user access success rate, reduces service latency, enhances resource utilization efficiency and system reliability, and has good dynamic adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462066B_ABST
    Figure CN121462066B_ABST
Patent Text Reader

Abstract

The application provides a multi-satellite hopping beam scheduling method for random access, and the core lies in proposing a constraint reinforcement learning framework based on a security diffusion model, wherein a beam is modeled as an intelligent agent, a multi-agent deterministic policy gradient MADDPG algorithm is used to realize multi-satellite hopping beam collaborative scheduling, and a conditional generation diffusion GDM model is introduced as an intelligent agent strategy network, so that a high-quality, diversified and robust beam coverage scheme can be generated in a complex state space. A centralized training and distributed execution mechanism is used, a Lagrange multiplier is introduced through a primal-dual optimization PDO method to adjust the constraint cost, and a cell waiting delay is dynamically controlled, so that a fixed cell polling time delay constraint can be ensured. The application is significantly superior in improving the number of successful access users and reducing the service waiting time delay, and can more stably meet the service quality requirements under a dynamic load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to a multi-star hop beam scheduling method for random access. Background Technology

[0002] With the increasing demand for high-speed internet access globally, low-Earth orbit (LEO) satellite constellations, with their low latency and global coverage, are gradually becoming a crucial infrastructure for next-generation communication networks. In LEO satellite constellations, random access protocols are essential, enabling user terminals to initiate uplink communication requests autonomously without pre-allocated resources. However, traditional fixed-beam coverage or static polling schemes struggle to meet the efficiency and flexibility requirements of uplink random access in dynamic LEO environments, leading to degraded access performance and low resource utilization. To improve the flexibility of system scheduling and uplink access efficiency, beam hopping (BH) technology for random access has become a research hotspot, enabling rapid response to hotspot areas or sudden demands by dynamically adjusting beam pointing strategies.

[0003] While beam-carrier (BH) technology has been extensively studied in downlink scenarios, including greedy hybrid precoding methods, time slot allocation mechanisms based on genetic algorithms, and adaptive beam scheduling strategies incorporating deep reinforcement learning (DRL), most of these methods rely on the premise that the system is fully observable and the scheduling objective is clearly defined. In the uplink direction, existing research focuses primarily on planned access and centralized control scenarios, proposing beam-carrier allocation algorithms for Quality of Service (QoS) assurance and service priority scheduling strategies based on improved cuckoo search algorithms. However, these generally depend on centralized control and global information awareness, making them difficult to directly apply to random access scenarios where user-initiated behavior is highly random and information is partially considerable.

[0004] Furthermore, most BH methods remain limited to single-satellite systems, neglecting the cooperative interference and scheduling issues among multiple satellites in LEO constellations. More critically, uplink BH scheduling must meet strict QoS constraints, especially the maximum latency limit determined by the inherent polling mechanism of uplink access. Existing DRL methods mostly optimize average performance, lacking effective modeling and guarantees for security constraints. Although Safe Reinforcement Learning (Safe RL) has shown promising prospects in safety-critical scenarios such as unmanned systems, research on its application in combining QoS constraints with uplink beam hopping scheduling in LEO remains lacking. There is an urgent need for a more expressive and constraint-handling intelligent decision-making framework that can simultaneously optimize system performance and ensure operational safety and reliability in practical satellite communications. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a multi-star hop beam scheduling method for random access, which can significantly improve the user access success rate and effectively reduce service waiting latency.

[0006] In a first aspect, the present invention provides a multi-satellite beam hopping scheduling method for random access, comprising:

[0007] Obtain a pre-established low-Earth orbit satellite uplink random access communication system model; the system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multiple time slots. The system model also constructs an interactive environment for beam hopping scheduling based on beam hopping time schedules and user random access behavior rules.

[0008] The global state of the system model is determined in each discrete time slot, and the local observation data of the agent for the ground user cell is determined by taking the beam of the low-orbit satellite as an agent in each discrete time slot. The global state includes the number of active users and the cumulative service waiting time of all ground cells. The local observation data includes the number of active users and the cumulative service waiting time of each cell that can be served by each beam.

[0009] In each discrete time slot, a beam coverage strategy is generated and executed for the agent to complete the interaction with the environment. Each agent generates a dynamic beam coverage strategy based on local state information using a policy network constructed with a Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage strategy, and the environment updates its state according to the user access results and latency changes under the CRDSA protocol, and feeds back reward signals and constraint cost signals to the agent. The local state information, dynamic beam coverage strategy, reward signal, constraint cost signal, and the next local state information together constitute the training transition tuple.

[0010] The system optimizes the policy network through centralized training based on training transformation tuples, thereby improving the scheduling performance of multi-satellite multi-beam networks. Specifically, the system adopts the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework, constructs a reward and cost dual evaluation network during the centralized training phase, and samples transformation tuples from the experience pool for policy updates. At the same time, the primal-dual optimization method is introduced to adjust the policy objective through Lagrange multipliers to meet the time delay constraint.

[0011] In one implementation, obtaining a pre-established low-Earth orbit satellite uplink random access communication system model further includes:

[0012] Beam hopping time scheduling is based on TDM to dynamically schedule beam resources within discrete time slots in order to achieve periodic coverage of the target cell;

[0013] The user random access behavior adopts the contention-based access protocol CRDSA, which enables users to send replica packets through randomly selected access sub-slots during the beam dwell period, and improves decoding efficiency through serial interference cancellation.

[0014] In one implementation, a beam coverage strategy for the agent is generated and executed in each discrete time slot to complete environmental interaction, including:

[0015] The policy network constructed by the conditional generation diffusion model (GDM) generates a dynamic beam coverage strategy for each discrete time slot based on the agent's local observation data. The dynamic beam coverage strategy is a binary decision by which the agent performs beam coverage on the ground user cell it covers, and is used to indicate the position of each beam in each time slot.

[0016] The intelligent agent completes the interaction with the low-Earth orbit satellite communication system environment by executing the beam coverage strategy. The environment updates its status according to the user access results and latency changes under the CRDSA protocol.

[0017] The total number of users who successfully accessed the network in each time slot is determined by the statistical results of users who successfully completed uplink access through the random access protocol in the cells covered by the beam within each time slot.

[0018] Based on the service waiting time of users who failed to successfully access the service in each time slot, the total service waiting time of all users in each time slot is determined.

[0019] Calculate the global reward for all agents in each time slot based on the total number of successfully connected users and the total service waiting latency in each time slot.

[0020] The global constraint cost is calculated based on the statistical results of the number of cells whose service waiting time exceeds the inherent cell polling waiting time in each time slot;

[0021] Based on the state in each discrete time slot, the dynamic beam coverage strategy, the global reward, the global constraint cost, and the state in the next discrete time slot, a training transformation tuple is constructed for each discrete time slot.

[0022] In one implementation, the policy network is optimized using a centralized training method based on training transformation tuples, thereby improving the scheduling performance of multi-satellite multi-beam networks, including:

[0023] The framework adopts a centralized training method. During the training phase, the reward and cost evaluation network of each agent learns by utilizing the global state and the joint actions of all agents, thereby guiding the policy network to learn an efficient distributed policy.

[0024] Sample target training transformation tuples from the experience replay buffer; wherein, the experience replay buffer stores training transformation tuples under multiple discrete time slots;

[0025] The reward evaluation network and cost evaluation network are trained using target training transformation tuples with the goal of minimizing the mean square error function.

[0026] Furthermore, the policy network is trained using training transformation tuples, combined with the inherent small cell polling constraints, with the goal of maximizing the value of the Lagrange objective function; wherein, the Lagrange objective function is constructed by integrating the original reward optimization objective function and the constraint violation penalty term weighted by the dual variables, and the original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution.

[0027] In one implementation, the expression for the Lagrange objective function is as follows:

[0028] ;

[0029] in, The current strategy to be optimized is... Optimize the objective function for the original reward. To constrain violations and penalties, As dual variables, To accumulate the expected cost of constraints, For long-term constraint budgets, it is used to limit the expected value of the cumulative violation cost of the system throughout the entire policy execution process.

[0030] In one implementation, the expression for the original reward optimization objective function is as follows:

[0031] ,

[0032] ;

[0033] in, The current strategy to be optimized is... For strategy The long-term discount benefits, This represents the discount factor, used to control the degree to which a strategy focuses on future returns. For instant rewards, The specific expression for the immediate constraint cost item is as follows:

[0034] ,

[0035] ;

[0036] in, The weighting factor controls the balance between user access rate and service latency. To determine the number of users who can be successfully connected, The total number of users included in the ground cells covered by the intelligent agent. To account for total service delay, To maximize the allowable waiting delay, i.e. the inherent cell polling delay constraint, This is an indicator function; when the waiting time of any cell exceeds a threshold... When the time condition is met, the value is 1; otherwise, it is 0. Indicates the community In the time slot Service wait time.

[0037] In one implementation, the method further includes:

[0038] Based on the constrained cost and constrained budget, the primal-dual optimization method is used to dynamically update the dual variables and obtain new dual variables.

[0039] Secondly, the present invention also provides a multi-satellite beam hopping scheduling device for random access, comprising:

[0040] The model acquisition module is used to acquire a pre-established low-Earth orbit satellite uplink random access communication system model. The system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multiple time slots. The system model also constructs an interactive environment for beam hopping scheduling based on beam hopping time schedules and user random access behavior rules.

[0041] The data determination module is used to determine the global state of the system model in each discrete time slot, and to determine the local observation data of the agent for the ground user cell in each discrete time slot, using the beam of the low-orbit satellite as an agent; wherein, the global state includes the number of active users and the cumulative service waiting time of all ground cells; the local observation data includes the number of active users and the cumulative service waiting time of each cell that can be served by each beam.

[0042] The policy interaction module is used to generate and execute the beam coverage policy of the agent in each discrete time slot to complete the interaction with the environment. Each agent generates a dynamic beam coverage policy based on local state information using a policy network constructed with a Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage policy, and the environment updates its state according to the user access results and latency changes under the CRDSA protocol, and feeds back reward and constraint cost signals to the agent. The local state information, dynamic beam coverage policy, reward signal, constraint cost signal, and the next local state information together constitute the training transition tuple.

[0043] The model training module is used to optimize the policy network in a centralized training manner based on training transformation tuples, thereby improving the scheduling performance of multi-satellite multi-beam networks. The system adopts the multi-agent deep deterministic policy gradient (MADDPG) framework, constructs a reward and cost dual evaluation network in the centralized training phase, and samples transformation tuples in the experience pool for policy updates. At the same time, the primal-dual optimization method is introduced to adjust the policy objective through Lagrange multipliers to meet the time delay constraint.

[0044] Thirdly, the present invention also provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.

[0045] Fourthly, the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.

[0046] The present invention provides a multi-satellite hop beam scheduling method for random access. First, a pre-established uplink random access communication system model for low-Earth orbit satellites is obtained. The system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multiple time slots. The model also constructs an interactive environment for beam hopping scheduling based on rules such as beam hopping time schedules and user random access behavior. Then, it determines the global state of the system model in each discrete time slot, and uses the low-Earth orbit satellite beams as agents to determine the local observation data of the agents for ground user cells in each discrete time slot. The global state includes the number of active users and cumulative service waiting time for all ground cells; the local observation data includes the number of active users and cumulative service waiting time for each beam-servable cell. Next, it generates and executes the agent's beam coverage strategy in each discrete time slot, completing the environmental interaction. Each agent generates a dynamic beam coverage strategy based on local state information using a policy network constructed using the Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage strategy, and the environment updates its state according to user access results and latency changes under the CRDSA protocol, feeding back reward and constraint cost signals to the agent. The aforementioned states, actions, rewards, constraint costs, and the next state together constitute the training transition tuple. Finally, the policy network is optimized using a centralized training approach based on the training transition tuple, thereby improving the scheduling performance of the multi-satellite, multi-beam network. Specifically, the system employs a multi-agent deep deterministic policy gradient (MADDPG) framework, constructing a dual-evaluation network for rewards and costs during the centralized training phase, and sampling transition tuples from the experience pool for policy updates. Simultaneously, a primal-dual optimization method is introduced, adjusting the policy objective through Lagrange multipliers to meet latency constraints. This method generates dynamic beam allocation actions through a secure diffusion reinforcement learning model, effectively improving beam coverage flexibility and user access efficiency. Furthermore, by training the model with the objective of maximizing the number of successfully accessed users and minimizing total service latency, it ensures that while improving user access success rate, service quality requirements are strictly met, significantly improving resource utilization efficiency and system reliability.

[0047] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0048] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0050] Figure 1 A flowchart illustrating a multi-star hop beam scheduling method for random access provided in an embodiment of the present invention;

[0051] Figure 2 A schematic diagram illustrating a scenario for a multi-star hop beam scheduling method for random access provided in an embodiment of the present invention;

[0052] Figure 3 A diagram illustrating a TDM and CRDSA mechanism provided in an embodiment of the present invention;

[0053] Figure 4 A schematic diagram of a secure diffusion multi-agent reinforcement learning model framework provided in an embodiment of the present invention;

[0054] Figure 5 This is a schematic diagram of the structure of a multi-star hop beam scheduling device for random access provided in an embodiment of the present invention;

[0055] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Currently, traditional fixed-beam or static polling schemes suffer from low access efficiency and poor resource utilization when facing uplink communication scenarios in LEO constellations where user access behavior is highly random and resource requirements change dynamically.

[0058] To overcome the aforementioned limitations, beam hopping technology for random access has become a research hotspot, dynamically adjusting beam direction to quickly respond to hotspots or sudden demands. However, while there is considerable research on downlink beam hopping technology, such as greedy hybrid precoding, time slot allocation based on genetic algorithms, and adaptive beam scheduling incorporating deep reinforcement learning, these methods typically assume complete system observability. Uplink research, on the other hand, focuses primarily on centralized access scenarios, such as beam-carrier allocation to guarantee QoS and service priority scheduling improved by cuckoo search. These methods heavily rely on centralized control and global information, making them unsuitable for scenarios with highly random user behavior and incomplete information. Furthermore, most beam hopping methods remain at the single-satellite system level, neglecting the challenges of coordinated interference and scheduling among multiple satellites in LEO constellations. More critically, uplink beam hopping scheduling must meet strict QoS constraints, especially the maximum waiting delay determined by the polling mechanism. Existing deep reinforcement learning methods primarily optimize average performance, lacking effective modeling and guarantees for security and constraints. Although Safe Reinforcement Learning (SafeRL) has been applied in safety-critical areas such as unmanned systems, research on its integration with Quality of Service (QoS) constraints in LEO uplink hopping beam scheduling has not yet been carried out. This highlights the need for a new framework that can simultaneously optimize system performance and ensure operational safety and reliability in practical satellite communications.

[0059] Based on this, the present invention provides a multi-star hop beam scheduling method for random access, which can significantly improve the user access success rate and effectively reduce service waiting latency.

[0060] To facilitate understanding of this embodiment, a multi-satellite hop beam scheduling method for random access disclosed in this embodiment of the invention will first be described in detail. See [link to relevant documentation]. Figure 1 The diagram shows a flowchart of a multi-star hop beam scheduling method for random access, which mainly includes the following steps S102 to S108:

[0061] Step S102: Obtain the pre-established low-orbit satellite uplink random access communication system model.

[0062] The system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multi-time slots. The model also constructs an interactive environment for beam hopping scheduling based on rules such as beam hopping time schedules and user random access behavior. Users can achieve random access through contention-based resolution of the ALOHA (CRDSA) diversity time slot protocol.

[0063] Step S104: Determine the global state of the system model in each discrete time slot, and use the beam of the low-orbit satellite as an agent to determine the local observation data of the agent for the ground user cell in each discrete time slot.

[0064] The global status includes the number of active users and cumulative service waiting time for all ground cells; the local observation data includes the number of active users and cumulative service waiting time for each cell that can be served by each beam. In practical applications, low-Earth orbit satellites have multiple beams, and each beam acts as an independent intelligent agent.

[0065] Step S106: Generate and execute the beam coverage strategy of the agent in each discrete time slot to complete the environmental interaction.

[0066] Specifically, a policy network constructed using the Conditional Generative Diffusion Model (GDM) generates a dynamic beam coverage strategy for each agent in each discrete time slot based on local observation data. The agent executes the beam coverage strategy to interact with the low-Earth orbit satellite communication system environment, which updates its state based on user access results and latency changes under the CRDSA protocol. The total number of successfully accessed users in each time slot is determined based on the statistical results of users who successfully completed uplink access via the random access protocol in the beam-covered cells. The total service waiting latency for all users in each time slot is determined based on the service waiting time of users who failed to access successfully. The global reward for all agents in each time slot is calculated based on the total number of successfully accessed users and the total service waiting latency. The global constraint cost is calculated based on the statistical results of the number of cells in each time slot whose service waiting time exceeds the inherent cell polling waiting latency. Finally, training transformation tuples for each discrete time slot are constructed based on the state, dynamic beam coverage strategy, global reward, global constraint cost, and the state in the next discrete time slot.

[0067] Step S108: Optimize the policy network using a centralized training method based on the training transformation tuple, thereby improving the scheduling performance of the multi-satellite multi-beam network.

[0068] In one example, a centralized training approach is adopted. During the training phase, the reward and cost evaluation networks of each agent learn using the global state and the joint actions of all agents, thereby guiding the policy network to learn an efficient distributed policy. Following the aforementioned step S106, the dynamic beam coverage policy and its execution policy of the agents are obtained for different discrete time slots, yielding corresponding rewards and constraint costs. Training transformation tuples corresponding to different discrete time slots are then constructed, and all training transformation tuples are stored in the experience replay buffer. Target training transformation tuples are sampled from the experience replay buffer and used to train the beam coverage policy network, reward evaluation network, and cost evaluation network, respectively. When training the reward evaluation network and cost evaluation network, the objective is to minimize the mean square error function. When training the beam coverage policy network, the objective is to maximize the Lagrange objective function. The Lagrange objective function is constructed by integrating the original reward optimization objective function and a constraint violation penalty term weighted by dual variables. The original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution.

[0069] The multi-star hop beam scheduling method for random access provided in this invention generates dynamic beam coverage strategies through a policy network constructed by a conditional generation diffusion model (GDM), effectively improving the flexibility of beam coverage and user access efficiency. At the same time, by training the model with the goal of maximizing the number of successfully accessed users and minimizing service waiting latency, it ensures that while improving the user access success rate, it strictly meets the quality of service requirements, significantly improving resource utilization efficiency and system reliability.

[0070] In LEO satellite communication networks, traditional fixed-beam or static polling schemes suffer from low access efficiency and poor resource utilization in uplink communication scenarios where user access behavior is highly random and resource demands dynamically change. This invention proposes a dynamic beam resource optimization method based on secure diffusion multi-agent constrained reinforcement learning. Specifically, it includes: fusing multi-agent deep deterministic policy gradient (MADDPG) to achieve collaborative learning and resource allocation among multiple satellites; replacing the traditional deterministic Actor network with a conditional generative diffusion model (GDM) to generate diverse and robust beam allocation strategies, enhancing policy exploration capabilities; and introducing a primal-dual optimization method, using inherent QoS indicators such as cell polling latency as security constraints to ensure that the system strictly meets preset QoS security constraints while optimizing user access and reducing latency. This method can significantly improve user access success rate, reduce service latency, and continuously meet QoS security constraints, demonstrating good dynamic adaptability and promising practical application prospects.

[0071] For ease of understanding, this embodiment of the invention provides a specific implementation of a multi-star hop beam scheduling method for random access.

[0072] (a) System modeling and optimization goal definition.

[0073] See Figure 2 The diagram illustrates a scenario of a multi-satellite hop beam scheduling method for random access, showcasing a low-Earth orbit (LEO) multi-satellite hop beam communication network. It details the network environment of the LEO satellite communication system, including satellites, users, and the random access mechanism. This embodiment of the invention considers the uplink of a multi-LEO satellite communication system. The LEO satellite communication system model includes... A constellation of LEO satellites, for Coverage is provided by individual terrestrial cells (i.e., terrestrial user cells). Within a given service area, each satellite can generate [a certain number of cells]. Each independent steerable beam can simultaneously cover Each cell shares uplink resources. Therefore, the number of beams in the system is... Let the satellite set be denoted as The community is a collection of . No. The subset of cells covered by a satellite is denoted as The wide-area coverage and rapid movement of LEO satellites result in uneven geographical distribution of users, creating highly dynamic service demands. , No. The number of users in each community is represented as .

[0074] See Figure 3 The diagram illustrates a TDM and CRDSA mechanism. The system employs beam hopping technology to provide ubiquitous and flexible connectivity, and implements discrete-time operation through a Time Division Multiplexing (TDM) scheme. The Beam Hopping Time Schedule (BHTP) defines periodic frames with a frame length of... Divided into There are discrete time slots, each with a duration of [duration]. This represents the minimum dwell time of the beam within a cell. Each time slot is further divided into multiple sub-time slots to support random access. Within each time slot, the satellite system performs dynamic beam scheduling to determine which cells to serve. Due to limited beam resources and constantly changing user demands, not all access requests can be met immediately, resulting in service latency. Cell In the time slot Service wait time is recorded as This indicates the delay in the cell's wait for beam coverage service up to the current moment, and this indicator reflects the user's quality of service (QoS).

[0075] Within the covered cell, users employ the Contention-Solving Diversity Slotted ALOHA (CRDSA) protocol for random access. CRDSA improves system throughput and access success rate by introducing redundant packet copies and Iterative Interference Cancellation (SIC) mechanisms. Users generate a fixed number of packet copies and randomly place them in predefined access sub-slots. The receiver iteratively decodes packets using the SIC mechanism until no more packets can be decoded.

[0076] A user is considered successfully served when they successfully access the network via the CRDSA protocol while their cell is covered by the beam. Conversely, users who fail to access the network or need to wait for subsequent beam coverage will experience service latency. Therefore, the number of users successfully served and the service wait time are two key metrics for evaluating uplink system performance.

[0077] Based on the above system model, the main system objectives include maximizing the number of successfully connected users and minimizing user service latency. A direct optimization approach is to jointly optimize these two objectives.

[0078] Define time Total number of successfully connected users for: ;in, It is a binary variable, representing the time slot. satellite Can you provide services for the community? Provides beam coverage; Indicates in time slot residential area The number of users selected and successfully connected.

[0079] Define total service wait time for: ;in, Indicates the community In the time slot Service delays.

[0080] It should be noted that there is a conflict between these two objectives: maximizing the number of connected users may increase waiting time in some areas, while minimizing waiting time may reduce the overall number of connected users; moreover, their scales are inconsistent, making direct joint optimization difficult. Therefore, a weighted sum is introduced to flexibly adjust the trade-off between throughput and latency.

[0081] The optimization problem can be expressed as:

[0082]

[0083] In the objective function described above, the first term is the normalized number of successfully connected users, which is the sum of the number of successfully served users and the total number of nominal users. The ratio; the second term is the normalized service wait time penalty term; weight parameters Balancing the control of throughput and latency are two objectives.

[0084] The optimization problem is subject to the following constraints: C1 ensures This is a binary variable indicating whether the cell is covered. C2 is limited to any time slot. Each star can serve a maximum of no more than [number] cells. C3 ensures that every community In the time slot Service delay shall not exceed the maximum allowed delay Finally, C4 is defined. The effective weighting factor is between 0 and 1.

[0085] Although the optimization problem is described as a weighted problem, it remains difficult to solve due to the dynamic changes in user states, the randomness of CRDSA access results, and the large beam scheduling action space. Therefore, it is modeled as a temporal decision problem, and a deep reinforcement learning (DRL) framework is introduced, using the optimization objective function as a reward signal to guide the policy learning process.

[0086] (II) Building a reinforcement learning model for security diffusion.

[0087] To satisfy the latency constraints during training, this paper models the problem as a constrained Markov decision process (CMDP). The goal is to maximize the expected cumulative reward while ensuring that the long-term expected cumulative constraint cost does not exceed the system's long-term constraint budget.

[0088] The objective function after modeling can be expressed as:

[0089]

[0090] in, The current strategy to be optimized is... For strategy The expected long-term discount benefit. This is a discount factor used to control the strategy's focus on future returns. For long-term constraint budgets, this limits the expected value of the cumulative violation cost of the system throughout the entire policy execution process. Where, Corresponding to the optimization objective Instantaneous performance indicators used to reflect system utility; and This is the instantaneous constraint cost function, used to characterize delay violations; and These represent states and actions, respectively; detailed definitions are provided below.

[0091] To address CMDP, a primordial-dual optimization (PDO) approach is employed. Specifically, Lagrange multipliers are introduced to jointly optimize the beam scheduling strategy and time delay constraints. The time delay constraint is incorporated into the objective function through a dual variable, the Lagrange multiplier, and iteratively updated during training. Therefore, we reformulate the constraint optimization as the following Lagrange objective:

[0092]

[0093] in, These are Lagrange multipliers. This form allows the use of PDOs to learn a policy that maximizes reward while satisfying time delay constraints. .

[0094] Based on the above, to address the CMDP problem, we propose the SD-MACRL algorithm, which integrates PDO into the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework. The system employs a centralized training and distributed execution framework, where each beam acts as an agent making local decisions while sharing global feedback during training.

[0095] Each beam in the system is considered an independent intelligent agent, and there are a total of [number missing] beams in the system. Several agents are involved. The dynamic LEO satellite hopping beam uplink access system is modeled as a DRL environment, in which these agents operate. The agents interact with the environment by executing beam coverage strategies, and the environment updates its state based on user access results and latency changes under the CRDSA protocol, and feeds back reward and cost signals to the agents.

[0096] The detailed definition of the SD-MACRL component is as follows:

[0097] state space Intelligent agent Able to observe its visible cell set The current number of users in the current time slot and the service wait time in the medium-sized cell. An intelligent agent at time Local observations are recorded as ;in, For any , This indicates the inference obtained by combining recent CRDSA access results (such as slot-level collisions and successful transmissions). Number of active users in each community This indicates the cumulative service wait time for the corresponding cell. Global Status Defined as the joint state of all agents.

[0098] Action space Intelligent agent In each time slot Selecting a cell for beam coverage, the action is denoted as... The combined actions of all intelligent agents constitute ;

[0099] Instant Rewards All agents share the same immediate reward, which is used to characterize the overall system performance and is defined as follows: Among them, weighting factors By controlling the balance between user access rate and service latency, this shared reward design encourages collaborative learning among all beams.

[0100] Immediately constrained cost item The immediate constraint cost term is a binary variable used as a system-level default indication signal, defined as follows: ;in, This is an indicator function; when the waiting time of any cell exceeds a threshold... When the time condition is met, the value is 1; otherwise, it is 0. Indicates the community In the time slot Service wait time.

[0101] Strategy Each agent learns a strategy. Local observations are mapped to a beam scheduling strategy, which is implemented by a conditional generation and diffusion model (GDM).

[0102] In each time slot Each agent interacts with the environment. The environment updates its state based on user access results and latency changes under the CRDSA protocol and feeds back reward signals to the agents. With constraint cost signals Each intelligent agent Receive its local observations Select a community to provide services (action) ), and obtain global rewards based on environmental feedback. and cost .

[0103] To address the CMDP problem, a constrained reinforcement learning framework based on the security diffusion model was built using PDO. This framework includes an Actor network based on the conditional generative diffusion model (GDM) to generate beam coverage strategies. Additionally, a reward evaluation network and a cost evaluation network were constructed, and a primal-dual optimization PDO mechanism was integrated to enforce QoS security constraints.

[0104] This invention establishes a secure diffusion multi-agent reinforcement learning model architecture. It is built upon the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework, see [link to relevant documentation]. Figure 4 The diagram illustrates a security diffusion multi-agent reinforcement learning model framework, primarily comprising two key components: first, replacing the traditional deterministic Actor network with a conditional generative diffusion model (GDM) as the beam coverage policy network; and second, integrating security constraints to enforce latency sensitivity requirements. This design leverages the expressive power of the diffusion model for policy exploration while ensuring real-time feasibility through constrained reinforcement learning. Specifically, the model includes multiple agents, each possessing a beam coverage policy network. and a shared reward and evaluation network. and cost evaluation network ,in, , , These represent the model parameters of the policy network, reward evaluation network, and cost evaluation network, respectively. Based on this, a dynamic beam coverage policy for the agent under the current discrete time slot is generated using a policy network based on a conditional generation diffusion model, taking into account the agent's local observation data and noise. The dynamic beam coverage policy is a binary decision by the agent regarding whether to perform beam coverage on the ground user cells it covers. The agent interacts with the environment by executing the beam coverage policy, and the environment updates its state according to the user access results and latency changes under the CRDSA protocol, feeding back reward and constraint cost signals to the agent. The aforementioned state, action, reward, constraint cost, and next state together constitute the training transition tuple. The system adopts a multi-agent deep deterministic policy gradient (MADDPG) framework, constructing a dual-evaluation network of reward and cost during the centralized training phase, sampling transition tuples from the experience pool for policy updates; simultaneously, a primal-dual optimization method is introduced, adjusting the policy objective through Lagrange multipliers to satisfy latency constraints.

[0105] To enhance policy exploration capabilities and achieve constraint-aware decision-making, this embodiment of the invention replaces the traditional deterministic Actor in MADDPG with a Conditional Generative Diffusion Model (GDM), generating beam coverage policies through a conditional inverse denoising process. During policy generation, the policy network receives local observation data from the agent during the inverse process. As a conditional input, the noise is gradually refined into a dynamic beam coverage strategy that can balance reward maximization and constraint satisfaction through the iterative sampling process of the Denoising Diffusion Implicit Model (DDIM).

[0106] The policy generator in this embodiment of the invention employs the GDM framework and utilizes the Denoising Diffusion Implicit Model (DDIM) to accelerate the action sampling process. Each agent... The conditional generation diffusion model is used as its Actor network, denoted as . .in, Local observation The mapping is performed as a probability distribution of cell selection actions, a process based on a denoising sampling mechanism. Unlike traditional deterministic strategies, this modeling approach can generate diverse, high-quality actions with constraint awareness.

[0107] Specifically, GDM uses a reverse denoising process to transform a random noise vector... Sampling is performed as an action vector, and this process is based on local observations. As a condition. This denoising process follows an update rule similar to DDIM:

[0108] ,

[0109] ;

[0110] in, The initial action vector predicted based on the observations. It is a diffusion step The intermediate vector in the middle, It is determined by parameters The noise prediction model is represented. These are predefined noise scheduling parameters, and This is a local observation. After denoising, the final sampled value... Discretized into a beam coverage strategy This ensures its compatibility with the discrete action space of the environment. Therefore, As an intelligent agent The policy function implicitly represents this complete denoising generation process.

[0111] In addition, the model introduces a primordial-dual optimization (PDO) mechanism, which maintains a Lagrangian dual variable. This method is used to manage QoS security constraints. Specifically, this invention proposes a security reinforcement learning method based on the Constrained Markov Decision Process (CMDP) framework. Unlike standard MDPs, CMDPs include both reward and cost functions, and their goal is to optimize the policy. The goal is to maximize the expected cumulative reward while satisfying all safety constraints (referred to as constraints). Specifically, the expected cumulative constraint cost is... It must not exceed its constraint threshold. Among them, strategy This can be understood as a mapping from local observation data to a dynamic beam coverage strategy. Formally, this is expressed as:

[0112]

[0113] This is the optimal strategy. To solve the CMDP problem within the MADDPG framework, this embodiment of the invention employs Primitive-Dual Optimization (PDO). PDO introduces Lagrange multipliers (dual variables) for each constraint. This transforms the constrained problem into an unconstrained optimization task. The policy network is trained to maximize an augmented Lagrangian objective function that combines the original reward with a constraint violation penalty term weighted by the dual variable. Therefore, the objective function becomes:

[0114] ;

[0115] in, The current strategy to be optimized is... Optimize the objective function for the original reward. To constrain violations and penalties, As dual variables, To accumulate the expected cost of constraints, For long-term constraint budgets, it is used to limit the expected value of the cumulative violation cost of the system throughout the entire policy execution process.

[0116] The expressions for the original reward optimization objective function and the expected cumulative constraint cost are as follows:

[0117] ,

[0118] ;

[0119] in, The current strategy to be optimized is... The expected cost is estimated by the cost evaluation network. This represents the discount factor, used to control the degree to which a strategy focuses on future returns. For instant rewards, This is to constrain cost items immediately.

[0120] To estimate the expected value of the constraint cost This invention employs a dedicated Cost Critic Network. The output of this network directly determines the adaptive update of the dual variable. In one implementation, the update is based on the constraint cost and the long-run constraint budget. The dual variable is dynamically updated to obtain a new dual variable. For details, please refer to the following rules for updating the dual variable:

[0121] ;in For learning rate, The expected value of the expected constraint cost estimated by the cost evaluation network. The threshold is used as a constraint. This iterative primal-dual update mechanism ensures that the learned policy maximizes rewards while satisfying safety constraints.

[0122] (III) Iterative Training of the Policy Network. The proposed model is iteratively optimized through continuous interaction between the agent and the simulation environment. Through experience replay and network parameter updates, the model learns strategies for efficiently scheduling beams in dynamic environments and maximizing system performance while satisfying various QoS security constraints. The training process is conducted through continuous interaction between the agent and the environment; details are as follows:

[0123] (a) Before training begins, the system initializes all networks, including the GDM policy network, reward evaluation network and cost evaluation network, as well as their corresponding target networks, and initializes the Lagrange dual variables and experience replay buffers used for constraint management.

[0124] (b) Based on the global state in the current discrete time slot, the dynamic beam allocation action, the expected reward, the expected constraint cost and the global state in the next discrete time slot, construct the training transformation tuple in the current discrete time slot.

[0125] Specifically, the training process iterates in episodes. At each time step of each episode, the agent first generates a dynamic beam coverage action based on its local observation data and through its GDM policy network, combined with sampled noise. Once the joint action of all agents is executed, the environment will return an immediate reward. ,cost And new status information and will transform tuples Stored in the experience replay buffer for use in subsequent training.

[0126] (c) Sample target training transformation tuples from the experience replay buffer; wherein the experience replay buffer stores training transformation tuples in multiple discrete time slots.

[0127] Specifically, small batches of data are sampled from the buffer for network updates.

[0128] (d) Using the target training transformed tuples, the reward evaluation network and cost evaluation network are trained with the objective of minimizing the mean squared error function. Reward Evaluation Network and cost evaluation network Updates are performed by minimizing their respective mean squared error (MSE) loss functions.

[0129] Specifically, the target Q-values ​​of the reward evaluation network and the cost evaluation network are first calculated. These target values ​​utilize the target network to improve stability. The reward evaluation network and the cost evaluation network are updated by minimizing the mean square error between the predicted Q-value and the target Q-value. The target Q-value represents the estimate of the future reward after the agent takes an action under the current policy and is mainly used to update the network; it is calculated using the predicted Q-value of the next state after adding a discount to the current reward (Bellman equation).

[0130] (e) The policy network is trained using the target training transformation tuple, combined with security constraints, with the goal of maximizing the value of the Lagrange objective function; wherein the Lagrange objective function is constructed by integrating the original reward optimization objective function and the constraint violation penalty term weighted by the dual variables, and the original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution.

[0131] Specifically, the GDM policy network updates by maximizing an augmented Lagrangian objective function that combines rewards with a constraint cost penalty term weighted by a dual variable. The update of the GDM policy network aims to maximize an augmented Lagrangian objective function that not only considers rewards but also penalizes costs through a Lagrangian dual variable, thereby prompting the GDM policy network to learn a policy that maximizes performance while satisfying QoS security constraints.

[0132] (f) Based on the constrained costs and long-term constrained budgets, the dual variables are dynamically updated to obtain new dual variables. Dual variables The output of the cost evaluation network and the long-term constraint budget are dynamically updated to enforce compliance with QoS security constraints.

[0133] Specifically, the Lagrange dual variable is also dynamically updated based on the output of the cost evaluation network and the long-run constraint budget. If the cost caused by the current policy exceeds the long-run constraint budget, the corresponding dual variable will increase, thereby imposing a greater penalty on the optimization objective of the GDM policy network and forcing it to adjust its policy to better meet the constraints.

[0134] Furthermore, to enhance the stability and convergence of the training, the online parameters of each network are periodically soft-updated to their corresponding target network during the training process. Through iterative training, the model can learn to efficiently perform multi-satellite cooperative beam scheduling in the dynamically changing LEO satellite communication environment, and maximize user access success rate and minimize service latency while strictly meeting quality of service constraints.

[0135] In summary, this invention addresses the problems of low uplink user random access efficiency and increased service latency in traditional multi-beam systems under dynamic user traffic in low Earth orbit (LEO) satellite communication networks, and improves existing beam scheduling and resource allocation methods. Experimental verification shows that this invention achieves better performance than traditional methods in both improving user access success rate and reducing service latency. This invention proposes the following three core technical points, as follows: (1) A dynamic beam strategy generator based on a conditional generation diffusion model is proposed, replacing the traditional deterministic Actor network in the multi-agent deep deterministic policy gradient framework with a conditional generation diffusion model. (2) A QoS security constraint enforcement mechanism based on primordial-dual optimization (PDO) is designed, explicitly modeling key service quality indicators as security constraints in the reinforcement learning process. By introducing the primordial-dual optimization method, the Lagrange dual variable is dynamically adjusted during policy training to punish behaviors that violate the constraints. (3) A novel integrated framework of "safe diffusion multi-agent constrained reinforcement learning" was proposed and constructed, which creatively integrates the multi-agent collaborative decision-making capability of MADDPG, the policy generation diversity of conditional GDM and the strict safety constraint mechanism of PDO, so as to achieve the collaborative optimization of global system performance and safe and reliable operation.

[0136] Based on this, the technical effects of the embodiments of the present invention are as follows: By adopting a user-demand-centered goal setting, combined with multi-agent collaborative learning and dynamic beam coverage generation strategies, joint optimization of user access efficiency and service latency is achieved; in the beam coverage strategy generation part, the beam position is dynamically adjusted through a secure diffusion multi-agent constraint reinforcement learning framework, effectively improving the flexibility of beam coverage and user access efficiency; in the security constraint part, by introducing primordial-dual optimization to strictly constrain service waiting time, the system ensures that while improving user access success rate, it strictly meets service quality requirements, significantly improving system reliability. Compared with traditional static or heuristic beam allocation schemes, the embodiments of the present invention have stronger environmental adaptability and system scalability, and can effectively cope with the rapid movement of low-orbit satellites and the non-uniform dynamic changes in user traffic, continuously maintaining optimized performance. By fusing MADDPG to achieve multi-agent collaboration and using the generative diffusion model GDM to generate diverse strategies, the present invention achieves efficient collaboration and flexible scheduling of beams among multiple satellites, significantly improving system throughput (i.e., user access success rate) and user service capabilities while ensuring system security constraints.

[0137] Based on the foregoing embodiments, this invention provides a multi-satellite beam hopping scheduling device for random access, see [link to previous document]. Figure 5 The diagram shows a multi-satellite hop beam scheduling device for random access. The device mainly includes the following parts:

[0138] The model acquisition module 502 is used to acquire a pre-established low-Earth orbit satellite uplink random access communication system model. The system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite, used to simulate dynamic beam coverage under multi-time slots. The model also constructs an interactive environment for beam hopping scheduling based on rules such as beam hopping time schedules and user random access behavior.

[0139] The data determination module 504 is used to determine the global state of the system model in each discrete time slot, and to determine the local observation data of the agent for the ground user cell in each discrete time slot, taking the low-orbit satellite beam as an agent; wherein, the global state includes the number of active users and the cumulative service waiting time of all ground cells; the local observation data includes the number of active users and the cumulative service waiting time of each beam-servable cell.

[0140] The policy interaction module 506 is used to generate and execute the beam coverage policy of the agent in each discrete time slot to complete the environmental interaction. Each agent generates a dynamic beam coverage policy based on local state information using a policy network constructed with a Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage policy, and the environment updates its state according to the user access results and latency changes under the CRDSA protocol, and feeds back reward and constraint cost signals to the agent. The aforementioned state, action, reward, constraint cost, and next state together constitute the training transition tuple.

[0141] The model training module 508 is used to optimize the policy network in a centralized training manner based on training transformation tuples, thereby improving the scheduling performance of multi-satellite multi-beam networks. The system adopts the multi-agent deep deterministic policy gradient (MADDPG) framework, constructs a reward and cost dual evaluation network in the centralized training phase, and samples transformation tuples in the experience pool for policy updates. At the same time, the primal-dual optimization method is introduced to adjust the policy objective through Lagrange multipliers to meet the time delay constraint.

[0142] The multi-star hop beam scheduling device for random access provided in this invention generates dynamic beam allocation actions through a secure diffusion reinforcement learning model, which effectively improves the flexibility of beam coverage and user access efficiency. At the same time, by training the model with the goal of maximizing the number of successfully accessed users and minimizing service waiting latency, it ensures that while improving the user access success rate, it strictly meets the quality of service requirements, which significantly improves resource utilization efficiency and system reliability.

[0143] In one implementation, the model acquisition module 502 is specifically used for:

[0144] Beam hopping time scheduling is based on TDM to dynamically schedule beam resources within discrete time slots in order to achieve periodic coverage of the target cell;

[0145] The user random access behavior adopts the contention-based access protocol CRDSA, which enables users to send replica packets through randomly selected access sub-slots during the beam dwell period, and improves decoding efficiency through serial interference cancellation.

[0146] In one implementation, the strategy interaction module 506 is specifically used for:

[0147] The policy network constructed by the conditional generation diffusion model (GDM) generates a dynamic beam coverage strategy for each discrete time slot based on the agent's local observation data. The dynamic beam coverage strategy is a binary decision by which the agent performs beam coverage on the ground user cell it covers, and is used to indicate the position of each beam in each time slot.

[0148] The intelligent agent completes the interaction with the low-Earth orbit satellite communication system environment by executing the beam coverage strategy. The environment updates its status according to the user access results and latency changes under the CRDSA protocol.

[0149] The total number of users who successfully accessed the network in each time slot is determined by the statistical results of users who successfully completed uplink access through the random access protocol in the cells covered by the beam within each time slot.

[0150] Based on the service waiting time of users who failed to successfully access the service in each time slot, the total service waiting time of all users in each time slot is determined.

[0151] Calculate the global reward for all agents in each time slot based on the total number of successfully connected users and the total service waiting latency in each time slot.

[0152] The global constraint cost is calculated based on the statistical results of the number of cells whose service waiting time exceeds the inherent cell polling waiting time in each time slot;

[0153] Based on the state in each discrete time slot, the dynamic beam coverage strategy, the global reward, the global constraint cost, and the state in the next discrete time slot, a training transformation tuple is constructed for each discrete time slot.

[0154] In one implementation, the model training module 508 is specifically used for:

[0155] The framework adopts a centralized training method. During the training phase, the reward and cost evaluation network of each agent learns by utilizing the global state and the joint actions of all agents, thereby guiding the policy network to learn an efficient distributed policy.

[0156] Sample target training transformation tuples from the experience replay buffer; wherein, the experience replay buffer stores training transformation tuples under multiple discrete time slots;

[0157] The reward evaluation network and cost evaluation network are trained using target-trained transformation tuples with the goal of minimizing the mean square error function.

[0158] Furthermore, the policy network is trained using training transformation tuples, combined with the inherent small cell polling constraints, with the goal of maximizing the value of the Lagrange objective function; wherein, the Lagrange objective function is constructed by integrating the original reward optimization objective function and the constraint violation penalty term weighted by the dual variables, and the original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution.

[0159] In one implementation, the expression for the Lagrange objective function is as follows:

[0160] ;

[0161] in, The current strategy to be optimized is... Optimize the objective function for the original reward. To constrain violations and penalties, As dual variables, To accumulate the expected cost of constraints, For long-term constraint budgets, it is used to limit the expected value of the cumulative violation cost of the system throughout the entire policy execution process.

[0162] In one implementation, the expression for the original reward optimization objective function is as follows:

[0163] ,

[0164] ;

[0165] in, The current strategy to be optimized is... For strategy The long-term discount benefits, This represents the discount factor, used to control the degree to which a strategy focuses on future returns. For instant rewards, The specific expression for the immediate constraint cost item is as follows:

[0166] ,

[0167] ;

[0168] in, The weighting factor controls the balance between user access rate and service latency. To determine the number of users who can be successfully connected, The total number of users included in the ground cells covered by the intelligent agent. To account for total service delay, To maximize the allowable waiting delay, i.e. the inherent cell polling delay constraint, This is an indicator function; when the waiting time of any cell exceeds a threshold... When the time condition is met, the value is 1; otherwise, it is 0. Indicates the community In the time slot Service wait time.

[0169] In one implementation, the model training module 508 is further configured to:

[0170] Based on the constrained cost and constrained budget, the primal-dual optimization method is used to dynamically update the dual variables and obtain new dual variables.

[0171] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0172] This invention provides an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.

[0173] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected through the bus 62. The processor 60 is used to execute executable modules, such as computer programs, stored in the memory 61.

[0174] The memory 61 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0175] Bus 62 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0176] The memory 61 is used to store programs. After receiving an execution instruction, the processor 60 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.

[0177] Processor 60 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 60 or by instructions in software form. Processor 60 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 61. Processor 60 reads the information in memory 61 and, in conjunction with its hardware, completes the steps of the above method.

[0178] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.

[0179] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0180] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-satellite beam hopping scheduling method for random access, characterized in that, include: Obtain a pre-established low-Earth orbit satellite uplink random access communication system model; the system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multiple time slots. The system model also constructs an interactive environment for beam hopping scheduling based on beam hopping time schedules and user random access behavior rules. The global state of the system model is determined in each discrete time slot, as well as the beam configuration agent for the low-Earth orbit satellite. The local observation data of the agent for the ground user cell is determined in each discrete time slot. The global state includes the number of active users and the cumulative service waiting time of all ground cells. The local observation data includes the number of active users and the cumulative service waiting time of each beam-servable cell. In each discrete time slot, a beam coverage strategy for the agent is generated and executed. Each agent generates a dynamic beam coverage strategy based on local state information using a policy network constructed with a Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage strategy. The environment then updates its state according to the user access results and latency changes under the CRDSA protocol, and feeds back reward and constraint cost signals to the agent. Local state information, dynamic beam coverage strategy, reward signal, constraint cost signal, and the next local state information together constitute the training transition tuple. The policy network is optimized using a centralized training method based on training transformation tuples. A multi-agent deep deterministic policy gradient framework is adopted, and a reward and cost dual evaluation network is constructed during the centralized training phase. Transformation tuples in the experience pool are sampled for policy updates. At the same time, the primal-dual optimization method is introduced to adjust the policy through Lagrange multipliers to meet the time delay constraint. This paper optimizes the policy network using a centralized training approach based on training transformation tuples, thereby improving the scheduling performance of multi-satellite multi-beam networks, including: The framework employs a centralized training approach. During the training phase, the reward and cost evaluation network for each agent learns using the global state and the joint actions of all agents, thereby guiding the policy network to learn an efficient distributed policy. Sample target training transformation tuples from the experience replay buffer; wherein, the experience replay buffer stores multiple training transformation tuples under the discrete time slots; The target training transformation tuple is used to train the reward evaluation network and the cost evaluation network with the goal of minimizing the mean square error function; Furthermore, the policy network is trained using the training transformation tuples, combined with the cell polling constraint, with the goal of maximizing the value of the Lagrange objective function; wherein the Lagrange objective function is constructed by integrating the original reward optimization objective function and the constraint violation penalty term weighted by the dual variables, and the original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution. The expression for the Lagrange objective function is as follows: ; in, The current strategy to be optimized is... Optimize the objective function for the original reward. To constrain violations and penalties, For dual variables, i.e., Lagrange multipliers, To accumulate the expected cost of constraints, For long-term constraint budgets, it is used to limit the expected value of the cumulative violation cost of the system throughout the entire policy execution process.

2. The multi-satellite beam hopping scheduling method for random access according to claim 1, characterized in that, Obtain a pre-established model of a low-Earth orbit satellite uplink random access communication system, including: The beam hopping time schedule is based on TDM to dynamically schedule beam resources in discrete time slots in order to achieve periodic coverage of the target cell; The user random access behavior adopts the contention-based access protocol CRDSA, which enables users to send replica packets through randomly selected access sub-slots during the beam dwell period, and improves decoding efficiency through serial interference cancellation.

3. The multi-satellite hop beam scheduling method for random access according to claim 1, characterized in that, Generate and execute the agent's beam coverage strategy in each discrete time slot to complete environmental interactions, including: The policy network constructed by the conditional generation diffusion model (GDM) generates a dynamic beam coverage strategy for the agent in each discrete time slot based on the agent's local observation data. The dynamic beam coverage strategy is a binary decision by which the agent performs beam coverage on the ground user cell it covers, used to indicate the position of each beam in each time slot. The intelligent agent completes the interaction with the low-orbit satellite communication system environment by executing the beam coverage strategy. The environment updates its status according to the user access results and latency changes under the CRDSA protocol. The total number of users who successfully accessed the network in each time slot is determined by the statistical results of users who successfully completed uplink access through the random access protocol in the cells covered by the beam within each time slot. Based on the service waiting time of users who failed to successfully access the service in each time slot, the total service waiting time of all users in each time slot is determined. Calculate the global reward for all agents in each time slot based on the total number of successfully connected users and the total service waiting latency in each time slot. Based on the statistical results of the number of cells whose service waiting time exceeds the cell polling waiting time in each time slot, the global constraint cost is calculated. Based on the state in each discrete time slot, the dynamic beam coverage strategy, the global reward, the global constraint cost, and the state in the next discrete time slot, construct the training transformation tuple for each discrete time slot.

4. The multi-satellite beam hopping scheduling method for random access according to claim 1, characterized in that, The expressions for the original reward optimization objective function and the expected cumulative constraint cost are as follows: , ; in, The current strategy to be optimized is... For strategy The long-term discount benefits, This represents the discount factor, used to control the degree to which a strategy focuses on future returns. For instant rewards, The specific expression for the immediate constraint cost item is as follows: , ; in, The weighting factor controls the balance between user access rate and service latency. To determine the number of users who can be successfully connected, This refers to the total number of users included in the ground cells covered by the intelligent agent. To account for total service delay, To maximize the allowable waiting delay, i.e., the cell polling delay constraint, This is an indicator function; when the waiting time of any cell exceeds a threshold... When the time condition is met, the value is 1; otherwise, it is 0. Indicates the community In the time slot Service wait time.

5. The multi-satellite beam hopping scheduling method for random access according to claim 1, characterized in that, The method further includes: Based on the constraint cost and constraint budget, the primal-dual optimization method is used to dynamically update the dual variables to obtain new dual variables.

6. A multi-satellite beam hopping scheduling device for random access, characterized in that, include: The model acquisition module is used to acquire a pre-established low-Earth orbit satellite uplink random access communication system model. The system model includes multiple low-Earth orbit satellites, multiple ground user cells, and multiple independently oriented beams on each satellite to simulate dynamic beam coverage under multiple time slots. The system model also constructs an interactive environment for beam hopping scheduling based on beam hopping time schedules and user random access behavior rules. The data determination module is used to determine the global state of the system model in each discrete time slot, as well as the beam configuration agent for the low-Earth orbit satellite, and to determine the local observation data of the agent for the ground user cell in each discrete time slot; wherein, the global state includes the number of active users and the cumulative service waiting time of all ground cells; the local observation data includes the number of active users and the cumulative service waiting time of each beam-servable cell; The policy interaction module is used to generate and execute the beam coverage policy of the agent in each discrete time slot. Each agent generates a dynamic beam coverage policy based on local state information using a policy network constructed with a Conditional Generative Diffusion Model (GDM). The agent interacts with the environment by executing the beam coverage policy. The environment then updates its state according to the user access results and latency changes under the CRDSA protocol, and feeds back reward and constraint cost signals to the agent. The local state information, dynamic beam coverage policy, reward signal, constraint cost signal, and the next local state information together constitute the training transition tuple. The model training module is used to optimize the policy network in a centralized training manner based on training transformation tuples. The system adopts the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework, constructs a reward and cost dual evaluation network in the centralized training phase, and samples transformation tuples in the experience pool for policy updates. At the same time, the primal-dual optimization method is introduced to adjust the policy objective through Lagrange multipliers to meet the time delay constraint. The model training module is specifically used for: The framework employs a centralized training approach. During the training phase, the reward and cost evaluation network for each agent learns using the global state and the joint actions of all agents, thereby guiding the policy network to learn an efficient distributed policy. Sample target training transformation tuples from the experience replay buffer; wherein, the experience replay buffer stores multiple training transformation tuples under the discrete time slots; The target training transformation tuple is used to train the reward evaluation network and the cost evaluation network with the goal of minimizing the mean square error function; Furthermore, the policy network is trained using the training transformation tuples, combined with the cell polling constraint, with the goal of maximizing the value of the Lagrange objective function; wherein the Lagrange objective function is constructed by integrating the original reward optimization objective function and the constraint violation penalty term weighted by the dual variables, and the original reward optimization objective function is defined as the expected cumulative reward obtained by the policy during execution. The expression for the Lagrange objective function is as follows: ; in, The current strategy to be optimized is... Optimize the objective function for the original reward. To constrain violations and penalties, For dual variables, i.e., Lagrange multipliers, To accumulate the expected cost of constraints, For long-term constraint budgets, it is used to limit the expected value of the cumulative violation cost of the system throughout the entire policy execution process.

7. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Low earth orbit satellite constellation dynamic hopping beam variation optimization method based on deep reinforcement learning

    CN120342471A

  • Low earth orbit satellite-based communication method and computing apparatus for performing same

    WO2024106948A1