A satellite communication resource double-layer scheduling method and system considering long-term fairness and safety energy efficiency

CN122553975APending Publication Date: 2026-08-11BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]针对多波束卫星网络中安全容量需求与卫星能量可持续性之间的核心冲突,以及高动态窃听威胁下边缘及高危用户容易遭受“远-近”饥饿效应的问题,本发明旨在解决多波束卫星网络中安全与功耗的核心冲突、高窃听区域用户通信保障不足的问题,实现卫星网络低能耗、高安全、强公平的通信目标,保障高窃听风险区域用户获得充足通信资源,满足长期稳定通信需求

Benefits of technology

1.本发明提出多波束卫星信道、干扰、窃听、能效一体化物理建模方法,通过完整数学公式,实现能耗与安全相关参数的精准量化表征,为能耗与公平性协同优化提供基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122553975A_ABST
    Figure CN122553975A_ABST
Patent Text Reader

Abstract

This invention provides a two-layer scheduling method and system for satellite communication resources that balances long-term fairness, security, and energy efficiency. The method includes: establishing a downlink transmission system model for a multi-beam satellite network; collecting network status data based on the downlink transmission system model; pre-modeling a lower-layer optimization solver; constructing an upper-layer intelligent agent model; using the lower-layer optimization solver to perform instantaneous resource orchestration of strategic hyperparameters to obtain a trained upper-layer intelligent agent model; deploying the trained upper-layer intelligent agent model on a satellite to receive real-time network status data and infer the strategic hyperparameters of the current time slot; using the strategic hyperparameters of the current time slot to guide the lower-layer optimization solver in generating a final resource scheduling scheme, and distributing the instantaneous resource scheduling scheme to the multi-beam satellite for execution. This invention achieves the communication goals of low energy consumption, high security, and strong fairness in satellite networks, ensuring that users in high-risk eavesdropping areas receive sufficient communication resources and meeting long-term stable communication needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of 6G non-terrestrial communication network and satellite communication technology, specifically involving a two-layer scheduling method and system for satellite communication resources that takes into account both long-term fairness and security and energy efficiency. Background Technology

[0002] Multi-beam satellites are a core component of 6G integrated space-air-ground networks. However, their broadcast communication characteristics are prone to eavesdropping risks, and onboard power resources are strictly limited, creating a significant trade-off that is difficult to balance effectively. Simultaneously, due to complex inter-beam interference, beam, channel, and power resources are highly coupled. Existing instantaneous greedy resource allocation strategies generally favor low-eavesdropping-risk areas, leading to a severe "far-near starvation" effect for users in high-eavesdropping-risk areas or at the beam edge, resulting in inadequate communication support. Current industry optimization solutions either employ instantaneous static snapshot designs, which cannot adapt to scenarios where space threats change dynamically with satellite transit and cannot guarantee long-term temporal fairness; or rely on pure black-box AI scheduling, lacking the mathematical boundary support of physical constraints. None of these solutions can simultaneously meet the core requirements of low energy consumption, high security, and fairness, and cannot provide stable communication guarantees for users in high-eavesdropping-risk areas. Summary of the Invention

[0003] Addressing the core conflict between security capacity requirements and satellite energy sustainability in multi-beam satellite networks, and the problem of "far-near" starvation effects easily suffered by edge and high-risk users under high-dynamic eavesdropping threats, this invention aims to resolve the core conflict between security and power consumption in multi-beam satellite networks, as well as the problem of insufficient communication guarantees for users in high-eavesdropping areas. It seeks to achieve the communication goals of low energy consumption, high security, and strong fairness in satellite networks, ensuring that users in high-eavesdropping-risk areas have sufficient communication resources and meet their long-term stable communication needs.

[0004] To achieve the above objectives, the present invention provides the following solution: A two-tiered scheduling method for satellite communication resources that balances long-term fairness with security and energy efficiency includes: A downlink transmission system model for a multi-beam satellite network is established, and network status data is collected based on the model. The network status data includes user channel status information, eavesdropping threat probability map, security efficiency statistics, and system physical constraints. A pre-modeled lower-level optimization solver is used, which adopts a dynamic asymmetric Nash bargaining game model and sequentially performs secure beam pointing matching, secure sub-channel allocation and power allocation to generate an instantaneous resource scheduling scheme. A higher-level intelligent agent model is constructed, which generates strategic hyperparameters based on the network state data. The strategic hyperparameters include an energy penalty factor and a fairness sensitivity factor. The lower-level optimization solver is used to perform instantaneous resource orchestration on the strategic hyperparameters to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, thus obtaining a trained upper-level agent model. The trained upper-layer intelligent agent model is deployed on the satellite to receive real-time network status data and infer the strategic hyperparameters of the current time slot. The strategic hyperparameters of the current time slot are used to guide the lower-level optimization solver to generate the final resource scheduling scheme, and the instantaneous resource scheduling scheme is then sent to the multi-beam satellite for execution.

[0005] Preferably, the conditions for establishing a downlink transmission system model for a multi-beam satellite network include: known satellite operating altitude, orbital inclination, communication frequency band, and the distribution pattern of the space threat probability map; legitimate ground users are randomly distributed in the satellite beam coverage area, potential eavesdroppers are randomly distributed in space, and the positions of users and eavesdroppers are both random variables with known distribution ranges; The downlink transmission system model of the multi-beam satellite network includes M satellites, Q spot beams, and N legitimate ground users.

[0006] Preferably, the method for performing secure beam pointing matching includes: A binary conflict matrix is ​​constructed to characterize the spatial isolation constraint between any two candidate beam centers. If the distance between two candidate beam centers is less than a preset minimum beam spacing threshold, or if they are the same beam center, the element at the corresponding position in the binary conflict matrix is ​​set to 1; otherwise, it is set to 0. Based on the potential security utility score of each candidate beam center, and under the premise of satisfying the binary conflict matrix constraint, a greedy strategy is adopted to select conflict-free beam centers as the initial active beam set. Candidate beam centers are selected sequentially from the dynamic candidate pool to replace existing beam centers in the current active beam set. The replacement operation is performed only if the candidate beam center to be replaced does not conflict with any other beam centers in the current active beam set, and the replacement satisfies the preset global weighted security utility enhancement condition, until convergence. The security beam pointing matching result is used to maximize the weighted security coverage while minimizing information leakage to eavesdroppers.

[0007] Preferably, the method for performing secure sub-channel allocation includes: Subchannel allocation is modeled as a matching game between the user set and the subchannel set; Define user preference values ​​for sub-channels, where the preference values ​​are the gradient approximation of the globally weighted security utility with respect to the allocation variables; The user constructs a preference list in descending order based on the preference values; A delayed acceptance mechanism is used to perform matching, so that each unmatched user sends a matching request to the sub-channel ranked first in the preference list in turn; Each sub-channel collects all users who have sent matching requests, selects the user with the highest preference value, temporarily holds the matching request, and rejects other users; The rejected user continues to send a request to the next sub-channel in the preference list until all users have completed the matching or the preference list is exhausted, at which point a stable and secure sub-channel allocation matching result is obtained; the secure sub-channel allocation matching result is used to actively avoid eavesdropping interference while avoiding sub-channel collisions and mitigating inter-beam interference.

[0008] Preferably, the method for performing power allocation includes: A continuous convex approximation method is used to handle the non-convexity introduced by coupling interference in the signal-to-interference-plus-noise ratio expression. In each iteration, a concave lower bound surrogate function is constructed for the legal rate. The concave lower bound surrogate function adopts an adaptive coefficient form, wherein the adaptive coefficient is calculated based on the signal-to-interference-plus-noise ratio value of the current iteration point. A convex upper bound surrogate function is constructed using a first-order Taylor expansion to assess the leakage rate; In each iteration, a convex optimization problem consisting of the concave lower bound surrogate function and the convex upper bound surrogate function is solved, and the power vector is updated until the power vector converges; the goal of the power allocation is to achieve a balance between power flux density constraints, service quality requirements and energy consumption penalty factors.

[0009] Preferred methods for constructing upper-layer intelligent agent models include: A local threat perception state space is constructed based on the collected network status data. The state space includes the local maximum eavesdropping threat probability, the global average eavesdropping probability, historical security efficiency statistics, and the ratio of security rate to legality rate. Construct a policy network with the state space as input and energy penalty factor and fairness sensitivity factor as output; Design a risk-weighted reward function that includes a logarithmic safety and energy efficiency benefit term, a local threat-weighted information leakage penalty term, and a Jain fairness index reward term; Based on the policy network and the risk-weighted reward function, the upper-layer agent model is constructed.

[0010] Preferably, the optimization objective of iteratively optimizing the upper-layer intelligent agent model is the global weighted security utility, which is the weighted utility that integrates fairness and energy consumption constraints to maximize the average security energy efficiency over all time periods.

[0011] This invention also discloses a two-layer scheduling system for satellite communication resources that balances long-term fairness with security and energy efficiency, used to implement the method, comprising: The system model construction module is used to establish a downlink transmission system model of a multi-beam satellite network and collect network status data based on the downlink transmission system model of the multi-beam satellite network. The network status data includes user channel status information, eavesdropping threat probability map, security efficiency statistics and system physical constraints. The solver construction module is used to pre-model the lower-level optimization solver. The lower-level optimization solver adopts a dynamic asymmetric Nash bargaining game model and sequentially performs secure beam pointing matching, secure sub-channel allocation and power allocation to generate an instantaneous resource scheduling scheme. The intelligent agent construction module is used to construct an upper-layer intelligent agent model. The upper-layer intelligent agent model generates strategic hyperparameters based on the network state data. The strategic hyperparameters include an energy penalty factor and a fairness sensitivity factor. The agent training module is used to perform instantaneous resource orchestration on the strategic hyperparameters using the lower-level optimization solver to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, thereby obtaining a trained upper-level agent model. The strategic hyperparameter inference module is used to deploy the trained upper-layer intelligent agent model on the satellite, receive real-time network status data, and infer the strategic hyperparameters of the current time slot. The execution module is used to guide the lower-level optimization solver to generate the final resource scheduling scheme using the strategic hyperparameters of the current time slot, and to distribute the instantaneous resource scheduling scheme to the multi-beam satellite for execution.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes an integrated physical modeling method for multi-beam satellite channels, interference, eavesdropping, and energy efficiency. Through complete mathematical formulas, it achieves accurate quantitative characterization of energy consumption and security-related parameters, providing a foundation for the coordinated optimization of energy consumption and fairness.

[0013] 2. A two-layer collaborative resource orchestration method based on deep reinforcement learning and dynamic game theory is proposed. This method learns and outputs strategic hyperparameters (energy consumption penalty factors) through the upper-layer agent. and fairness weighting factor The resource allocation process at the lower level is dynamically controlled based on a dynamic asymmetric Nash bargaining game. Specifically, Dynamically balancing safety and energy efficiency gains with power penalties. This method achieves a dynamic and coordinated balance among multiple objectives—energy consumption constraints, safety and efficiency, and user fairness—by decoupling upper-level strategic decision-making from lower-level execution optimization. Then adjust the user's bargaining weight in real time. This approach tilts resource allocation towards users with historically poor performance or those in high-risk eavesdropping areas, thus mathematically ensuring long-term proportional fairness. The entire process is supported by detailed mathematical models and formulas, resulting in a practical and protectable technical solution. Attached Figure Description

[0014] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 This is a system scenario diagram of a multi-beam satellite network flying over a dynamic threat area according to an embodiment of the present invention; Figure 3 This is a diagram of the two-layer optimization framework of an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] Example 1: like Figure 1 As shown, a two-tiered scheduling method for satellite communication resources that balances long-term fairness with security and energy efficiency includes: S1: Establish a downlink transmission system model for a multi-beam satellite network and collect network status data based on this model. This data includes user channel status information, eavesdropping threat probability maps, security efficiency statistics, and system physical constraints. In a system scenario where the multi-beam satellite network flies over a dynamic threat area, legitimate users are distributed throughout the area, while potential eavesdroppers are distributed according to the threat probability map. A further implementation involves establishing the downlink transmission system model based on the following conditions: known satellite altitude, orbital inclination, communication frequency band, and the distribution pattern of the space threat probability map; legitimate ground users are randomly distributed within the satellite beam coverage area, while potential eavesdroppers are randomly distributed in space, with both user and eavesdropper positions being random variables and their distribution ranges known; the downlink transmission system model includes M satellites, Q point beams, and N legitimate ground users. Figure 2 As shown.

[0019] In this embodiment, the satellite transmitting antenna gain is considered. User receive gain The fixed value range, the calculation model of free space loss and atmospheric attenuation, and the modeling method of antenna sidelobe gain are defined. The channel gain of beam q and user n is defined. The formula is calculated by combining transmit antenna gain, receive gain and path loss: In the formula, the path loss calculation formula is: In the formula, the atmospheric attenuation term Based on exponential modeling of cloud and rain attenuation coefficients, cross-beam interference characteristics are accurately characterized. Based on the binary variable definition rules of user-beam association and sub-channel allocation, on-board transmit power... The upper limit of the value of Gaussian noise power spectral density Sub-channel bandwidth Based on the fixed parameters and superposition rules of cross-beam aggregation interference, combined with the correlation caused by the random distribution of users and the randomness of allocation requirements, the formula for calculating the signal-to-interference-plus-noise ratio (SINR) of legitimate users is derived: In the formula, This represents the cross-beam aggregation interference experienced by user n on subchannel k. This is combined with a spatial threat probability map. The specific representation methods, the construction logic of the worst-case eavesdropping model, and the basic calculation methods of the legal rate and eavesdropping rate are explained. The eavesdropper's signal-to-interference-plus-noise ratio and eavesdropping rate are derived. The user's achievable safe rate is defined as the non-negative difference between the legal rate and the probability-weighted leakage rate, as shown in the formula: In the formula Where x is a placeholder, referring to The overall difference expression. It represents the maximum value between the difference and 0. For user n, the legal rate on beam q and subchannel k. Let q be the eavesdropping rate of the eavesdropper on the beam and sub-channel k. This represents the probability of eavesdropping threats within the coverage area of ​​beam q. Based on power amplifier efficiency. Fixed circuit power consumption With fixed values, define the single-user SEE formula: In the formula, For power amplifier efficiency, Let n be the instantaneous transmit power. To ensure fixed circuit power consumption, network state data was collected, including the following specific components and computational boundaries: (1) User channel state information: The specific data consists of channel gain data between the legitimate user and each beam, which is the basic input for calculating the signal-to-interference-plus-noise ratio of the legitimate user; (2) Eavesdropping threat probability map: Specifically, it consists of a set of spatial probability distributions constructed based on historical monitoring data, which is used to allocate space candidate locations within the service area. The probability value of eavesdropping occurring within the range; (3) SEE statistic; (4) System physical constraints: The specific calculation methods and boundary conditions include: the constraint that the sum of the power of all beams is less than or equal to the upper limit of the maximum transmission power of the satellite, the upper and lower limits of the power of a single beam, the upper limit of the power flux density mask triggered when the probability of eavesdropping in the area is higher than a specified threshold, and the spatial frequency reuse constraint based on the minimum security isolation distance between beams.

[0020] Based on the aforementioned network state data, a training dataset containing positive physical mapping relationships is constructed to provide comprehensive data support for subsequent two-layer resource orchestration optimization.

[0021] The long-term proportional fairness optimization problem is modeled as a dynamic asymmetric Nash bargaining game model, and the long-term objective is decoupled into instantaneous optimization subproblems in each time slot based on stochastic network optimization theory.

[0022] S2: Pre-modeled lower-level optimization solver. The lower-level optimization solver adopts a dynamic asymmetric Nash bargaining game model, sequentially performing secure beam pointing matching, secure sub-channel allocation, and power allocation to generate an instantaneous resource scheduling scheme. This solver uses the strategic hyperparameters (energy consumption penalty coefficient) output by the upper-level agent. And fairness sensitivity factors Guided by this principle, a generalized asymmetric Nash bargaining product, incorporating a logarithmic proxy utility function and dynamic bargaining weights, is constructed as the instantaneous optimization objective. The solver includes a spatially constrained extended matching algorithm (Algorithm 1), a gradient-based delayed acceptance matching algorithm (Algorithm 2), and a continuous convex approximation algorithm (Algorithm 3). Through systematic mathematical derivation and modeling of the solver, it is ensured that it meets hard physical constraints such as on-board power and spectral isolation, while also guaranteeing the convergence of the algorithms. The solver's execution flow is decomposed into the following three security-aware sub-problems: Subproblem 1 (SP1): Methods for performing secure beam pointing matching include: A binary conflict matrix is ​​constructed to represent the spatial isolation constraint between any two candidate beam centers. If the distance between two candidate beam centers is less than a preset minimum beam spacing threshold, or if they are the same beam center, the corresponding element in the binary conflict matrix is ​​set to 1; otherwise, it is set to 0. Specifically, to maximize secure coverage and suppress information leakage, a single-beam dimension local pre-evaluation is performed on the overall system objective function to calculate the value of each candidate beam center. Potential safety utility score This score represents the expected weighted security gain that would result from activating the center, and its calculation formula is as follows: In the formula, For the candidate beam center The set of legitimate users within the coverage area; The dynamic fairness bargaining weight is output in real time by the upper-level intelligent agent; For this beam to users The potential legal transmission rate; The probability of spatial eavesdropping threat at this center; Let represent the potential worst-case eavesdropping rate at this center. After obtaining the potential security utility scores of all candidate beam centers, a greedy strategy is first used to select conflict-free beam centers as the initial active beam set, based on the scores and under the constraint of the binary conflict matrix. Subsequently, the remaining unselected high-potential centers are constructed into a dynamic candidate pool. Candidate beam centers are sequentially selected from the dynamic candidate pool to attempt to replace existing beam centers in the current active beam set. The replacement operation is performed only when the candidate beam center to be replaced has no conflict with all other beam centers in the current active beam set, and the replacement satisfies the preset monotonic increase condition of the global weighted security utility, until the algorithm converges. The security beam pointing matching result is ultimately used to maximize the weighted security coverage while minimizing information leakage to the eavesdropper.

[0023] Specifically, to satisfy beam space isolation constraints, a binary conflict matrix is ​​constructed. ,in This represents the set of candidate beam centers. The matrix elements are defined as follows: In the formula, as candidate center and distance, The minimum required beam spacing is defined. During algorithm initialization, conflict-free beam centers are greedily selected based on potential security utility scores. In the iterative phase, existing active centers are attempted to be replaced from the dynamic candidate pool, only if the new center is conflict-free with all other active beams and substantially improves the global weighted utility. The replacement is performed until convergence. The goal is to select the beam center that maximizes weighted security coverage and minimizes leakage to eavesdroppers.

[0024] Sub-problem 2 (SP2): Methods for performing secure sub-channel allocation include: Subchannel allocation is modeled as a matching game between the user set and the subchannel set; Define user preference values ​​for sub-channels, where preference values ​​are the gradient approximation of the globally weighted security utility with respect to the allocation variables. Construct a preference list based on users' preference values ​​in descending order. A delayed acceptance mechanism is used to perform matching, causing each unmatched user to sequentially send a matching request to the sub-channel at the top of the preference list. Each sub-channel collects all users who have sent matching requests, temporarily retains the matching request from the user with the highest preference value, and rejects other users. Rejected users continue to send requests to the next sub-channel in the preference list until all users have completed matching or the preference list is exhausted, at which point a stable secure sub-channel allocation matching result is obtained. The secure sub-channel allocation matching result is used to actively avoid eavesdropping interference while avoiding sub-channel collisions and mitigating inter-beam interference.

[0025] Specifically, sub-channel allocation is modeled as a user set. and sub-channel set Matching game between players. To ensure the matching result aligns with the optimization objective... Gradient direction alignment, defining the preference value of user n for sub-channel k as... For allocation variables Gradient approximation: In the formula, the terms in square brackets represent terms in the sub-channel. The instantaneous safety rate gain. User-defined... The preference list is constructed in descending order. A delayed acceptance (DA) mechanism is employed: unmatched users sequentially send requests to the sub-channel at the top of the preference list. The sub-channel temporarily holds the request of the highest bidder and rejects others. Rejected users then request the next preference sub-channel until all users are matched or the list is exhausted, thus obtaining a stable matching result. The goal is to actively avoid eavesdropping interference while avoiding collisions and mitigating inter-beam interference.

[0026] Subproblem 3 (SP3): The method for performing power allocation includes: using a continuous convex approximation method to handle the non-convexity introduced by coupling interference in the signal-to-interference-plus-noise ratio (SIR) expression; in each iteration, constructing a concave lower bound surrogate function for the legal rate, which adopts an adaptive coefficient form, where the adaptive coefficients are calculated based on the SIR value at the current iteration point; constructing a convex upper bound surrogate function for the leakage rate using a first-order Taylor expansion; in each iteration, solving the convex optimization problem composed of the concave lower bound surrogate function and the convex upper bound surrogate function, updating the power vector until the power vector converges; the goal of power allocation is to achieve a balance between power flux density constraints, service quality requirements, and energy consumption penalty factors.

[0027] Specifically, the Continuous Convex Approximation (SCA) method is used to handle the non-convexity introduced by coupling interference in the SINR expression. During iteration, the legal rate is... Constructing a concave lower bound proxy function The adaptive coefficient , Regarding the leakage rate A convex upper bound is constructed using a first-order Taylor expansion. In each iteration, the convex optimization problem constituted by the aforementioned surrogate function is solved, and the power vector P is updated until the iteration converges. The objective is to address constraints related to power flux density (PFD), quality of service (QoS) requirements, and energy consumption penalty factors. A balance is achieved between them.

[0028] Combining the system model and dataset constructed above, the solver takes hyperparameters from the upper layer as input and outputs instantaneous beam steering, subchannel allocation, and power control results. It also integrates engineering constraints such as total power, beam power, power shielding in high-threat areas, minimum communication rate guarantee for users in high-eavesdropping areas, subchannel orthogonality, and beam spatial isolation, providing high-quality feedback support for the optimization of the upper-layer agent.

[0029] S3: Construct an upper-layer intelligent agent model. The upper-layer intelligent agent model generates strategic hyperparameters based on network state data. The strategic hyperparameters include an energy penalty factor and a fairness sensitivity factor. A further implementation method for constructing the upper-layer intelligent agent model includes: constructing a local threat perception state space based on collected network state data. The state space includes the local maximum eavesdropping threat probability, the global average eavesdropping probability, historical security efficiency statistics, and the ratio of security rate to legality rate; constructing a policy network with the state space as input and the energy penalty factor and fairness sensitivity factor as output; designing a risk-weighted reward function that includes a logarithmic security efficiency benefit term, a local threat-weighted information leakage penalty term, and a Jain fairness index reward term; and completing the construction of the upper-layer intelligent agent model based on the policy network and the risk-weighted reward function.

[0030] Specifically, the core function of the model is clearly defined as dynamically adjusting hyperparameters, such as... Figure 3 As shown, the upper-level policyr learns a meta-policy with historical awareness. , This is used to guide the optimization of lower-level generalized Nash negotiations. Among them... This represents the energy consumption penalty factor, used to constrain the total power consumption on the satellite and achieve precise energy consumption control. Used to regulate fairness weights The fairness weight is used to ensure fair communication between users in high-risk eavesdropping areas and those on the margins. The weight value adaptively adjusts based on a user's historical communication performance and the level of eavesdropping threat. The reward function for the agent is designed to incorporate this fairness weight. Energy consumption penalty factor , and optimization objectives In response, taking into account fairness, utility gain, eavesdropping risks, and energy consumption control, the formula is: The collected network state dataset is divided into training and test sets, which are then input into the agent model. A pre-modeled lower-level optimization solver is used to apply an energy consumption penalty factor to the agent model's output. With fairness weighting factor Instantaneous resource scheduling is performed, with long-term safety, energy efficiency, and fairness indicators used as the reward function. The agent model is then iteratively optimized until the model policy converges and reaches a set threshold. These are the power, sub-channel, and beam assignment vectors, respectively. For fairness weights with wavy lines, For users to achieve a safe speed, For user transmission power; in the formula For single-user safety and energy efficiency, For the rate of user information leakage, This is a fairness-based reward. Simultaneously, a dynamic adaptive bargaining weight model based on user eavesdropping threat level and historical communication performance is introduced. This bargaining model adopts a dynamic asymmetric Nash bargaining framework, as shown in Table 1. This framework adjusts resource allocation priorities, tilting resources towards users in high-eavesdropping-risk areas to ensure their communication needs and achieve a synergistic balance between energy consumption and fairness.

[0031] Table 1 S4: Utilize the lower-level optimization solver to perform instantaneous resource orchestration on the strategic hyperparameters to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, obtaining a well-trained upper-level agent model. A further implementation method involves using a globally weighted security utility as the optimization objective for iteratively optimizing the upper-level agent model. The globally weighted security utility is a weighted utility that integrates fairness and energy consumption constraints, maximizing the average security energy efficiency across all time periods. The formula is: S5: Deploy the trained upper-layer intelligent agent model on the satellite, receive real-time network status data, and infer the strategic hyperparameters of the current time slot; S6: Use the strategic hyperparameters of the current time slot to guide the lower-level optimization solver to generate the final resource scheduling scheme, and then send the instantaneous resource scheduling scheme to the multi-beam satellite for execution.

[0032] Specifically, real-time local threat perception information and historical state data are input into the trained intelligent agent model. Its output hyperparameters are then fed into the lower-level optimization solver, which outputs the actual beam, sub-channel, and power allocation results. These allocation results are substituted into the system model to calculate the actual system security efficiency, the weakest user security rate improvement ratio, and the information leakage suppression level. These results are then compared and verified with preset communication requirements to confirm that the allocation results meet the core objectives of low energy consumption, high security, and strong fairness. Simultaneously, the effectiveness of dynamic bargaining weight modeling is verified, ensuring that users in high-risk eavesdropping areas receive sufficient communication resources. Finally, the entire modeling and optimization process is completed, verifying the feasibility and engineering implementation of the solution.

[0033] Example 2: This invention also discloses a two-layer scheduling system for satellite communication resources that balances long-term fairness with security and energy efficiency, for implementing the method of Embodiment 1, comprising: The system model building module is used to establish a downlink transmission system model for a multi-beam satellite network and collect network status data based on the downlink transmission system model. The network status data includes user channel status information, eavesdropping threat probability map, security efficiency statistics, and system physical constraints. The solver construction module is used to pre-model the lower-level optimization solver. The lower-level optimization solver adopts a dynamic asymmetric Nash bargaining game model and sequentially performs secure beam pointing matching, secure sub-channel allocation and power allocation to generate an instantaneous resource scheduling scheme. The agent construction module is used to build the upper-layer agent model. The upper-layer agent model generates strategic hyperparameters based on network state data. The strategic hyperparameters include energy penalty factors and fairness sensitivity factors. The agent training module is used to perform instantaneous resource orchestration of strategic hyperparameters using the lower-level optimization solver to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, thus obtaining a trained upper-level agent model. The strategic hyperparameter inference module is used to deploy the trained upper-layer intelligent agent model on the satellite, receive real-time network status data, and infer the strategic hyperparameters of the current time slot. The execution module is used to guide the lower-level optimization solver to generate the final resource scheduling scheme using the strategic hyperparameters of the current time slot, and to distribute the instantaneous resource scheduling scheme to the multi-beam satellite for execution.

[0034] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A two-layer scheduling method for satellite communication resources that balances long-term fairness with security and energy efficiency, characterized in that, include: A downlink transmission system model for a multi-beam satellite network is established, and network status data is collected based on the model. The network status data includes user channel status information, eavesdropping threat probability map, security efficiency statistics, and system physical constraints. A pre-modeled lower-level optimization solver is used, which adopts a dynamic asymmetric Nash bargaining game model and sequentially performs secure beam pointing matching, secure sub-channel allocation and power allocation to generate an instantaneous resource scheduling scheme. A higher-level intelligent agent model is constructed, which generates strategic hyperparameters based on the network state data. The strategic hyperparameters include an energy penalty factor and a fairness sensitivity factor. The lower-level optimization solver is used to perform instantaneous resource orchestration on the strategic hyperparameters to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, thus obtaining a trained upper-level agent model. The trained upper-layer intelligent agent model is deployed on the satellite to receive real-time network status data and infer the strategic hyperparameters of the current time slot. The strategic hyperparameters of the current time slot are used to guide the lower-level optimization solver to generate the final resource scheduling scheme, and the instantaneous resource scheduling scheme is then sent to the multi-beam satellite for execution.

2. The method according to claim 1, characterized in that, The conditions for establishing a downlink transmission system model for a multi-beam satellite network include: known satellite operating altitude, orbital inclination, communication frequency band, and the distribution pattern of the space threat probability map; legitimate ground users are randomly distributed in the satellite beam coverage area, and potential eavesdroppers are randomly distributed in space, with the locations of users and eavesdroppers being random variables and their distribution ranges known; The downlink transmission system model of the multi-beam satellite network includes M satellites, Q spot beams, and N legitimate ground users.

3. The method according to claim 1, characterized in that, Methods for performing secure beam pointing matching include: A binary conflict matrix is ​​constructed to characterize the spatial isolation constraint between any two candidate beam centers. If the distance between two candidate beam centers is less than a preset minimum beam spacing threshold, or if they are the same beam center, the element at the corresponding position in the binary conflict matrix is ​​set to 1; otherwise, it is set to 0. Based on the potential security utility score of each candidate beam center, and under the premise of satisfying the binary conflict matrix constraint, a greedy strategy is adopted to select conflict-free beam centers as the initial active beam set. Candidate beam centers are selected sequentially from the dynamic candidate pool to replace existing beam centers in the current active beam set. The replacement operation is performed only if the candidate beam center to be replaced does not conflict with any other beam centers in the current active beam set, and the replacement satisfies the preset global weighted security utility enhancement condition, until convergence. The security beam pointing matching result is used to maximize the weighted security coverage while minimizing information leakage to eavesdroppers.

4. The method according to claim 1, characterized in that, Methods for performing secure subchannel allocation include: Subchannel allocation is modeled as a matching game between the user set and the subchannel set; Define user preference values ​​for sub-channels, where the preference values ​​are the gradient approximation of the globally weighted security utility with respect to the allocation variables; The user constructs a preference list in descending order based on the preference values; A delayed acceptance mechanism is used to perform matching, so that each unmatched user sends a matching request to the sub-channel ranked first in the preference list in turn; Each sub-channel collects all users who have sent matching requests, selects the user with the highest preference value, temporarily holds the matching request, and rejects other users; The rejected user continues to send a request to the next sub-channel in the preference list until all users have completed the matching or the preference list is exhausted, at which point a stable and secure sub-channel allocation matching result is obtained; the secure sub-channel allocation matching result is used to actively avoid eavesdropping interference while avoiding sub-channel collisions and mitigating inter-beam interference.

5. The method according to claim 1, characterized in that, Methods for performing power allocation include: A continuous convex approximation method is used to handle the non-convexity introduced by coupling interference in the signal-to-interference-plus-noise ratio expression. In each iteration, a concave lower bound surrogate function is constructed for the legal rate. The concave lower bound surrogate function adopts an adaptive coefficient form, wherein the adaptive coefficient is calculated based on the signal-to-interference-plus-noise ratio value of the current iteration point. A convex upper bound surrogate function is constructed using a first-order Taylor expansion to assess the leakage rate; In each iteration, a convex optimization problem consisting of the concave lower bound surrogate function and the convex upper bound surrogate function is solved, and the power vector is updated until the power vector converges; the goal of the power allocation is to achieve a balance between power flux density constraints, service quality requirements and energy consumption penalty factors.

6. The method according to claim 1, characterized in that, Methods for constructing upper-layer intelligent agent models include: A local threat perception state space is constructed based on the collected network status data. The state space includes the local maximum eavesdropping threat probability, the global average eavesdropping probability, historical security efficiency statistics, and the ratio of security rate to legality rate. Construct a policy network with the state space as input and energy penalty factor and fairness sensitivity factor as output; Design a risk-weighted reward function that includes a logarithmic safety and energy efficiency benefit term, a local threat-weighted information leakage penalty term, and a Jain fairness index reward term; Based on the policy network and the risk-weighted reward function, the upper-layer agent model is constructed.

7. The method according to claim 1, characterized in that, The optimization objective of iteratively optimizing the upper-layer intelligent agent model is the global weighted security utility, which is the weighted utility that integrates fairness and energy consumption constraints to maximize the average security energy efficiency over all time periods.

8. A two-layer scheduling system for satellite communication resources that balances long-term fairness with security and energy efficiency, used to implement the method described in any one of claims 1-7, characterized in that, include: The system model construction module is used to establish a downlink transmission system model of a multi-beam satellite network and collect network status data based on the downlink transmission system model of the multi-beam satellite network. The network status data includes user channel status information, eavesdropping threat probability map, security efficiency statistics and system physical constraints. The solver construction module is used to pre-model the lower-level optimization solver. The lower-level optimization solver adopts a dynamic asymmetric Nash bargaining game model and sequentially performs secure beam pointing matching, secure sub-channel allocation and power allocation to generate an instantaneous resource scheduling scheme. The intelligent agent construction module is used to construct an upper-layer intelligent agent model. The upper-layer intelligent agent model generates strategic hyperparameters based on the network state data. The strategic hyperparameters include an energy penalty factor and a fairness sensitivity factor. The agent training module is used to perform instantaneous resource orchestration on the strategic hyperparameters using the lower-level optimization solver to maximize long-term cumulative rewards and iteratively optimize the upper-level agent model until the model policy converges, thereby obtaining a trained upper-level agent model. The strategic hyperparameter inference module is used to deploy the trained upper-layer intelligent agent model on the satellite, receive real-time network status data, and infer the strategic hyperparameters of the current time slot. The execution module is used to guide the lower-level optimization solver to generate the final resource scheduling scheme using the strategic hyperparameters of the current time slot, and to distribute the instantaneous resource scheduling scheme to the multi-beam satellite for execution.