A broadband satellite communication verification system and method
By employing a distributed asynchronous Q-learning and dynamic coordination mechanism, the latency jitter problem caused by multi-user resource contention in the broadband satellite communication verification system was resolved, achieving rapid response and improved stability, and adapting to resource allocation optimization in various scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AEROSPACE FUTURE (SHENZHEN) AEROSPACE TECHNOLOGY CO LTD
- Filing Date
- 2025-10-20
- Publication Date
- 2026-04-24
AI Technical Summary
Existing broadband satellite communication verification systems suffer from latency jitter due to multi-user resource contention in space-ground integrated networks. Existing solutions suffer from decision delays due to high algorithm complexity, affecting real-time service quality.
A distributed asynchronous Q-learning algorithm is adopted, which independently generates resource allocation decisions through local decision agents and generates a globally consistent resource allocation scheme through a coordination module. Combined with a penalty function to suppress conflicts and a topology change metric to dynamically adjust the update cycle, a fast response is achieved.
It significantly improves the response efficiency and stability of the broadband satellite communication verification system, reduces decision-making delay, enhances the fairness of resource allocation and system scalability, and adapts to various scenarios such as UAV communication and edge computing.
Smart Images

Figure CN121001205B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of broadband satellite communication technology, and in particular to a broadband satellite communication verification system and method. Background Technology
[0002] Broadband satellite communication verification systems play an important role in the development of space-ground converged networks. With the advancement of 6G technology, the integration of low-Earth orbit satellites and terrestrial networks has become a focus. This integration aims to achieve global coverage and support diverse user access, including airborne mobile devices and ground terminals. Verification systems need to test performance in multi-user scenarios to ensure communication reliability. The application background involves urban, rural, and mobile environments, among which resource allocation mechanisms are a key testing point.
[0003] In space-ground converged networks, the problem of multi-user resource contention is becoming increasingly apparent. When airborne users, such as drones, share base station resources with ground users, the resource allocation algorithm faces challenges due to the dynamic changes in the number of users. Existing verification systems use joint optimization methods to handle resource allocation, but such methods may cause latency fluctuations in scenarios with dense user populations or frequent movement. Latency jitter affects real-time applications such as video transmission, leading to unstable service quality. The root cause of the problem is that the algorithm's response speed is insufficient to adapt to rapidly changing network conditions.
[0004] Most existing solutions optimize resource allocation through intelligent algorithms. For example, some verification systems introduce deep reinforcement learning frameworks to model the resource allocation problem as an optimization task and learn the best policy by training agents. These methods aim to balance the needs of different users and reduce conflicts. However, the algorithms are complex and the decision-making process may be prolonged during sudden traffic surges or user spikes, exacerbating latency issues. Other methods, such as distributed coordination mechanisms, have also been explored, but they rely on global information synchronization and face bottlenecks when scaling up. Summary of the Invention
[0005] In view of the aforementioned existing problems, the present invention is proposed.
[0006] This invention provides a broadband satellite communication verification system and method to solve the problem of latency jitter caused by multi-user resource contention in satellite-ground integrated networks, and the problem that existing solutions suffer from decision delays due to algorithm complexity, which affects the quality of real-time services.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] In a first aspect, embodiments of the present invention provide a broadband satellite communication verification system, comprising:
[0009] Multiple user node simulation units are used to simulate user equipment in a space-ground converged network;
[0010] A resource management unit is communicatively connected to the plurality of user node simulation units;
[0011] The user node simulation unit includes a local decision agent, which is configured to independently generate resource allocation decisions for the node based on an asynchronous Q-learning algorithm. The resource management unit is configured to coordinate decisions from multiple user node simulation units and generate a globally consistent resource allocation scheme.
[0012] In a preferred embodiment of the broadband satellite communication verification system described in this invention, the local decision agent includes a state awareness module, which is configured to collect channel quality information and user demand information of the local node in real time and use them as decision inputs.
[0013] As a preferred embodiment of the broadband satellite communication verification system of the present invention, the state awareness module is further configured to receive partial state information from at least one other user node simulation unit, and the local decision agent is configured to suppress conflicts caused by multiple nodes competing for the same resource simultaneously based on a penalty function.
[0014] The steps for suppressing conflicts caused by multiple nodes simultaneously competing for the same resource include:
[0015] a) Within each user node simulation unit, the state awareness module collects and normalizes the data through a sliding window to obtain five cost elements: conflict rate, interference, queuing delay cost, jitter cost, and signaling overhead cost; the normalization interval is uniformly mapped to [0,1].
[0016] b) To suppress high contention, high interference, and high overhead actions at the node side, at time... Construct a comprehensive penalty:
[0017] ,
[0018] in, Indicates time Comprehensive punishment, Indicates time Conflict rate Indicates time Interference level, Indicates time The cost of queuing delay Indicates time The cost of shaking Indicates time Signaling overhead cost, These are the non-negative weights of the corresponding elements, and the sum of the weights is 1, which is used to reflect the emphasis on the business side; this penalty is used as an endogenous inhibition of local action selection in the strategy evaluation.
[0019] c) Combine business utility with comprehensive penalties to generate immediate rewards for value updates and action optimization in local asynchronous Q-learning:
[0020] ,
[0021] in, Indicates time Instant rewards Indicates time The business utility is obtained by normalizing either throughput satisfaction or latency satisfaction, and is dimensionless. Let be the penalty coefficient, non-negative, which modulates the sensitivity to the penalty; based on The value update and strategy improvement of asynchronous Q-learning are carried out independently on each node. The resource management unit only collects statistics for network-side coordination and does not block local updates.
[0022] d) When increased competition for local hotspot resources is observed, the coordination module can issue a congestion indication within a short time window, and the nodes can adjust their settings slightly upwards according to the indication. or corresponding The system is designed to accelerate the suppression of hotspot behaviors, and then returns to the normal configuration at a preset annealing step size after the competition subsides.
[0023] As a preferred embodiment of the broadband satellite communication verification system of the present invention, the resource management unit includes a coordination module, which is configured to receive resource allocation decisions from each user node simulation unit asynchronously, and generate a globally consistent resource allocation scheme without blocking local decision updates of each node.
[0024] As a preferred embodiment of the broadband satellite communication verification system of the present invention, the coordination module is configured to dynamically adjust the update cycle of the resource allocation scheme based on network topology change measurement, so as to shorten the update cycle when the topology changes rapidly and extend the update cycle when the topology is relatively stable.
[0025] The method for dynamically adjusting the update cycle of the resource allocation scheme is as follows:
[0026] The resource management unit obtains and normalizes three types of elements from the network simulation unit and reports from each node: satellite / beam switching rate. rate of change of adjacent sets Link hold time cost All three are mapped to a sliding window and anti-mutation pruning. ;
[0027] To obtain a single schedulable signal over time, a comprehensive topology change metric is defined:
[0028] ,
[0029] in, Indicates time A comprehensive measure of topological change. Indicates time Satellite / beam switching rate, Indicates time The rate of change of the adjacent set, Indicates time The cost of link hold-up time, For the corresponding weights, non-negative and This is used to reflect the emphasis different business slices place on the three types of changes;
[0030] To shorten the update cycle during topological changes and lengthen it during stable conditions, a smooth logarithmic mapping is introduced:
[0031] ,
[0032] in, Indicates time The resource allocation scheme update cycle, Indicates the shortest allowed update cycle. Indicates the longest allowed update cycle. This is the kurtosis coefficient of the curve, which is non-negative. The inflection point of the change measure is located at... The mapping is in When it rises Towards Convergence, in When reduced Towards near;
[0033] Set three types of triggers:
[0034] 1. Triggering, when Exceeding the ascent threshold At that time, an early update will be triggered immediately and during the cooldown period. It will not be triggered again within the specified time.
[0035] II. Component triggering, when Exceeding the threshold or Exceeding the threshold Perform the same action at the same time;
[0036] III. Stable triggering, when continuous The number of windows is below the threshold. At that time, According to step size Slowly pull towards Simultaneously set hysteresis: rise threshold Downlink threshold Separation.
[0037] As a preferred embodiment of the broadband satellite communication verification system of the present invention, the system further includes a network simulation unit configured to simulate the communication link between a low-Earth orbit satellite constellation and a ground base station, and to provide the resource management unit with link status and switching events.
[0038] Secondly, this invention provides a broadband satellite communication verification method, comprising,
[0039] Step S1: In each user node simulation unit, the local decision agent independently generates resource allocation decisions for the node based on the asynchronous Q-learning algorithm.
[0040] Step S2: In the resource management unit, the resource allocation decisions of multiple user node simulation units are coordinated asynchronously to form a network-oriented resource allocation scheme.
[0041] As a preferred embodiment of the broadband satellite communication verification method of the present invention, step S1 includes: real-time evaluation of the channel state and user priority of the local node, which is used as one of the constituent elements of the state and reward of asynchronous Q-learning.
[0042] As a preferred embodiment of the broadband satellite communication verification method of the present invention, step S1 further includes: exchanging partial state information between adjacent user node simulation units, and adjusting local decisions based on constraints on exchange overhead.
[0043] As a preferred embodiment of the broadband satellite communication verification method of the present invention, step S2 includes: adaptively determining the update cycle of the resource allocation scheme based on the network topology change metric, and triggering an early update when the topology change is detected to exceed a preset threshold.
[0044] The beneficial effects of this invention are as follows: By employing a distributed asynchronous Q-learning framework and a dynamic coordination mechanism, this invention significantly improves the response efficiency and stability of a broadband satellite communication verification system. In the highly dynamic environment of a satellite-ground integrated network, the local decision agent operates independently, rapidly generating resource allocation decisions based on real-time channel status and user needs, avoiding the computational bottleneck caused by centralized optimization. This distributed architecture reduces inter-node communication overhead and decision latency, thereby effectively suppressing jitter caused by multi-user competition. The penalty function mechanism incorporates multi-dimensional costs such as conflict rate and interference into the learning process, enabling nodes to autonomously avoid highly competitive actions and enhancing the fairness of resource allocation. Simultaneously, the coordination module adaptively adjusts the update cycle through topology change metrics, shortening the response interval when satellites move rapidly to ensure seamless link switching, and extending the cycle when the network is stable to save control resources. This method does not rely on complex global information, achieving global coordination only through local interactions, thus improving the system's scalability and robustness.
[0045] In summary, this invention optimizes resource utilization, ensures low-latency service requirements, provides a lightweight verification foundation for 6G integrated terrestrial and space networks, and is suitable for various scenarios such as drone communication and edge computing. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.
[0047] Figure 1 This is a schematic diagram of the framework of the broadband satellite communication verification system in the embodiment.
[0048] Figure 2 This is a flowchart illustrating the broadband satellite communication verification method in the embodiment. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0050] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0051] For example, the terms “first” and “second” used in this application are only used to distinguish and describe similar objects, to differentiate the first object from another object, and are not used to describe a specific order or sequence, nor should they be interpreted as indicating or implying relative importance.
[0052] This application proposes a broadband satellite communication verification system, combined with Figure 1 As shown, it includes:
[0053] Multiple user node simulation units are used to simulate user equipment in a space-ground converged network;
[0054] A resource management unit communicates and connects with multiple user node simulation units;
[0055] The user node simulation unit includes a local decision agent, which is configured to independently generate resource allocation decisions for the node based on an asynchronous Q-learning algorithm. The resource management unit is configured to coordinate decisions from multiple user node simulation units to form a network-oriented resource allocation scheme. In this embodiment, the user node simulation unit is defined as a user-side simulation entity with protocol stack abstraction and a service generator, which exchanges observations and commands with the control plane / data channel of the resource management unit through a simulation interface. Adjacent user node simulation units refer to a set of nodes within the same beam coverage or connected by edges in the interference graph. The default sampling period is 10ms, adjustable from 1ms to 50ms, set according to the service delay level and protocol slot granularity. The minimum duration of control coordination is an integer multiple of one sampling period, with a default of 100ms. When the above determinations are involved, if interference graph information is missing, the received reference signal strength threshold is used instead for adjacent determination. If control plane congestion leads to observation loss, the available decisions from the previous moment are used and updated in the next valid period.
[0056] In one embodiment, the local decision agent includes a state-aware module configured to collect channel quality information and user demand information of the local node in real time and use them as decision inputs. Specifically, the acquisition of channel quality information mainly relies on SINR or CQI obtained from link estimation, and incorporates statistics of reference signal received power and interference grid when necessary. The acquisition of user demand information mainly relies on the target bit rate / delay level and queue length declared by the service initiator. The default quantization resolution is CQI level 15 and delay level 3, and the quantization resolution can be adjusted within the standard range to match the simulation accuracy. The acquisition window defaults to approximately 100 sampling points, and an exponential moving average is used to suppress abrupt changes. When any source information is missing, the last valid value is maintained for no more than 500ms, and a one-time smooth transition is performed after recovery to avoid transient jumps.
[0057] In one embodiment, the state-aware module is further configured to receive partial state information from at least one other user node simulation unit, and the local decision agent is configured to suppress conflicts caused by multiple nodes competing for the same resource simultaneously based on a penalty function, thereby reducing the probability of resource allocation conflicts.
[0058] The steps to suppress conflicts caused by multiple nodes competing for the same resource simultaneously include:
[0059] a) Within each user node simulation unit, the state awareness module collects and normalizes the data through a sliding window to obtain five cost elements: conflict rate (the proportion of the same resource requested concurrently by multiple nodes), interference degree (a dimensionless value obtained by inverse quantization of the interference intensity at the receiving end or SINR), queuing delay cost (normalized deviation relative to the target delay), jitter cost (normalized root mean square of the delay difference between adjacent packets), and signaling overhead cost (normalized proportion of control bits to service bits). The normalization interval is uniformly mapped to [0,1], which facilitates subsequent multiplication and superposition with weights.
[0060] For example, the conflict rate is measured by the percentage of concurrent requests for the same resource from multiple nodes within the sliding window; the interference rate is dequantized from the current SINR using an empirical mapping table to a dimensionless cost; the queuing delay cost is normalized by the positive deviation of the queuing delay from the target delay; the jitter cost is normalized by the root mean square of the end-to-end delay difference between adjacent packets; and the signaling overhead cost is normalized by the proportion of control bits to service bits. The default sliding window length is 100 to 200 sampling points, which can be adjusted between 50 and 500 sampling points depending on the service burstiness. To resist anomalies, each original quantity is pruned by 1% above and below the normalization threshold, which can be slightly adjusted between 0.5% and 2%. If there are insufficient valid samples within the window, the cost will not be updated and the previous normalization result will be inherited.
[0061] b) To suppress high contention, high interference, and high overhead actions at the node side, at time... Construct a comprehensive penalty:
[0062] ,
[0063] in, Indicates time Comprehensive punishment, Indicates time The conflict rate, normalized, dimensionless. Indicates time The interference degree, normalization, dimensionless, Indicates time Queuing delay cost, normalized, dimensionless. Indicates time The jitter cost, normalized, dimensionless. Indicates time Signaling overhead cost, normalized, dimensionless. These are the non-negative weights of the corresponding elements, with the sum of the weights being 1, used to reflect the business-side emphasis; this penalty serves as an endogenous inhibition term for local action selection in strategy evaluation; similarly, the weights... The settings are based on business slice preferences: latency-sensitive slices prioritize increasing the weights related to latency, while throughput-sensitive slices moderately increase the weights related to interference / conflict. The weights are set to around 0.2 by default, allowing adjustment within [0,1] while maintaining a sum of 1. The weight update cycle is synchronized with the node-side learning step size, defaulting to no faster than 100ms to avoid frequent oscillations. When a cost dimension is temporarily unavailable due to sensor loss, the corresponding weights for that dimension are proportionally redistributed to other dimensions until that dimension is restored. To avoid instantaneous over-suppression, a soft limit is set for the comprehensive penalty; the limit boundary is associated with the business level and is uniformly issued by the coordination module.
[0064] c) Combine business utility with comprehensive penalties to generate immediate rewards for value updates and action optimization in local asynchronous Q-learning:
[0065] ,
[0066] in, Indicates time Instant rewards Indicates time The business utility is obtained by normalizing either throughput satisfaction or latency satisfaction, and is dimensionless. Let be the penalty coefficient, non-negative, which modulates the sensitivity to the penalty; based on Asynchronous Q-learning is used for value updates and strategy improvements. This process is performed independently on each node. The resource management unit only collects statistics for network-side coordination and does not block local updates. Optionally, The acquisition follows the service quality satisfaction criteria: throughput satisfaction is obtained by saturation mapping of actual average throughput relative to target throughput, and latency satisfaction is obtained by monotonic mapping of target latency and actual latency. One of the two or the slicing strategy is preset in the implementation. The default value is set to the median range of 0.3 to 0.7, and then tuned within [0,1] according to the scenario. The tuning is based on stabilizing the variance of the return and minimizing the convergence time under representative business mix. To improve robustness under asynchronous updates, a slight smoothing can be performed before the return enters the value update; when any input that the return calculation depends on is not yet complete, the previous valid return is kept to no more than one learning step and marked as a non-update step to avoid erroneous learning.
[0067] d) When increased competition for local hotspot resources is observed, the coordination module can issue a congestion indication within a short time window, and the nodes can adjust their settings slightly upwards according to the indication. or corresponding The set is used to accelerate the suppression of hotspot behaviors, and after the competition is relieved, it falls back to the normal configuration according to the preset annealing step size; the mechanism does not change the state or action dimension, but only makes time-varying fine-tuning of the coefficients in steps b) and c), which is convenient for linkage with asynchronous coordination and topology change measurement.
[0068] Furthermore, the congestion indication is generated primarily based on the rising slope of the collision rate or interference level within a short-term window, with a default comparison window of 20 to 50 sampling points; the coefficient is adjusted upwards by 5% to 15% of the original value, with the duration consistent with the cooling-off time, and the default cooling-off time is [not specified]. The duration is 1 to 3 seconds. To avoid frequent adjustments, the congestion indication employs an uplink and downlink hysteresis strategy, with the uplink threshold being one fixed offset higher than the downlink threshold. If excessive conflicts still occur within two consecutive congestion cycles, the upper limit of the interference-related weights can be temporarily increased to 0.4 without changing the state or action dimension, and then adjusted according to the annealing step size after congestion is relieved. Recovery will be phased in.
[0069] Specifically, using the five observable costs of nodes as primitives, a comprehensive penalty is formed through unified normalization and non-negative weighted aggregation, thereby explicitly suppressing high-competition, high-interference, and high-cost actions during the policy evaluation phase. The penalty is incorporated into the reward in a linear manner, which maintains the interpretability of the reward signal and reduces the dependence of the learning process on the expansion of the state dimension, adapting to the parallel update characteristics of asynchronous Q-learning. The weights and penalty coefficients serve as the external carriers of business preferences: the weights control the relative importance between cost dimensions, and the penalty coefficients control the overall suppression strength. The hierarchical adjustment of the two facilitates configuration migration between different slices or scenarios. The coordination module performs short-cycle fine-tuning of the coefficients through congestion indicators, which can provide external stimuli when the topology changes rapidly and competition increases sharply, shortening the convergence time and reducing the duration of conflict.
[0070] In one embodiment, the resource management unit includes a coordination module, which is configured to receive resource allocation decisions from each user node simulation unit asynchronously and generate a globally consistent resource allocation scheme without blocking local decision updates at each node. In this embodiment, the coordination module performs consistency arbitration after time-aligning the local decisions and statistics reported by each node, and outputs a set of resource-to-node mapping instructions. The default alignment error does not exceed one sampling period; if it does, nearest-neighbor alignment is performed based on the timestamps reported by each node. The coordination execution period is consistent with the current effective update period of the system. When in the early update triggering period mentioned in the method, the coordination module enters a cooling state after completing this arbitration to avoid repeated triggering.
[0071] In one embodiment, the coordination module is configured to dynamically adjust the update cycle of the resource allocation scheme based on network topology change metrics, so as to shorten the update cycle when the topology changes rapidly and extend the update cycle when the topology is relatively stable.
[0072] The method for dynamically adjusting the update cycle of the resource allocation scheme is as follows:
[0073] The resource management unit obtains and normalizes three types of elements from the network simulation unit and reports from each node: satellite / beam switching rate. (The proportion of switching occurrences per unit time to the reference upper limit), Adjacency set change rate (The ratio of the symmetric difference between the sets of adjacent reachable nodes to the reference capacity), link hold-up time cost (A dimensionless quantity obtained by normalization with the inverse of the average retention time; the longer the retention time, the smaller the value); all three are mapped to a sliding window and anti-mutation pruning. ;
[0074] To obtain a single schedulable signal over time, a comprehensive topology change metric is defined:
[0075] ,
[0076] in, Indicates time A comprehensive measure of topological change. Indicates time Satellite / beam switching rate, normalized, dimensionless. Indicates time The rate of change of the adjacent set, normalized, dimensionless. Indicates time Link hold time cost, normalized, dimensionless. For the corresponding weights, non-negative and This is used to reflect the emphasis different business slices place on the three types of changes;
[0077] Specifically, It is obtained by normalizing the number of satellite / beam switching events per unit time; It is obtained by normalizing the symmetric difference size of the set of adjacent reachable nodes relative to the reference capacity; The value is obtained by normalizing the inverse of the average link retention time. The sliding window and anti-mutation pruning for all three parameters follow the aforementioned settings; weights... Deployed by the business slicing strategy and kept non-negative, summed to 1, with default balanced configuration, improving performance in low-latency slices. or Appropriately increase the throughput in high-throughput slices To accelerate the response to frequent switching. When any component is temporarily unavailable, its weight is proportionally redistributed to the remaining components, and a smooth regression is performed in the first period after that component is restored.
[0078] To shorten the update cycle during topological changes and lengthen it during stable conditions, a smooth logarithmic mapping is introduced:
[0079] ,
[0080] in, Indicates time The resource allocation scheme update cycle, Indicates the shortest allowed update cycle. Indicates the longest allowed update cycle. This is the kurtosis coefficient of the curve, which is non-negative. The inflection point of the change measure is located at... The mapping is in When it rises Towards Convergence, in When reduced Towards near;
[0081] For example, Used to limit the shortest coordination update period, with a default value of 50ms to 200ms; Used to limit the maximum coordination update period, with a default value of 500ms to 2000ms; Used to adjust the steepness of the curve, with a default value of 1 to 6; The inflection point for setting the change metric is set, with a default value of 0.4 to 0.6. These values are based on offline simulations at representative orbits and cell densities, resulting in higher sensitivity in moderate change ranges and a saturation buffer zone in extremely stable or drastic change ranges. When the system is under maintenance or batch reconfiguration, the mapping output can be temporarily constrained at... Fixed points within the system are used to avoid frequent scheduling.
[0082] Set three types of triggers:
[0083] 1. Triggering, when Exceeding the ascent threshold At that time, an early update will be triggered immediately and during the cooldown period. It will not be triggered again within the specified time.
[0084] II. Component triggering, when Exceeding the threshold or Exceeding the threshold Perform the same action at the same time;
[0085] III. Stable triggering, when continuous The number of windows is below the threshold. At that time, According to step size Slowly pull towards Simultaneously set hysteresis: rise threshold Downlink threshold Separate to avoid frequent oscillations;
[0086] Similarly, threshold , The default values are 0.7 and 0.3. The default value is 0.6. The default value is 0.5; cooldown time Default 1s; stride The default time is between 20ms and 50ms; the stability window M defaults to 5 to 10 update cycles. These values can be slightly adjusted based on the scenario without changing the trigger type; when the number of consecutive triggers exceeds 3 and... While still in the high range, it is permissible to [do something] during the cooling period. Temporarily increase the discrete setting to improve response sensitivity, and automatically revert after cooling is complete.
[0087] Specifically, the multi-source changes in the topology are compressed into a single metric, and the update cycle is obtained using a monotonically controllable mapping, forming a direct channel from change intensity to scheduling frequency. Replacing the linear relationship with a logarithmic curve can achieve higher sensitivity in the medium change range and maintain saturation in the extremely stable or drastic change range, reducing cycle jitter. Three types of triggering mechanisms provide event-driven fast response and smooth rollback outside the curve: transition triggering is for short-term mutations, component triggering covers extreme values of key dimensions, and stability triggering is used to gradually lengthen the cycle in the low change phase. The combination of hysteresis and cooling time suppresses the control plane burden caused by frequent switching, is compatible with asynchronous coordination frameworks, and does not block the node-side learning process. Weights and thresholds can be configured and distributed according to business slices, realizing cross-scenario migration without changing the state or action space, which is convenient for working in conjunction with the aforementioned penalty function and reward construction.
[0088] In one embodiment, the system further includes a network simulation unit configured to simulate the communication link between the low-Earth orbit satellite constellation and the ground base station, and to provide the resource management unit with link status and switching events to support the coordination module's decision-making.
[0089] Optionally, link status includes, but is not limited to, instantaneous or windowed SINR, reachable adjacency set, and expected handover time window. Handover events include event type, occurrence time, and target beam / satellite identifier. The above information is provided externally at a pace consistent with the sampling period, with a minimum external push period of 50ms allowed by default. When the network simulation unit cannot output all fields on schedule due to resource constraints, at least the availability of handover events and basic SINR is guaranteed. The remaining fields are substituted with the most recent available value and uniformly corrected within the first cycle after output recovery.
[0090] This application also proposes a broadband satellite communication verification method, combined with Figure 2 As shown, the method includes:
[0091] Step S1: In each user node simulation unit, the local decision agent independently generates resource allocation decisions for the node based on the asynchronous Q-learning algorithm.
[0092] Step S2: In the resource management unit, the resource allocation decisions of multiple user node simulation units are coordinated asynchronously to form a network-oriented resource allocation scheme.
[0093] Step S1 includes: evaluating the channel state and user priority of the local node in real time, using this as one of the components of the state and reward of asynchronous Q-learning, thereby obtaining local resource allocation decisions;
[0094] Step S1 further includes: exchanging partial state information between adjacent user node simulation units, and adjusting local decisions based on constraints on exchange overhead, so as to reduce control signaling load and suppress resource allocation conflicts.
[0095] Step S2 includes: adaptively determining the update cycle of the resource allocation scheme based on the network topology change metric, and triggering an early update when the topology change exceeds a preset threshold.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0097] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
Claims
1. A broadband satellite communication verification system, characterized in that, include: Multiple user node simulation units are used to simulate user equipment in a space-ground converged network; A resource management unit is communicatively connected to the plurality of user node simulation units; The user node simulation unit includes a local decision agent, which is configured to independently generate resource allocation decisions for the node based on an asynchronous Q-learning algorithm. The resource management unit is configured to coordinate decisions from multiple user node simulation units and generate a globally consistent resource allocation scheme. The local decision agent includes a state awareness module, which is configured to collect channel quality information and user demand information of the local node in real time and use them as decision input. The resource management unit includes a coordination module, which is configured to receive resource allocation decisions from each user node simulation unit asynchronously, and generate a globally consistent resource allocation scheme without blocking local decision updates at each node. The system also includes a network simulation unit, which is configured to simulate the communication link between the low-Earth orbit satellite constellation and the ground base station, and provide the resource management unit with link status and handover events; The coordination module is configured to dynamically adjust the update cycle of the resource allocation scheme based on network topology change metrics, so as to shorten the update cycle when the topology changes rapidly and extend the update cycle when the topology is relatively stable.
2. The broadband satellite communication verification system as described in claim 1, characterized in that, The state awareness module is also configured to receive partial state information from at least one other user node simulation unit, and the local decision agent is configured to suppress conflicts caused by multiple nodes competing for the same resource simultaneously based on a penalty function. The steps for suppressing conflicts caused by multiple nodes simultaneously competing for the same resource include: a) Within each user node simulation unit, the state awareness module collects and normalizes the data through a sliding window to obtain five cost elements: conflict rate, interference, queuing delay cost, jitter cost, and signaling overhead cost; the normalization interval is uniformly mapped to [0,1]. b) To suppress high contention, high interference, and high overhead actions at the node side, at time... Construct a comprehensive penalty: , in, Indicates time Comprehensive punishment, Indicates time Conflict rate Indicates time Interference level, Indicates time The cost of queuing delay Indicates time The cost of shaking Indicates time Signaling overhead cost, These are the non-negative weights of the corresponding elements, and the sum of the weights is 1, which is used to reflect the emphasis on the business side; this penalty is used as an endogenous inhibition of local action selection in the strategy evaluation. c) Combine business utility with comprehensive penalties to generate immediate rewards for value updates and action optimization in local asynchronous Q-learning: , in, Indicates time Instant rewards Indicates time The business utility is obtained by normalizing either throughput satisfaction or latency satisfaction, and is dimensionless. Let be the penalty coefficient, non-negative, which modulates the sensitivity to the penalty; based on The value update and strategy improvement of asynchronous Q-learning are carried out independently on each node. The resource management unit only collects statistics for network-side coordination and does not block local updates. d) When increased competition for local hotspot resources is observed, the coordination module can issue a congestion indication within a short time window, and the nodes can adjust their settings slightly upwards according to the indication. or corresponding The system is designed to accelerate the suppression of hotspot behaviors, and then returns to the normal configuration at a preset annealing step size after the competition subsides.
3. The broadband satellite communication verification system as described in claim 2, characterized in that, The method for dynamically adjusting the update cycle of the resource allocation scheme is as follows: The resource management unit obtains and normalizes three types of elements from the network simulation unit and reports from each node: satellite / beam switching rate. rate of change of adjacent sets Link hold time cost All three are mapped to a sliding window and anti-mutation pruning. ; To obtain a single schedulable signal over time, a comprehensive topology change metric is defined: , in, Indicates time A comprehensive measure of topological change. Indicates time Satellite / beam switching rate, Indicates time The rate of change of the adjacent set, Indicates time The cost of link hold-up time, For the corresponding weights, non-negative and This is used to reflect the emphasis different business slices place on the three types of changes; To shorten the update cycle during topological changes and lengthen it during stable conditions, a smooth logarithmic mapping is introduced: , in, Indicates time The resource allocation scheme update cycle, Indicates the shortest allowed update cycle. Indicates the longest allowed update cycle. This is the kurtosis coefficient of the curve, which is non-negative. The inflection point of the change measure is located at... The mapping is in When it rises Towards Convergence, in When reduced Towards near; Set three types of triggers:
1. Triggering, when Exceeding the ascent threshold At that time, an early update will be triggered immediately and during the cooldown period. It will not be triggered again within the specified time. II. Component triggering, when Exceeding the threshold or Exceeding the threshold Perform the same action at the same time; III. Stable triggering, when continuous The number of windows is below the threshold. At that time, According to step size Slowly pull towards Simultaneously set hysteresis: rise threshold Downlink threshold Separation.
4. A broadband satellite communication verification method, based on the broadband satellite communication verification system according to any one of claims 1 to 3, characterized in that, include: Step S1: In each user node simulation unit, the local decision agent independently generates resource allocation decisions for the node based on the asynchronous Q-learning algorithm. Step S2: In the resource management unit, the resource allocation decisions of multiple user node simulation units are coordinated asynchronously to form a network-oriented resource allocation scheme.
5. The broadband satellite communication verification method as described in claim 4, characterized in that, Step S1 includes: real-time evaluation of the channel state and user priority of this node, which serves as one of the components of the state and reward of asynchronous Q-learning.
6. A broadband satellite communication verification method as described in any one of claims 4 and 5, characterized in that, Step S1 further includes: exchanging partial state information between adjacent user node simulation units, and adjusting local decisions based on constraints on exchange overhead.
7. The broadband satellite communication verification method as described in claim 6, characterized in that, Step S2 includes: adaptively determining the update cycle of the resource allocation scheme based on the network topology change metric, and triggering an early update when the topology change exceeds a preset threshold.
Citation Information
Patent Citations
Wireless communication resource allocation method based on federated learning and optimization theory
CN117793928A
Spectrum resource allocation method based on non-ground network and related equipment
CN120601939A