A heterogeneous satellite network access control method based on block semantic-aware parameterized quantum circuits and hierarchical knowledge distillation

CN122577971APending Publication Date: 2026-08-14NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,现有蒸馏方法大多面向通用任务设计,缺乏针对PAoI优化目标的专门约束,难以保证蒸馏后的轻量化模型仍能保持对信息时效性的有效优化能力

Benefits of technology

[0018]本发明方法通过引入分块语义感知参数化量子电路,将系统状态按决策历史、信道拓扑、时效队列三类语义分块独立处理,量子电路的可训练参数量仅随量子比特数线性增长,从根源上解决了经典全连接特征提取器参数量随卫星数量平方级增长的问题,为资源受限的卫星地面网关提供了可行的高维状态特征提取方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122577971A_ABST
    Figure CN122577971A_ABST
Patent Text Reader

Abstract

This invention discloses a heterogeneous satellite network access control method based on Block Semantic Aware Parameterized Quantum Circuit (BS-PQC) and hierarchical knowledge distillation, belonging to the field of satellite communication technology. Addressing the Peak Age of Information (PAoI) optimization problem in GEO / LEO heterogeneous satellite networks, the system state is semantically segmented. BS-PQC is used to extract state features, constructing a Quantum-Enhanced Multi-Head Attention and Post-Decision State Enhanced Double Dueling Deep Q-Network (QMAPS-D3QN) quantum teacher network. Through hierarchical knowledge distillation, the state awareness capability and PAoI decision logic of the teacher network are transferred to a lightweight student network, achieving satellite network access control oriented towards PAoI optimization. The proposed method achieves an average PAoI of 3.27 during the training phase, significantly lower than DQN's 6.31. During the testing phase, the average PAoIs for teacher and student models are 2.69 and 2.80, respectively, and the number of parameters for students is only 29% of that for teachers, achieving a balance between performance and lightweight design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses a heterogeneous satellite network access control method based on block semantic-aware parameterized quantum circuits and hierarchical knowledge distillation, belonging to the field of satellite communication technology. Background Technology

[0002] In applications sensitive to information timeliness, information freshness is paramount. Traditional average latency metrics primarily focus on long-term average performance, failing to reflect the instantaneous peak characteristics of information age within the system. Age of Information (AoI), a key indicator for measuring information freshness, is defined as the time interval from packet generation to successful reception, effectively characterizing the timeliness of information within the system. Peak Age of Information (PAoI), as an important supplementary indicator to AoI, captures the maximum age value before each packet transmission is complete, more intuitively reflecting the worst-case scenario of information freshness. This is of significant importance for critical application scenarios such as emergency response, disaster early warning, and real-time monitoring.

[0003] In recent years, geosynchronous Earth Orbit (GEO) / low Earth Orbit (LEO) heterogeneous satellite networks have been widely studied and applied in data transmission scenarios where information timeliness is critical. GEO satellites offer advantages such as wide coverage and stable links, but suffer from higher propagation latency; LEO satellites offer advantages such as low propagation latency and flexible access, but are prone to frequent handover issues due to high-speed movement. Finding a reasonable choice between these two types of satellites, achieving a balance between high-latency stability and low-latency dynamism, is a key challenge in improving the system's information freshness.

[0004] For the PAoI optimization problem, current mainstream methods are mainly based on deep reinforcement learning techniques. However, with the continuous expansion of heterogeneous satellite networks, traditional deep reinforcement learning methods still have several shortcomings. First, traditional methods mainly focus on the value function approximation and policy learning of a single agent, generally only considering the transition process from the current state to the next state, ignoring the important intermediate state information of the Post-Decision State (PDS), making it difficult to fully utilize the predictable state evolution laws in the system. Second, in multi-gateway scenarios, there is a lack of effective information interaction and cooperation mechanisms. Independent decision-making by multiple gateways often leads to multiple gateways competing for the same satellite resources simultaneously, causing link congestion and system performance degradation. Third, with the increase in the number of satellites and the increase in the dimensionality of the state space, the number of parameters of traditional classical neural network feature extractors grows rapidly, making it difficult to deploy efficiently on resource-constrained satellite ground gateways or edge nodes.

[0005] In recent years, the development of quantum machine learning technology has provided new ideas for solving the above problems. Parameterized Quantum Circuits (PQCs) construct trainable models through parameterized quantum gates, possessing strong high-dimensional nonlinear feature representation capabilities and showing potential advantages in parameter scale control. Especially in heterogeneous satellite network access control scenarios, the system state typically contains multiple heterogeneous information such as decision history, channel topology, and time-sensitive queues. Directly using a unified classical feature extraction method can easily lead to aliasing between different semantic information, affecting feature learning performance. Therefore, how to design more efficient feature extraction methods for different semantic states has become an important issue in improving access control performance.

[0006] Meanwhile, Knowledge Distillation (KD), as an effective model compression method, can transfer knowledge from complex models to lightweight models, thus balancing performance and deployment feasibility. However, most existing distillation methods are designed for general tasks and lack specific constraints for PAoI optimization objectives, making it difficult to guarantee that the distilled lightweight model can still maintain effective optimization capabilities for information timeliness.

[0007] Therefore, this invention proposes a heterogeneous satellite network access control method optimized for PAoI, which can combine a block semantic awareness quantum feature extraction mechanism with a hierarchical knowledge distillation deployment strategy to improve the collaborative decision-making capabilities of multiple gateways while achieving efficient deployment of the model on resource-constrained nodes. Summary of the Invention

[0008] To address the shortcomings of classical feature extractors in access control scenarios with the goal of minimizing PAoI in large-scale heterogeneous satellite constellations, where the number of parameters increases quadratically with the state dimension and is difficult to deploy at edge gateways, this invention designs a block-based semantically aware parameterized quantum circuit. This circuit processes states independently in semantic blocks, achieving efficient feature extraction with the number of parameters increasing linearly only with the number of qubits. Furthermore, a hierarchical knowledge distillation scheme is designed to transfer the state awareness capabilities and PAoI optimization decision logic of the quantum teacher network to a purely classical lightweight student network, forming a deployable lightweight access control method.

[0009] To achieve the above objectives, the technical solution provided by the present invention is as follows:

[0010] A heterogeneous satellite network access control method based on block semantic-aware parameterized quantum circuits and hierarchical knowledge distillation includes the following steps:

[0011] Step 1: Construct a GEO / LEO heterogeneous satellite network architecture, set up GEO and LEO satellite constellations, initialize ground gateway equipment and data transmission parameters, and build a channel model and satellite orbit simulation configuration adapted to actual application scenarios to provide basic support for subsequent dynamic access decisions that conform to natural environmental conditions.

[0012] Step 2: Define the extended system state space and explicitly split the gateway state vector into three functional blocks according to semantics: decision history block, channel topology block, and time-sensitive queue block. This provides a direct basis for the independent processing of the subsequent block-based semantically aware parameterized quantum circuit.

[0013] Step 3: Design BS-PQC: For each semantic block, after compression by its own independent encoder, angle encoding is performed using a parameterized quantum rotation gate to prepare the initial quantum state; then, quantum state evolution is performed through a parameterized quantum circuit composed of alternating stacked rotation layers and entanglement layers; Pauli Z-basis measurement is performed on the final state to obtain the classical feature vector; finally, through a cross-block fusion layer and a classical post-processing network, a high-dimensional feature representation aligned with the multi-head attention collaborative decision-making module is output; each gateway is configured with an independent BS-PQC instance to avoid gradient competition among multiple gateways.

[0014] Step 4: Using BS-PQC as the core feature extractor, and combining the multi-head attention collaboration module and the PDS state-assisted learning module, construct a complete QMAPS-D3QN quantum teacher network; under the framework of centralized training and distributed execution, with minimizing PAoI as the optimization objective, train the quantum teacher network through a dual loss function jointly driven by gateway collaborative perception based on multi-head attention mechanism and post-decision state-assisted learning.

[0015] Step 5: To address the engineering bottleneck that prevents the direct deployment of the quantum teacher network to the edge nodes of a computing-constrained gateway, a hierarchical knowledge distillation scheme is designed. The student network is constructed by replacing the BS-PQC feature extractor with a lightweight, purely classical multilayer perceptron. The total distillation loss is a triple-weighted combination of feature layer alignment loss, Q-value soft label loss, and task-specific temporal difference error loss. This allows the high-dimensional state perception capability and PAoI decision logic of the quantum teacher network to be fully transferred to the student network. After distillation training and performance verification, a lightweight model with compressed parameters is provided for subsequent deployment at the gateway edge.

[0016] Step 6: In the actual system, the distilled student network is deployed to various ground gateway nodes. Online access decisions are made using pure classical multilayer perceptron inference, without the need for quantum computing hardware. Satellite access actions are output based on the real-time network status, realizing dynamic and lightweight satellite access control.

[0017] Beneficial effects

[0018] The method of this invention introduces a segmented semantic-aware parameterized quantum circuit, which independently processes the system state according to three semantic blocks: decision history, channel topology, and time-sensitive queue. The number of trainable parameters of the quantum circuit increases only linearly with the number of qubits, fundamentally solving the problem that the number of parameters of the classical fully connected feature extractor increases quadratically with the number of satellites. This provides a feasible high-dimensional state feature extraction scheme for resource-constrained satellite ground gateways.

[0019] The method of this invention adopts a distributed quantum feature extraction architecture, configuring an independent BS-PQC instance for each gateway, effectively avoiding the optimization competition problem of mutual interference of gradient signals when multiple gateways share quantum parameters. In collaboration with the multi-head attention collaborative decision module, it realizes high-quality extraction of local features and effective perception of the global competitive situation, significantly reducing the system PAOI in multi-gateway high-load heterogeneous scenarios.

[0020] This invention innovatively designs a triple-layered knowledge distillation scheme. Through a weighted combination of feature layer alignment loss, Q-value soft label loss, and task-specific temporal difference error loss, it ensures the integrity of quantum feature perception capability transfer while introducing the student network's own PAoI task constraint to prevent performance degradation. The distilled lightweight student network has approximately 29% of the parameters of the teacher network, but with only a slight performance degradation. It can run efficiently on purely classical embedded processors without quantum computing hardware, providing a lightweight deployment solution that balances performance and feasibility for satellite ground gateways with limited computing power. Attached Figure Description

[0021] Figure 1 This is a diagram of the GEO / LEO heterogeneous satellite network architecture used in this invention;

[0022] Figure 2 This is a diagram of the overall architecture for access control of heterogeneous satellite networks based on BS-PQC and hierarchical knowledge distillation.

[0023] Figure 3 A detailed structural diagram of the BS-PQC feature extraction module;

[0024] Figure 4 This is a schematic diagram of the multi-head attention module structure;

[0025] Figure 5 This is a schematic diagram of the PDS-assisted training module structure;

[0026] Figure 6 A comparison of the PAoI convergence curves of QMAPS-D3QN and DQN during the training phase;

[0027] Figure 7 A comparison of the PAoI variation curves of QMAPS-D3QN and DQN during the testing phase;

[0028] Figure 8 A comparison of the PAoI change curves of the QMAPS-D3QN teacher model and the distilled student model during the testing phase; Detailed Implementation

[0029] This embodiment uses a marine Internet of Things (IoT) application scenario as an example for illustration. The GEO / LEO heterogeneous satellite network architecture is as follows: Figure 1 As shown. Considering the special requirements of marine IoT applications for information timeliness, traditional single satellite systems cannot simultaneously meet the requirements of low latency and wide coverage. GEO satellites have stable coverage but relatively high propagation latency, while LEO satellites have low latency but suffer from frequent switching issues. Therefore, this invention adopts a GEO / LEO heterogeneous satellite network architecture, making full use of the complementary advantages of the two types of satellites, and solves the parameter efficiency and lightweight deployment challenges in large-scale constellation scenarios through block semantic quantum feature extraction and hierarchical knowledge distillation. The overall architecture of the model is as follows. Figure 2 As shown.

[0030] The system comprises a constellation of several GEO and LEO satellites. The LEO constellation uses a multi-orbit configuration, with multiple satellites deployed on each orbital plane, ensuring that multiple LEO satellites are visible simultaneously within the target sea area, providing the gateway with diverse access options.

[0031] set up This refers to the collection of Internet of Remote Things (IoRT) devices. This represents the set of Gateways (GAs). Let... This represents the set of satellites, or SAs, where... Indicates GEO satellite, Let n represent the LEO satellite; in this experiment, n = 1.

[0032] To characterize the temporal evolution of satellite visibility and link status, a satellite orbital model needs to be established first. For GEO satellites, since their position relative to the ground is approximately fixed, they can be considered as reference links providing stable wide-area coverage; for LEO satellites, the position changes caused by their orbital motion need to be explicitly described. Therefore, this invention uses a geocentric rectangular coordinate system to describe the position of LEO satellites in each time slot and calculates their trajectory over time based on orbital parameters. For any LEO satellite, its orbital simulation uses the following mathematical model: The formula for calculating the position of an LEO satellite in the geocentric rectangular coordinate system is:

[0033] x = R orb cos(μ)cos(Ω)-R orb sin(μ)sin(Ω)cos(i)

[0034] y = R orb cos(μ)sin(Ω)-R orb sin(μ)cos(Ω)cos(i)

[0035] z = R orb sin(μ)sin(i)

[0036] Where R orb Let be the orbital radius, μ be the mean anomaly, Ω be the right ascension of the ascending node, and i be the orbital inclination. These orbital parameters are chosen based on typical configurations of actual LEO constellations to ensure simulation accuracy. The time evolution equation for the mean anomaly is:

[0037]

[0038] Where μ0 is the initial mean anterior angle, and T orbital Δt represents the orbital period and Δt represents the time slot length. Satellite visibility is determined using an elevation angle threshold method.

[0039]

[0040] in Let P be the unit zenith vector at the location of gateway m. n (t) represents the position vector of satellite n in time slot t, P m Let l be the location vector of gateway m. m,n (t) represents the distance between gateway m and satellite n. When the elevation angle e m,n (t) is higher than the preset minimum elevation angle threshold e min That is, e m,n (t)≥e min At that time, satellite n is deemed visible to gateway m, and the visibility flag v is recorded. m,n (t) = 1, otherwise v m,n (t) = 0.

[0041] Building upon visibility, a channel model is needed to describe the actual transmission capabilities of different satellite links. Considering the strong line-of-sight component typically present over the ocean, coupled with the effects of sea surface reflection and scattering, this invention employs a Rician fading model to describe Ka-band satellite links. This model reflects the propagation characteristics dominated by the direct sunlight component in the marine environment while preserving the impact of random fading on link quality. The channel coefficients from gateway m to satellite n in time slot t are:

[0042]

[0043] Where η m,n (t)=(λ / 4πl m,n (t)) 2Here, λ is the free-space path loss coefficient, and l is the carrier wavelength. m,n (t) represents the distance from gateway m to satellite n. The Rician fading coefficient h... Rician (t) is:

[0044]

[0045] Where K is the Rician factor, used to characterize the relative strength of the line-of-sight component and the scattering component, φ is the initial phase of the line-of-sight component, X and Y are independent standard Gaussian random variables, and j is the imaginary unit. This modeling method ensures that the link quality is affected by both changes in orbital distance and the instantaneous fluctuations caused by random fading, thus more closely reflecting the actual marine satellite communication environment.

[0046] Based on this, the signal-to-noise ratio (SNR) and achievable data rate of the gateway when transmitting via the satellite link are further calculated. The SNR of gateway m when transmitting via satellite n in time slot t can be expressed as:

[0047]

[0048] Where P m,n G represents the transmission power of gateway m to satellite n. m,n For antenna gain, σ 2 This refers to noise power. Calculating noise power requires considering system noise temperature and signal bandwidth.

[0049] σ 2 =k B T eff B n

[0050] Where k B T is the Boltzmann constant. eff For effective noise temperature, B n Let n be the bandwidth of satellite n.

[0051] According to Shannon's formula, the achievable data transmission rate of gateway m through satellite n in time slot t is:

[0052]

[0053] Where N n (t) represents the number of gateways that simultaneously select to access satellite n in time slot t, and c m,n (t) represents the data transmission rate of gateway m to satellite n in time slot t. This formula shows that when multiple gateways share the same satellite resource, the available bandwidth needs to be allocated among them. In other words, satellite access selection depends not only on the physical quality of a single link but also on the concurrent behavior of other gateways. This is an important reason why this invention needs to introduce a collaborative decision-making mechanism.

[0054] In marine IoT data backhaul scenarios, system performance depends not only on successful data delivery but also on the freshness of the information received at the receiving end. For services such as environmental monitoring, emergency warning, and situational awareness, even with high system throughput, if the data acquired at the receiving end is significantly outdated compared to the actual state at the source, it will still affect the effectiveness of subsequent decisions. Therefore, this invention requires the use of timeliness indicators that can characterize information freshness, rather than relying solely on traditional average latency or throughput indicators.

[0055] To describe the freshness of the information currently held by the receiving end, this invention first introduces AoI (Aspect of Age) to measure the time difference between the generation time of the latest data packet at the receiving end and the current time. This reflects the staleness of the information held by the receiving end relative to the actual state at the source end. For the IoRT device c, its information age at the receiving end is denoted as Δ. c (t), defined as the difference between the current time slot t and the time slot generated by the most recently successfully received data packet from this device:

[0056] Δ c (t)=tU c (t)

[0057] in K represents the generation time of the latest data packet from device c held by the receiving end in time slot t. c (t)=max{k|t′ c,k ≤t} is the index of the latest received data packet up to time slot t, t′ c,k The AoI represents the data packet reception time. This definition indicates that if no new state updates arrive within a certain period, the AoI will continue to increase over time; while when a new data packet is successfully received, the AoI will decrease to the system dwell time experienced by that data packet from generation to reception. Therefore, AoI can dynamically depict the entire process of information gradually becoming old and being successfully refreshed, making it more suitable than traditional average latency metrics for describing the timeliness of information in state update systems.

[0058] However, using AoI alone is insufficient to fully reflect the timeliness requirements addressed in this invention. This is because AoI focuses more on describing the instantaneous age of information received at a given moment, while for services such as emergency monitoring and disaster early warning, it is more important to consider the most outdated state of the system information before each successful update. In other words, in these scenarios, a single, large spike in timeliness often causes performance degradation more easily than the long-term average. Therefore, based on AoI, this invention further introduces PAoI as a core optimization metric.

[0059] PAoI represents the peak AoI reached by each data packet before successful reception, more directly reflecting the worst-case scenario of information freshness. It is more suitable for marine emergency scenarios where data timeliness is critical. The PAoI value of the k-th data packet generated by device c is calculated as follows:

[0060]

[0061] Where Δ c (0) represents the initial age value, I c,k =t c,k -t c,k-1 T is the interval for generating data packets. c,k t represents the time the data packet resides in the system. c,k Let t′ be the time when device c generates the k-th data packet. c,k This represents the time it takes for the data packet to be received by the receiving end. As this definition shows, PAoI is simultaneously affected by the data generation rhythm, queuing process, and link transmission capacity, making it more suitable for measuring the timeliness of information access control in heterogeneous satellite networks.

[0062] Since marine IoT data packets are queued at the gateway on a first-come, first-served basis, the index k′ of the data packet in the gateway m queue must satisfy... The time when the data packet left. The cumulative number of transmitted packets up to gateway m in time slot τ is given by the actual number of transmitted packets in time slot t. q is the theoretical maximum number of packets transmitted. m (t) represents the total number of data packets queued in time slot t.

[0063] Theoretical maximum number of packets transmitted Affected by channel conditions and handover delay:

[0064]

[0065] Where c m (t) represents the data transmission rate of gateway m, T represents the time slot length, and Γ represents the data transmission rate of gateway m. m (t) represents the handover delay, and D represents the data packet size.

[0066] The long-term average PAoI of device c is:

[0067]

[0068] Therefore, the present invention aims to minimize the overall average PAoI of all IoRT devices:

[0069]

[0070] This long-term optimization objective is achieved through the immediate reward function in reinforcement learning:

[0071]

[0072] in This is the set of data packets that were successfully transmitted in time slot t. A represents the number of elements in a set. p This is the PAoI value of data packet p.

[0073] The multi-gateway PAoI optimization access control problem is formalized as a Markov Decision Process (MDP). In terms of state space, the observation state vector s of gateway m in time slot t is... m (t) Semantically divided into three functional blocks:

[0074]

[0075] in Represents historical decision-making information. This represents channel topology information related to satellite visibility, link quality, and distance. This represents queue state information directly related to data timeliness. This state construction method can provide a well-structured input for subsequent block-based semantic-aware parameterized quantum circuits and reduce mutual interference between different semantic features in the unified mapping.

[0076] Figure 3 for Figure 2 The diagram shows a detailed structural diagram of the BS-PQC module in the overall architecture of heterogeneous satellite network access control, which includes modules such as block coding, quantum evolution, quantum measurement, cross-block fusion, and classical post-processing.

[0077] To address the issue of gradient signal interference caused by neglecting state heterogeneity in existing general-purpose angle encoding schemes, BS-PQC adopts a three-stage architecture: semantic block division, independent encoding of each block, and cross-block fusion. Through block encoding, quantum evolution, quantum measurement, cross-block fusion, and classical post-processing, it outputs locally enhanced features.

[0078] Then, after passing through the second hidden layer, the final feature representation matrix f is obtained. m It is then input into the multi-head attention layer for inter-gateway collaborative decision-making.

[0079] In the block-based angle encoding stage, for the b-th semantic block (b∈{D, C, T}), it is first encoded by its own independent encoder f. b (·) Compress the features to the corresponding number of qubits n b Then, through the learnable scaling parameter λ (b)Amplitude modulation is performed to obtain the coded rotation angle vector:

[0080]

[0081] Where ⊙ represents the Hadamard product of element-wise multiplication. Let n b The ground state of a qubit is the initial state, and R is applied to each qubit. Y The rotating door performs angle encoding to prepare the input quantum state of the b-th block:

[0082]

[0083] in This represents the Kronecker product between quantum state vectors or unitary matrices. Choose R. Y Gates are used because the amplitude of the single-bit state obtained after their operation remains in the real domain. They are more suitable for carrying encoded information mapped from real-valued states such as visibility, distance, and queue backlog, effectively avoiding phase instability caused by complex domain computation. R Y For a rotation gate about the Y-axis, its unitary matrix representation is:

[0084]

[0085] In the parameterized quantum circuit evolution stage, each semantic block corresponds to an independent parameterized quantum circuit of depth L, which evolves to the final state through L layers of parameterized unitary transformation; the complete unitary transformation of the l-th layer is a combination of a rotation layer and an entanglement layer, denoted as U. l :

[0086] Each semantic block corresponds to an independent PQC of depth L, and the evolution of this circuit is based on the quantum encoded states prepared earlier. As input, it evolves to the final state |ψ| through L layers of parameterized unitary transformation. (b) (Θ (b) )>,Θ (b) Let be the set of all trainable parameters for the PQC corresponding to the b-th semantic block.

[0087] When constructing a parametric rotating gate and performing the final measurement, it is necessary to first introduce Pauli matrices. The three Pauli matrices are as follows:

[0088]

[0089] Using these three matrices as generators, a single-qubit rotation gate is defined as follows:

[0090]

[0091] Where I is a 2×2 identity matrix. From this definition, R can be directly obtained. X R Y and R ZThe explicit matrix form; these rotation gates control the quantum states to rotate about the x, y, and z axes on the Bloch sphere, respectively. After expansion, R... X With R Z The matrix representation is as follows:

[0092]

[0093] The rotating layer sequentially applies R to each qubit i of sub-block b. X R Y and R Z The three-gate rotation, whose overall unitary operator can be represented as the Kronecker product of the unitary matrices corresponding to the single-bit rotation gates of each qubit:

[0094]

[0095] in Let R be the trainable rotation angle parameter of the j-th rotation gate on the i-th qubit of the l-th layer, where j∈{1,2,3} corresponds to R. X R Y and R Z Three types of rotation gates; combinations of the three types of rotation gates can generate arbitrary single-qubit unitary transformations in the special unitary group SU(2); the entanglement layer applies a CNOT gate sequence with a ring topology:

[0096]

[0097] CNOT c→t This represents a controlled NOT gate with qubit c as the control bit and qubit t as the target bit, and its function is... This is modulo-2 addition. It enables quantum circuits to further characterize the correlations between different feature dimensions.

[0098] Therefore, the complete unitary transformation of the l-th layer can be written as:

[0099]

[0100] Among them U ent This represents the unitary transform corresponding to the entanglement layer. Let l represent the unitary transformation corresponding to the l-th rotation layer, where l = 1, 2, ..., L. After L layers of alternating stacking, the final quantum state of sub-block b is:

[0101]

[0102] in This indicates that the circuit operates sequentially according to the order l = 1, 2, ..., L.

[0103] During the quantum measurement phase, the final state is projected along the Pauli z-basis to obtain the expected value of each qubit. The output vector of the entire sub-block is:

[0104]

[0105] in Let be the Pauli Z-observable acting on the i-th qubit.

[0106] In the cross-block fusion and post-processing stage, the three quantum blocks independently process the features of their respective semantics, but the three types of information are not independent of each other: the time-sensitivity queue information related to PAoI accumulation speed is closely related to the channel topology information related to the current link quality, while historical selection information affects the calculation of handover cost. Therefore, a cross-block fusion layer is designed to cross-mix the quantum measurement outputs of the three blocks. After concatenating the output vectors of the three quantum blocks, the cross-block fusion layer captures the nonlinear coupling relationships between different semantic states:

[0107]

[0108] Where ReLU(·) is the modified linear unit activation function, and LayerNorm(·) is the layer normalization function. W fuse and b fuse These are the weights and biases for the cross-block processing layer, respectively. (D) o (C) o (T) This is for quantum output. The role of the cross-block fusion layer is to open up the information channels between the three semantic blocks, so that the subsequent Q-value calculation can simultaneously perceive the joint influence of link quality, timeliness status, and historical decisions. After fusion, it is mapped to the feature space consistent with the multi-head attention layer through two layers of classical post-processing network, and finally outputs the local enhanced feature h. m The multi-head attention layer, global fusion layer, and Dueling Q branch each receive this feature to complete subsequent collaborative decision-making and action value estimation.

[0109] h m =LayerNorm(ReLU(W2LayerNorm(ReLU(Dropout(W1f m +b1)))+b2))

[0110] Where W1 and W2 are the weights of the post-processing layer, b1 and b2 are the learnable bias parameters of the post-processing layer, and Dropout(·) is the random deactivation regularization function.

[0111] Figure 4 This is a schematic diagram of the multi-head attention module structure, which is responsible for capturing the dependencies between gateways and generating attention-enhancing features. Figure 5This is a schematic diagram of the PDS-assisted training module, which is used to provide additional deterministic state supervision signals during the training phase.

[0112] The collaborative decision-making component is used to model the collaborative relationships between multiple gateways. It uses the BS-PQC output features h from the M gateways. m Stacking them row by row yields the shared feature matrix F = [h1; h2; ...; h M The multi-head attention layer contains H parallel attention heads, each independently capturing the dependencies between gateways within a different feature subspace.

[0113]

[0114] Where d k For each attention head, the feature dimensions, Let W be the query matrix, key matrix, and value matrix of the h-th attention head, respectively. The outputs of the H attention heads are concatenated along the feature dimension and projected through the output matrix W. O After performing a linear transformation, followed by residual connections and layer normalization, we obtain the attention-enhanced collaborative features:

[0115] F (attn) =LayerNorm(F+Dropout((H1,H2,...,H H W O ))

[0116] right That is, F (attn) The m-th row undergoes a gateway-by-gateway nonlinear fusion transformation to obtain the fusion features that integrate local information and global collaborative information. The data is fed into a Dueling Q-network for action value estimation. The Dueling architecture decomposes the Q-value into state value and action advantage.

[0117]

[0118] To improve the training efficiency of quantum teacher networks in heterogeneous satellite network access control scenarios, this invention introduces a PDS-assisted learning mechanism in addition to the main Q-learning task. The core idea of ​​this mechanism is that the state transition process in a multi-gateway access control system is not entirely random. One part of the state changes is directly determined by the current action, constituting predictable deterministic changes; the other part arises from environmental disturbances such as random packet arrivals, channel fading, and dynamic changes in satellite links, constituting stochastic changes. If the state transition can be separated into deterministic and stochastic components, the deterministic component can provide additional supervision signals in each time slot, thereby improving training convergence speed and enhancing policy learning stability.

[0119] Specifically, let the actual state of the current time slot be s, the action to be performed be a, and the intermediate state after the action is performed and before the random disturbance occurs be s. The actual state of the next time slot is s′. The state transition process can then be formally represented as:

[0120]

[0121] in PDS stands for the intermediate state between the current state s and the random disturbance, after action a is performed. The probability of a state transition is deterministic and is directly determined by the action. The random state transition probability is used to describe the state after the decision. The random change process of the actual state s′ in the next time slot reflects environmental uncertainties such as channel fading and random arrival of queues.

[0122] Based on the above decomposition method, the goal of the PDS prediction module in this invention is to learn a deterministic transfer function. It can be calculated immediately in each time slot by the deterministic transfer function without waiting for reward feedback, thus breaking through the signal bottleneck of sparse rewards and providing a denser training signal outside the main Q learning task.

[0123] In the quantum teacher network of this invention, the PDS prediction module works in parallel with the main Q network. Considering that the front-end features have already been extracted by BS-PQC and then processed through a cross-block fusion layer and a classical post-processing network to obtain the local enhanced features h... m Therefore, the PDS prediction module uses the local augmentation feature h of gateway m. m and current action a m The one-hot encoding of (t) is one_hot(a m (t) is used as input. First, the two are concatenated along the feature dimension and then input into the hidden layer for non-linear mapping:

[0124]

[0125] Where [f m one_hot(a m [(t))] represents vector concatenation. This is the first-layer weight matrix of the PDS prediction module. This is the corresponding bias vector.

[0126] The predicted PDS is then obtained through a linear mapping of the output layer.

[0127]

[0128] in This is the weight matrix for the second layer. This is the bias vector for the second layer.

[0129] To enable the PDS module and the main Q-learning task to be optimized collaboratively, this invention employs a dual loss function to train the quantum teacher network, with the total loss expressed as:

[0130] L total =L Q +αL PDS

[0131] Here, α is a weighting factor responsible for balancing the Q-learning task and the PDS prediction auxiliary task. The main Q-learning loss L... Q The difference between the predicted Q value and the TD target value is calculated using smoothed L1 loss:

[0132]

[0133] in This represents the mathematical expectation operation. Let Q be the time-series difference objective value, representing the expected cumulative return of the current estimate. target (·) represents the action value function of the target network, r m (t) represents the instant reward, s m (t+1) represents the state of gateway m in time slot t+1, and a′ represents the action space. All possible actions in the middle, To smooth the L1 loss function, it degenerates into L2 loss when the error is small and into L1 loss when the error is large:

[0134]

[0135] PDS Auxiliary Learning Loss L PDS MSE is used to measure the difference between predicted PDS and actual PDS:

[0136]

[0137] in This refers to the actual PDS observed directly from the environment.

[0138] The hierarchical knowledge distillation scheme is the core of this invention for achieving lightweight deployment. The quantum teacher network contains multiple BS-PQC instances, a multi-head attention layer, and a Dueling Q branch, with approximately 980,000 parameters; the student network replaces BS-PQC with a single hidden layer MLP, with an overall parameter count of approximately 280,000, which is about 29% of that of the teacher network, making it more suitable for the storage and computing power constraints of gateway edge processors.

[0139] Let the characteristic output of the teacher network be Action value output is The corresponding output of the student network is and The triple distillation loss function ensures the integrity of knowledge transfer at different levels:

[0140] Feature layer alignment loss By constraining the mean squared error between student and teacher features, the system encourages student networks to learn intermediate feature representations consistent with teacher networks. This ensures that student networks learn state representations with genuine generalization capabilities, rather than shortcut solutions that are merely similar in distribution at the output layer.

[0141]

[0142] When the teacher network and student network are aligned at the feature layer, the student network only updates its own parameters, while the teacher network parameters remain frozen.

[0143] Q-value soft tag output loss By using teachers' soft labels to convey the relative superiority or inferiority of actions, and requiring students' Q-value distribution to be numerically close to the teachers' Q-value distribution, consistency in decision-making logic is ensured.

[0144]

[0145] Compared to hard labels, Q-value distribution carries information about the relative merits of each candidate satellite, making it a richer monitoring signal.

[0146] Task-specific PAOI loss As a core extension of the general distillation framework in this invention, the bottom line of PAoI optimization is protected by introducing the temporal difference constraint of the student network itself:

[0147]

[0148] The student network optimizes itself using its own temporal difference objective, enabling it to maintain its independent optimization capability for the PAoI objective while learning the teacher's decision knowledge.

[0149] The total distillation loss is a weighted combination of the three:

[0150]

[0151] Where w f w o w t It is an adjustable weight parameter.

[0152] This invention verifies the effectiveness of the proposed method through simulation experiments. The experimental scenario includes one GEO satellite and 36 LEO satellites, four marine gateway devices, and 12 IoRT data sources. The carrier frequency is 30 GHz, the GEO bandwidth is 50 MHz, the LEO bandwidth is 30 MHz, the data packet size is 12 Mbits, the simulation duration is 2000 time slots, and the quantum circuit uses n b =4 qubits, L=3 circuit depth.

[0153] Figure 6 The PAoI convergence curves of the QMAPS-D3QN model and the traditional Deep Q-Network (DQN) model during the training phase are presented. It can be seen that the QMAPS-D3QN model rapidly decreases and stabilizes at a low PAoI level in the early stages of training, while the DQN model decreases more slowly and still fluctuates significantly in the later stages of training. Ultimately, the average PAoI of QMAPS-D3QN drops to 3.27, significantly lower than that of DQN (6.31). This indicates that the QMAPS-D3QN proposed in this invention can accelerate training convergence and improve the network's feature extraction and decision-making capabilities under multi-gateway collaboration.

[0154] Figure 7 The average PAoI variation curves of QMAPS-D3QN and DQN during the testing phase are shown. The QMAPS-D3QN teacher model maintains a lower PAoI, with an average value of 2.69, while the DQN model has an average PAoI of 8.26. The curves show that the teacher model maintains a low PAoI peak in most time slots, while the DQN model shows a higher peak in some time slots, verifying the more robust control capability of the QMAPS-D3QN teacher model in dynamic network environments.

[0155] Figure 8 The comparison of PAoI between the distilled student model and the teacher model in the current time slot during the testing phase is presented. The average PAoI of the teacher model is 2.69, while the average PAoI of the distilled student model is 2.80. The overall trends of the two are similar, indicating that the student model inherits the key decision logic of the teacher model well and maintains better performance under the condition of lightweight parameters, thus verifying the effectiveness of the hierarchical knowledge distillation method.

[0156] Experimental results show that the proposed QMAPS-D3QN method effectively improves PAoI optimization performance; the BS-PQC-based distributed quantum feature extraction architecture effectively avoids the gradient competition problem among multiple gateways by configuring independent instances for each gateway; and the hierarchical knowledge distillation scheme maintains superior performance while significantly compressing model parameters. The proposed method significantly improves parameter efficiency and PAoI optimization performance in heterogeneous satellite network scenarios, providing an efficient and lightweight solution for information time-sensitive applications.

Claims

1. This invention discloses a heterogeneous satellite network access control method based on block semantic-aware parameterized quantum circuits and hierarchical knowledge distillation, with PAoI minimization as the optimization objective, characterized in that, Includes the following steps: Step 1: Construct a heterogeneous satellite network architecture of Geosynchronous Earth Orbit / Low Earth Orbit (GEO / LEO), set up GEO and LEO satellite constellations, initialize ground gateway equipment and data transmission parameters, and build a channel model and satellite orbit simulation configuration adapted to the actual application scenario to provide basic support for subsequent dynamic access decisions that conform to natural environmental conditions. Step 2: Define the extended system state space and semantically and explicitly split the gateway state vector into three functional blocks: decision history block, channel topology block, and time-sensitive queue block. This provides a direct basis for the independent processing of the subsequent block-based semantically aware parameterized quantum circuit. Step 3: Design BS-PQC: For each semantic block, after compression by its own independent encoder, angle encoding is performed using a parameterized quantum rotation gate to prepare the initial quantum state; then, quantum state evolution is performed through a parameterized quantum circuit composed of alternating stacked rotation layers and entanglement layers; Pauli Z-basis measurement is performed on the final state to obtain the classical feature vector; finally, through a cross-block fusion layer and a classical post-processing network, a high-dimensional feature representation aligned with the multi-head attention collaborative decision-making module is output. Each gateway is configured with an independent BS-PQC instance to avoid gradient contention among multiple gateways; Step 4: Using BS-PQC as the core feature extractor, and combining it with the multi-head attention collaboration module and the post-Decision State (PDS) state-assisted learning module, a complete QMAPS-D3QN quantum teacher network is constructed. Under the framework of centralized training and distributed execution, with minimizing PAoI as the optimization objective, the quantum teacher network is trained through a dual loss function jointly driven by gateway collaborative perception based on multi-head attention mechanism and post-Decision State-assisted learning. Step 5: To address the engineering bottleneck that prevents the quantum teacher network from being directly deployed to the edge nodes of the computing-constrained gateway, a hierarchical knowledge distillation scheme is designed; a student network is constructed by replacing the BS-PQC feature extractor with a lightweight, purely classical multilayer perceptron. A distillation total loss is adopted, which is a weighted combination of feature layer alignment loss, Q-value soft label loss and task-specific temporal difference error loss. The high-dimensional state awareness capability and PAoI decision logic of the quantum teacher network are completely transferred to the student network. After distillation training and performance verification, a lightweight model with compressed parameters is provided for subsequent gateway edge deployment. Step 6: In the actual system, the distilled student network is deployed to various ground gateway nodes. Online access decisions are made using pure classical multilayer perceptron inference, without the need for quantum computing hardware. Satellite access actions are output based on the real-time network status, realizing dynamic and lightweight satellite access control.

2. The heterogeneous satellite network access control method based on block semantic-aware parameterized quantum circuits and hierarchical knowledge distillation according to claim 1, characterized in that... The block coding and quantum circuit structure of BS-PQC in step 3 specifically includes: Step 3-1: Transfer the state vector s of gateway m in time slot t. m (t) Semantically split into three functional blocks: Decision History Block Select a one-hot coding vector for satellite access in the previous time slot for gateway m; channel topology block Includes the visibility flags v of each satellite to the gateway m (t), Received signal strength r m (t) and relative distance d m (t); Time-sensitive queue block Indicates the number of packets currently backlogged in the gateway queue; Step 3-2: For the b-th semantic block (b∈{D, C, T}), use their respective independent classical encoders f b (·) Compress the features to the corresponding number of qubits n b Learnable scaling parameter λ (b) Modulation yields the encoded rotation angle vector: in For the gateway's state semantic block in the time slot, ⊙ represents the Hadamard product of element-wise multiplication, with n... b The ground state of a qubit is the initial state, and R is applied to each qubit. Y The rotating door performs angle encoding to prepare the input quantum state of the b-th block: in The Kronecker product of quantum states, For vector θ (b) For the i-th element, choose R. Y The gate is used because the amplitude of the single-bit state obtained after its action remains in the real number domain, making it more suitable for carrying encoded information mapped from real-valued states such as visibility, distance, and queue backlog. Step 3-3: Each semantic block corresponds to an independent parameterized quantum circuit of depth L, which evolves to the final state through L layers of parameterized unitary transformation; the complete unitary transformation of the l-th layer is a combination of a rotation layer and an entanglement layer, denoted as U. l : Among them U ent This represents the unitary transform corresponding to the entanglement layer. This represents the unitary transformation corresponding to the l-th rotating layer, where l = 1, 2, ..., L; The rotating layer sequentially applies R to each qubit i of sub-block b. X R Y and R Z The three-gate rotation, whose overall unitary operator can be represented as the Kronecker product of the unitary matrices corresponding to the single-bit rotation gates of each qubit: in Let R be the trainable rotation angle parameter of the j-th rotation gate on the i-th qubit of the l-th layer, where j∈{1,2,3} corresponds to R. X R Y and R Z Three types of rotation gates; combinations of the three types of rotation gates can generate arbitrary single-qubit unitary transformations in the SU(2) group; the entangled layer applies a controlled-NOT (CNOT) gate sequence with a ring topology: The total number of trainable PQC parameters for sub-block b is 3Ln. b The PQC parameter efficiency advantage is directly reflected in the fact that it is not directly bound to the input feature dimension. Steps 3-4: After L layers of alternating stacking, the final quantum state of sub-block b before measurement is: in This indicates that the circuit operates sequentially according to the order l = 1, 2, ..., L, Θ (b) Let be the set of all trainable parameters for the PQC corresponding to the b-th semantic block. For the input quantum state of the b-th semantic block, To measure the previous and final states. By projecting measurements along the Pauli Z basis onto the final state, the expected values ​​of each qubit are obtained, and the output vector of the entire sub-block is: in Θ (b) This represents the set of all trainable parameters for the PQC corresponding to the b-th semantic block. The Pauli Z-observable is applied to the i-th qubit; the output vectors of the three qubits are concatenated and then passed through a cross-block fusion layer to capture the nonlinear coupling between different semantic states: Among them W fuse and b fuse These are the weights and biases for the cross-block processing layer, o (D) o (C) o (T) The quantum output is then mapped through two layers of classical post-processing networks to a feature space consistent with the multi-head attention layer, ultimately outputting the local augmented feature h. m The data is then fed into the multi-head attention collaborative decision-making module. h m =LayerNorm(ReLU(W2LayerNorm(ReLU(Dropout(W1f m +b1)))+b2)) Where W1 and W2 are the weights of the post-processing layer, and b1 and b2 are the learnable bias parameters of the post-processing layer.

3. The method according to claim 1, characterized in that, Based on the QMAPS-D3QN quantum teacher network, step 5 introduces a hierarchical knowledge distillation mechanism to significantly reduce network parameters, specifically including: Step 5-1: Construct a lightweight student network, using a single-hidden-layer multi-layer perceptron (MLP) as the shared classic feature extractor to replace the BS-PQC deployed independently by each gateway in the teacher network; retain the structurally identical multi-head attention collaborative decision layer and the independent Dueling Q branch of each gateway in the teacher network to ensure the fairness of the lightweight distillation comparison; load the trained teacher network and freeze all its parameters. Step 5-2: Design a triple-layered distillation loss function to ensure the integrity of knowledge transfer from three levels: feature representation, Q-value distribution, and PAoI task objective; Feature layer alignment loss By forcing student features to align with teacher features, we ensure that the student network learns a state representation with generalization capabilities. Where M is the total number of gateways, and MSE(·) represents the mean square error function. The feature vector output by the teacher network to gateway m has fixed parameters that are not updated. Let m be the corresponding feature vector output by the student network to gateway m. Q-value soft tag loss By using Q-value soft labels from the teacher network to convey the relative merits of actions, consistency in decision-making logic can be ensured. in and Estimate the Q-values ​​of gateway m for the student and teacher networks, respectively; Task-specific PAOI loss As a core extension to the distillation framework, a temporal difference constraint is introduced from the student network itself to protect the PAoI optimization baseline: in Let the time-series difference target value be the student network. The student master network is used to select the action that maximizes the Q-value. For the student target network, the parameters are periodically soft-updated to evaluate the value of the action, r. m Let γ be the immediate PAoI reward obtained by gateway m in the current time slot, and γ be a discount factor. Here, the student target network, rather than the teacher network, is used to calculate the Temporal Difference (TD) objective. The TD objective value is a supervised objective constructed based on the Bellman optimal equation, used to measure the difference between the current student network Q-value estimate and the target reward jointly determined by the immediate reward and the next-state target network estimate. This ensures that the student network learns from the teacher while independently optimizing the PAoI objective, preventing over-reliance on the teacher from causing PAoI performance degradation. Step 5-3: Jointly optimize the total distillation loss as a weighted combination of the three losses: Where w f w o w t As adjustable weight parameters, the setting of the three weights reflects the priority of each loss: Q-value, soft-label loss, and weight w. o At its highest level, it ensures a direct and accurate match to the decision-making logic; the feature layer loss weights w f Secondly, ensure the generalization ability of the underlying feature representation; task loss weight w t At the lowest level, a bottom-line protection mechanism is implemented as the optimization objective of PAoI; in each distillation training iteration, feature vectors are obtained using frozen quantum teacher network inference. and Q value Total distillation loss The gradient updates the student network parameters.