Multi-track satellite cooperative communication link optimization method and system
Patent Information
- Application Number
- CN202610861190.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-15
AI Technical Summary
[0004]本申请实施例提供了一种多轨卫星协同通信链路优化方法和系统,能够解决现有卫星通信方案缺乏对动态遮挡物的实时响应能力,导致切换时延超时,高频段NTN通信质量严重下降的问题
[0008]The multi-track satellite cooperative communication link optimization method and system provided in this application can improve the accuracy of terahertz band obstruction detection and reduce the signal attenuation false judgment rate by adopting multimodal frequency domain attention sensing; it can achieve ultra-high-speed orbit optimization by adopting quantum reinforcement learning to meet the requirements of 6G NTN sub-millisecond handover and ultra-reliable low-latency communication (URLLC); it can achieve beam, orbit and thrust joint optimization to reduce link handover latency and improve dynamic environment adaptability, thereby effectively solving problems such as dynamic cloud obstruction, terahertz signal attenuation, excessive handover latency, excessive fuel consumption and insufficient system stability, and significantly improving the reliability and real-time performance of 6G NTN communication links.
Smart Images

Figure CN122764291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of satellite communication technology, and in particular to a method and system for optimizing multi-track satellite cooperative communication links. Background Technology
[0002] Currently, sixth-generation mobile communication technology (6G) th Generation Mobile Networks (6G) are a global focus. 6G Non-Terrestrial Networks (NTNs) are a key technology of 6G. By integrating satellite communications (Low Earth Orbit, Medium Earth Orbit, and Geostationary Earth Orbit, GEO satellites) with terrestrial mobile networks, 6G NTNs can achieve seamless global coverage, providing broadband connectivity to remote areas, oceans, and airspace—regions traditionally difficult for terrestrial networks to reach.
[0003] However, 6G NTN needs to support millions of terminal connections per square kilometer, and dense cloud cover causes a 37% probability of terahertz link disruption. Traditional satellite communication solutions rely on ground base station-assisted handover or satellite trajectory prediction based on preset orbit parameters or static obstruction models, lacking real-time response capabilities to dynamic obstructions. In addition, 6G high-frequency bands (such as terahertz) are susceptible to atmospheric attenuation, and traditional obstruction detection algorithms have insufficient resolution, leading to handover delay timeouts and a severe deterioration in the communication quality of high-frequency NTN. Summary of the Invention
[0004] This application provides a method and system for optimizing multi-track satellite cooperative communication links, which can solve the problem that existing satellite communication schemes lack the ability to respond to dynamic obstructions in real time, resulting in handover delay timeouts and a serious deterioration in the quality of high-frequency NTN communication.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows: Firstly, a method for optimizing multi-orbit satellite cooperative communication links is provided, including: A high-orbit satellite acquires multimodal sensing data of a target area and obtains transmission parameters corresponding to the target area based on the multimodal sensing data; wherein the multimodal sensing data includes at least one of the following: frequency domain data of communication signals and spatial domain imaging data of physical fields; the transmission parameters are used to characterize the degree of attenuation of communication signals caused by obstructions; Based on the current environmental state, the high-orbit satellite uses a quantum reinforcement learning strategy to select the optimal target orbit from multiple satellite orbits covering the target area, and then distributes the optimal target orbit to the low-orbit satellites serving the target area; wherein, the current environmental state is obtained based on the transmission parameters corresponding to the target area; The low-Earth orbit satellite uses the optimal target orbit as its initial orbit. Based on the coupling relationship between orbit change and beam pointing, the relationship between metasurface phase and beam pointing adjustment, and the dynamic relationship between thrust and orbit change, the orbit adjustment parameters are determined. The orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters. The low-orbit satellite performs beam calibration and orbit adjustment on the communication link facing the target area based on the orbit adjustment parameters.
[0006] Secondly, a multi-orbit satellite cooperative communication link optimization system is provided, which includes at least: high-orbit satellites and multiple low-orbit satellites; The high-orbit satellite is used to acquire multimodal sensing data of the target area and obtain the transmission parameters corresponding to the target area based on the multimodal sensing data; wherein, the multimodal sensing data includes at least one of the following: communication signal frequency domain data and physical field spatial domain imaging data; the transmission parameters are used to characterize the degree of communication signal attenuation caused by obstructions; The high-orbit satellite is also used to select the optimal target orbit from multiple satellite orbits covering the target area based on the current environmental state using a quantum reinforcement learning strategy, and to distribute the optimal target orbit to the low-orbit satellites serving the target area; wherein, the current environmental state is obtained based on the transmission parameters corresponding to the target area; Each of the low-Earth orbit satellites serving the target area is used to determine orbit adjustment parameters based on the optimal target orbit as the initial orbit, the coupling relationship between orbit change and beam pointing, the regulation relationship between metasurface phase and beam pointing, and the dynamic relationship between thrust and orbit change. The orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters.
[0007] The low-orbit satellite is also used to perform beam calibration and orbit adjustment on the communication link facing the target area based on the orbit adjustment parameters and the initial orbit.
[0008] The multi-track satellite cooperative communication link optimization method and system provided in this application can improve the accuracy of terahertz band obstruction detection and reduce the signal attenuation false judgment rate by adopting multimodal frequency domain attention sensing; it can achieve ultra-high-speed orbit optimization by adopting quantum reinforcement learning to meet the requirements of 6G NTN sub-millisecond handover and ultra-reliable low-latency communication (URLLC); it can achieve beam, orbit and thrust joint optimization to reduce link handover latency and improve dynamic environment adaptability, thereby effectively solving problems such as dynamic cloud obstruction, terahertz signal attenuation, excessive handover latency, excessive fuel consumption and insufficient system stability, and significantly improving the reliability and real-time performance of 6G NTN communication links.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0011] Figure 1 This paper shows a structural block diagram of the multi-track satellite cooperative communication link optimization system provided in an embodiment of this application; Figure 2 A flowchart of the multi-track satellite cooperative communication link optimization method provided in an embodiment of this application is shown; Figure 3 This document illustrates a flowchart illustrating the process of obtaining transmission parameters corresponding to a target region based on multimodal sensing data, as provided in an embodiment of this application. Figure 4 This document illustrates a flowchart illustrating the use of quantum reinforcement learning to select the optimal target trajectory, as provided in an embodiment of this application. Figure 5 The flowcharts illustrating the trajectory selection evaluation, experience screening, and strategy optimization provided in the embodiments of this application are shown. Figure 6 This document illustrates a flowchart illustrating the determination of the optimal fuel control strategy based on disturbance scenarios and orbital adjustment parameters, as provided in an embodiment of this application. Figure 7 A flowchart of the context-aware differential privacy federated learning training process provided in an embodiment of this application is shown; Figure 8 The flowchart of adaptive Lyapunov neural differential stability verification and rollback control provided in the embodiments of this application is shown; Figure 9 A structural block diagram of an electronic device provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0012] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0013] To address the problem that existing satellite communication solutions lack real-time response capabilities to dynamic obstructions, leading to handover delays and a severe degradation in the quality of high-frequency NTN communication, this application provides a method and system for optimizing multi-track satellite cooperative communication links.
[0014] To facilitate understanding of the multi-track satellite cooperative communication link optimization method provided in the embodiments of this application, a brief introduction to the multi-track satellite cooperative communication link optimization system provided in the embodiments of this application will be given first.
[0015] Figure 1 This diagram illustrates the structure of a multi-track satellite cooperative communication link optimization system 10, as shown in an exemplary embodiment of this application. Figure 1 As shown, the system includes high-orbit satellites, multiple medium-orbit satellites, and multiple low-orbit satellites. The high-orbit satellites are used for global perception, quantum decision-making, and model aggregation; the low-orbit satellites are used for link execution, beam control, and orbit adjustment; and the medium-orbit satellites are used for regional coverage and service relay. The satellites in the system can communicate with each other and coordinate scheduling to jointly optimize communication links in the target area. These embodiments are only for explaining this application and do not constitute a limitation thereof.
[0016] Below, in conjunction with Figures 2 to 8 The multi-orbit satellite cooperative communication link optimization method provided in this application is explained and described in detail. It is understood that each low-orbit satellite in the system can have its communication link optimized according to the following implementation methods. These embodiments are only used to explain this application and do not constitute a limitation thereof.
[0017] Figure 2 A flowchart illustrating the multi-track satellite cooperative communication link optimization method provided in an embodiment of this application is shown. Figure 2 As shown, the multi-track satellite cooperative communication link optimization method mainly includes the following steps (S101-S104): S101: High-orbit satellites acquire multimodal sensing data of the target area and obtain the corresponding transmission parameters of the target area based on the multimodal sensing data; In this embodiment, the multimodal sensing data includes at least one of the following: frequency domain data of communication signals and spatial domain imaging data of physical fields.
[0018] In some embodiments, the frequency domain data of the communication signal includes the spectrum, power spectrum, bit error rate, etc. of the received signal, which is used to reflect the changes in the signal itself. For example, terahertz image data collected by high-orbit satellites or ground stations is an image taken by a terahertz camera, which contains frequency domain information and reflects signal attenuation.
[0019] In some embodiments, physical field spatial imaging data includes cloud images from meteorological satellites, infrared imaging, radar data, etc., used to reflect environmental changes. For example, humidity, temperature, cloud images, rainfall intensity, etc., provided by meteorological satellites, infrared imagers, or radar can reflect cloud temperature, thickness, and structure.
[0020] In this embodiment, the transmission parameter is used to characterize the degree of communication signal attenuation caused by obstructions. To improve the accuracy of dynamic obstruction sensing and address the problems of low efficiency and inaccurate attenuation modeling in traditional dual-modal networks for frequency domain feature processing, this embodiment further provides a refined method for calculating the transmission parameter. In some embodiments, Figure 3 This document illustrates a flowchart illustrating how transmission parameters of a target region are obtained based on multimodal sensing data, as provided in an embodiment of this application. Figure 3 As shown, the transmission parameters corresponding to the target area are obtained based on multimodal sensing data, including the following steps (S201-S204): S201. Perform frequency domain transformation on the multimodal sensing data to obtain the corresponding frequency domain features; For example, frequency domain transformations are performed on the acquired terahertz image data and infrared image data respectively to obtain frequency domain features. and .
[0021] S202. Use a frequency domain attention mechanism to learn the attention weights for different frequency domain features; For example, a frequency domain attention module is introduced to adaptively learn the attention weights for different frequency domain components.
[0022] For example, the frequency domain features can be calculated using the following formula. and attention weights :
[0023]
[0024] in, and These are learnable parameters of the attention mechanism. Represents the attention calculation function. Used for normalizing weights. and These are the attention weight matrices for terahertz and infrared frequency domain features, respectively.
[0025] S203. Weighted fusion based on attention weights and frequency domain features yields the fused frequency domain features; For example, the attention weights and frequency domain features are weighted and fused using the following formula to obtain the fused frequency domain features. :
[0026] in, This indicates that the matrix elements are multiplied together. This indicates channel splicing.
[0027] S204. Determine the transmission parameters corresponding to the target region based on the fused frequency domain characteristics.
[0028] In some embodiments, the transmission parameter is a transmission matrix.
[0029] For example, the fused frequency domain features The input is fed into a convolutional capsule encoding layer to extract high-level feature representations. :
[0030] in, This represents convolutional capsule encoding operations, including convolution, activation functions, and dynamic routing.
[0031] High-level feature representation is achieved through a capsule decoding layer and a sigmoid activation function. Decoded into a transmittance matrix .
[0032]
[0033] in, These are the learnable weights of the decoding layer. Limiting the output value to between 0 and 1 represents the transmittance.
[0034] In this embodiment, high-precision perception of obstructions such as clouds and atmosphere is achieved through dual-modal data fusion, providing environmental basis for subsequent track selection.
[0035] S102, based on the current environmental conditions, the high-orbit satellite uses a quantum reinforcement learning strategy to select the optimal target orbit from multiple satellite orbits covering the target area, and then distributes the optimal target orbit to the low-orbit satellites serving the target area; In this embodiment, the current environmental state is obtained based on the transmission parameters corresponding to the target area, and is used to characterize the occlusion attenuation and serviceable orbit status of the target area.
[0036] Existing technologies often rely on static rules or traditional reinforcement learning for orbit selection, which suffers from slow decision-making, high latency, and inability to adapt to the ultra-dense 6G NTN constellation scheduling. To achieve ultra-high-speed orbit matching and reduce handover latency, this application employs quantum reinforcement learning for orbit optimization.
[0037] In some embodiments, the multiple satellite orbits covering the target area are pre-defined candidate orbits, each of which can achieve signal coverage of the target area. For example, for a remote target area, four low-Earth orbit satellite orbits that can stably cover the area are preset, and the selection is performed only from this set of candidate orbits.
[0038] In some embodiments, Figure 4 A flowchart illustrating the use of quantum reinforcement learning to select the optimal target trajectory, as provided in an embodiment of this application, is shown. Figure 4 As shown, a quantum reinforcement learning strategy is used to select the optimal target orbit from multiple satellite orbits covering the target area, including the following steps (S301-S303): S301. Encode multiple satellite orbits into a quantum superposition state, and iteratively update the quantum superposition state through quantum circuitry; For example, will The candidate satellite orbit codes are as follows: Superposition of qubits Cj .
[0039] For example, the initial state can be a uniform superposition state:
[0040] Through parameterized quantum rotating gates The quantum circuit composed of the controlled NOT gate CONT updates the quantum state:
[0041] For example, in each iteration, through a parameterized quantum rotation gate Update quantum state This enables iterative strategy implementation. A quantum circuit structure based on RY rotating gates and CNOT gates is employed to enhance the policy's expressive power and optimization efficiency.
[0042]
[0043]
[0044] in, It acts on the first RY rotation gates on qubits The control bit is The target bit is The CNOT scandal. These are learnable quantum gate parameters.
[0045] S302. The orbital probability distribution of each satellite orbit is obtained by quantum measurement based on the updated quantum superposition state; For example, the probability of each orbital being selected is obtained by measurement for quantum states. Measurements were performed to obtain the probability distribution. Select the current optimal orbit based on probability distribution sampling. .
[0046] S303. Select the optimal target orbit from multiple satellite orbits based on the orbit probability distribution.
[0047] For example, the orbit with the highest probability value is selected as the current optimal target orbit. .
[0048] In this embodiment, the parallelism of quantum computing accelerates the orbit search, enabling sub-millisecond-level decision-making and meeting the ultra-low latency communication requirements of 6GNTN.
[0049] In some embodiments, the method provided in this application further includes steps for evaluating the effectiveness of track selection and learning from experience. Figure 5 A flowchart illustrating the trajectory selection evaluation, experience screening, and strategy optimization provided in an embodiment of this application is shown. Figure 5 As shown, the method provided in this application embodiment further includes the following steps (S401-S403): S401. Calculate the reward value corresponding to the optimal target orbit based on link communication quality, switching delay, orbit adjustment speed and transmission parameters. For example, low-Earth orbit satellites select the optimal target orbit. Interact with the environment and calculate the reward value according to the following formula: ,in, It is the average value of the transmission parameters obtained in step S101 over the target area where the user terminal is located, used to characterize the link quality. It is the link switching delay. It is the track adjustment speed. These are preset weighting coefficients. The higher the reward value, the better the track selection, the higher the link quality, and the better the latency and energy consumption.
[0050] S402. Determine the fidelity and quantum importance of the updated quantum superposition state; For example, the updated quantum superposition state is compared with the target quantum state to calculate the quantum state fidelity:
[0051] Simultaneously, the quantum importance (QI) of the current empirical sample is calculated:
[0052] in, Let J be the density matrix corresponding to the j-th empirical sample. Let Tr be the target density matrix. The ) denotes the matrix trace operation. Quantum importance is used to characterize the contribution of a sample to policy optimization.
[0053] S403. When both the fidelity and quantum importance are greater than the preset threshold, the updated quantum superposition state, the optimal target orbit, and the reward value are used as the current interaction data and stored in the experience set.
[0054] For example, a quantum gating mechanism is introduced to control the updating and sampling of the experience set D. Experience data is only added to the experience set or sampled for training when it meets specific quantum conditions. For instance, a quantum threshold is set. Only when the fidelity of the quantum state is... (greater than) Only then will the experience data be added to the experience set D.
[0055] For example, the sampling of the experience set D also adopts a quantum gating mechanism, prioritizing the sampling of samples with higher quantum importance.
[0056] In some embodiments, the method provided in this application further includes: extracting historical interaction data from an experience set, and using the extracted historical interaction data to update the quantum gate parameters in the quantum circuit, so as to iteratively update the quantum superposition state through the quantum circuit.
[0057] For example, using empirical data in the empirical set D, the quantum gate parameters are updated through optimization algorithms such as gradient descent. To maximize the expected cumulative reward, a gradient descent algorithm can be used to optimize the quantum gate parameters.
[0058] in, Let the quantum action value function be... Let D be the learning rate and D be the experience set. By updating the quantum gate parameters, the orbit selection strategy is iteratively optimized, improving the accuracy and convergence speed of subsequent orbit decisions.
[0059] In this embodiment, the quantum-gated experience replay mechanism can effectively utilize the parallel search capability and quantum state superposition characteristics of quantum computing to accelerate the training process of reinforcement learning, improve the stability and convergence speed of the algorithm, ensure the efficient use of experience data, and further enhance the intelligence and real-time performance of track selection.
[0060] S103. Low-Earth orbit satellites use the optimal target orbit as the initial orbit. Based on the coupling relationship between orbit change and beam pointing, the control relationship between metasurface phase and beam pointing, and the dynamic relationship between thrust and orbit change, orbit adjustment parameters are determined. In some embodiments, the orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters.
[0061] The optimal beam pointing parameters refer to the set of optimal pointing angles, phase configurations, and beam pointing angles when a reconfigurable metasurface array of a low-Earth orbit satellite is communicating with a target area. These parameters are used to align the beam center with the target area, reduce signal attenuation caused by obstruction, and improve the signal-to-noise ratio of communication. For example, the optimal beam pointing parameters may include azimuth angle, elevation angle, phase gradient, etc.
[0062] Optimal orbital parameters refer to the optimal orbital semi-major axis, orbital inclination, and orbital position required for a low-Earth orbit satellite to avoid obstruction and ensure coverage of the target area, so that the satellite's trajectory maintains a stable communication geometric relationship with the target area.
[0063] Optimal thrust parameters refer to the magnitude, direction, and duration of thrust required for a low-Earth orbit satellite to achieve optimal orbital adjustments, used to complete orbital maneuvers with minimal fuel consumption.
[0064] In some embodiments, the coupling relationship between orbital variation and beam pointing can be derived and established by beam pointing. With the semi-major axis of the track The coupled dynamics equations are implemented to describe the beam scanning speed. and the rate of change of the semi-major axis of the orbit The impact on beam pointing changes.
[0065] For example, the coupling relationship between orbital changes and beam pointing can be observed through the beam. The orbital coupling dynamics equations describe:
[0066] in, For beam direction, For the semi-major axis of the track, This is the influence coefficient of the orbit on the beam pointing, which is related to the orbital altitude and satellite attitude, and can be calculated using an orbital dynamics model. It can be approximated as a function related to the semi-major axis of the orbit, for example... ,in and It is a constant.
[0067] In some embodiments, the correlation between metasurface phase and beam pointing is realized through a metasurface beamforming model, which describes the mapping relationship between metasurface phase control array elements and beam pointing. By adjusting the phase distribution, beam pointing can be controlled quickly and accurately.
[0068] For example, the correlation between metasurface phase and beam pointing modulation can be described by a phase-shifting metasurface beamforming model:
[0069] in, For beam direction, For metasurface phase control array elements, The phase difference between adjacent metasurface elements. The wavelength of the current communication signal. For the metasurface beamforming function, the beam direction is proportional to the metasurface phase gradient.
[0070] In some embodiments, the dynamic relationship between thrust and orbital changes is realized through a satellite orbital dynamics model, which describes the physical constraints between the magnitude of satellite thrust, satellite mass and the rate of change of orbital parameters, and is the core basis for precise orbital adjustment.
[0071] For example, the dynamic relationship between thrust and orbital variation can be described based on a two-body orbital dynamics model:
[0072] in, For the semi-major axis of the track, The rate of change of the semi-major axis of the orbit. For the magnitude of satellite thrust, The mass of the low-Earth orbit satellite is given by G, where G is the gravitational constant and M is the mass of the Earth. This is the orbital dynamics function. The rate of change of the semi-major axis. With thrust magnitude Proportional to satellite mass Inversely proportional.
[0073] To achieve optimal coordinated control of beam, orbit, and thrust, this application integrates the beam-orbit coupled dynamic equation, the metasurface beamforming model, and the orbit dynamic model into a differentiable computational graph, enabling end-to-end gradient backpropagation and synchronous optimization of all control variables.
[0074] In some embodiments, the beam-orbit coupling dynamics equations, the metasurface beamforming model, and the orbital dynamics model are integrated into a differentiable computational graph to achieve end-to-end gradient backpropagation. A loss function is defined. For example, the optimization objectives are to minimize switching delay and maximize signal-to-interference-plus-noise ratio (SINR).
[0075]
[0076] Where, Δt( SINR is the link switching delay. ) represents the signal-to-interference-plus-noise ratio (SINR), and α and β are weighting coefficients used to balance the optimization objectives of latency and communication quality.
[0077] Based on the gradient descent algorithm, the metasurface phase control array element u and the thrust magnitude F are simultaneously optimized to minimize the loss function. :
[0078]
[0079] in, It's the learning rate. , These are the gradients of the loss function with respect to the metasurface phase parameters and thrust parameters, respectively.
[0080] In this embodiment, differentiable end-to-end joint optimization avoids the local optimum problem caused by traditional step-by-step optimization, and achieves global synchronous optimization of beam pointing, orbit parameters, and thrust magnitude, further reducing handover latency and improving communication link quality.
[0081] In existing technologies, traditional Bayesian dynamic programming (BDP) algorithms suffer from high computational complexity and low efficiency when handling continuous state spaces and high-dimensional spatial perturbations. Furthermore, they fail to incorporate contextual information such as beam pointing, orbital parameters, and thrust magnitude into the fuel optimization process, making it difficult to achieve refined fuel control. To address these issues, this application proposes a context-aware variational inference Bayesian dynamic programming fuel control algorithm. This algorithm deeply integrates link control and fuel control strategies, enabling more refined optimization of satellite fuel consumption while ensuring communication link quality.
[0082] In some embodiments, Figure 6 A flowchart illustrating the determination of the optimal fuel control strategy based on disturbance scenarios and orbital adjustment parameters is shown, such as... Figure 6 As shown, the process includes the following steps (S501-S502): S501. Generate disturbance samples corresponding to different disturbance scenarios based on the disturbance scenario generation model; This step utilizes a variational autoencoder (VAE) to learn spatial environmental perturbations. prior distribution This generates disturbance scenarios that conform to real-world spatial environments. During training, historical beam pointing is further incorporated. The orbital parameters a and thrust magnitude F are used to learn disturbance patterns that better reflect the system's operating conditions.
[0083] For example, the VAE model is composed of an encoder. and decoder The training is completed by minimizing the Evidence Lower Bound (ELBO) loss function.
[0084] After training, the decoder As a perturbation scene generator, it samples latent variables. Generate different disturbance scenarios Examples of disturbances include solar wind, atmospheric turbulence, and orbital perturbations.
[0085] S502. Based on each disturbance sample and the optimal beam pointing parameter, optimal orbit parameter, and optimal thrust parameter, determine the optimal fuel control strategy.
[0086] In this embodiment, the optimal fuel control strategy includes at least the minimum fuel consumption. Fuel consumption rate It may also include optimal fuel allocation strategies, orbit maintenance strategies, and disturbance adaptation strategies, etc., which are not limited in the embodiments of this application. In some embodiments, the orbit adjustment parameters further include: optimal fuel control strategy.
[0087] In some embodiments, step S502 includes the following steps: S5021. Input each disturbance sample, along with the optimal beam pointing parameter, optimal orbit parameter, and optimal thrust parameter, into the trained temporal feature learning network to predict the fuel consumption distribution features corresponding to different disturbance scenarios. In this embodiment, the fuel consumption is pre-established. Regarding contextual information and disturbance intensity Conditional probability model The optimal beam pointing parameters output in step S103 are... Optimal orbital parameters and optimal thrust parameters As contextual information The input time-series feature learning network, combined with the perturbation samples generated in step S501, predicts the fuel consumption distribution parameters. and .
[0088] For example, a fuel consumption probability model It follows a Gaussian distribution, but its parameters are context-aware.
[0089]
[0090] in, This represents the average fuel consumption. The variance of fuel consumption is determined by both the disturbance intensity and the context information, reflecting the context-aware characteristics of fuel prediction, i.e., its dependence on the output of step S103.
[0091] S5022. The expected approximation solution is performed on the fuel consumption distribution characteristics under different disturbance scenarios, and iterative optimization is performed based on the value function and the system state transition relationship, finally converging to obtain the optimal fuel control strategy.
[0092] In some embodiments, variational Bayesian dynamic programming is used to approximate the fuel consumption distribution and then iteratively optimizes the value function and state transition equation to obtain the optimal fuel control strategy.
[0093] For example, variational inference methods are used to approximate the posterior distribution of Bayesian dynamic programming. Define the value function The minimum fuel consumption in state s is represented by the state transition equation: The reward function is the negative value of fuel consumption. (Fuel mass change rate).
[0094] By minimizing the KL divergence Iteratively update the variational parameters and select the optimal fuel consumption that minimizes the desired fuel consumption. This maximizes the expected value function.
[0095]
[0096] in, This refers to solar wind pressure. As the reference pressure, These are the weighting coefficients.
[0097] Due to fuel consumption probability model The optimal fuel consumption is thus obtained by incorporating contextual information such as beam pointing, orbital parameters, and thrust magnitude. It also implicitly takes into account the beam pointing output in step S103. Orbital parameters And thrust magnitude This enables context-aware fuel optimization capabilities, allowing for deep integration and synergistic optimization of fuel control and link control strategies.
[0098] In this embodiment, context-aware modeling is used to beam the beam. track The integration of thrust joint control information into the fuel optimization process significantly improves the accuracy of fuel control under high-dimensional disturbance scenarios, reduces fuel consumption, and extends the satellite's on-orbit lifespan.
[0099] S104, low-orbit satellites perform beam calibration and orbit adjustment on communication links oriented towards the target area based on orbit adjustment parameters.
[0100] In this embodiment, the low-orbit satellite completes beam direction calibration based on the optimal beam pointing parameters and performs orbit fine-tuning based on the optimal orbit parameters and optimal thrust parameters, so that the communication link avoids obstructed areas and improves signal quality.
[0101] For example, a low-Earth orbit satellite receives the optimal target orbit from a high-Earth orbit satellite, uses it as the initial orbit, controls the reconfigurable metasurface array to adjust the beam direction according to the orbit adjustment parameters, and simultaneously controls the Hall electric propulsion system to perform minor orbit corrections, so that the beam is accurately pointed to the target area, thereby achieving rapid and stable reconfiguration of the communication link.
[0102] In some embodiments, the track adjustment parameters further include an optimal fuel control strategy. For example, the optimal fuel control strategy includes at least the minimum fuel consumption. With fuel consumption rate Low-Earth orbit satellites are designed based on optimal fuel consumption. With fuel consumption rate By constraining the magnitude and duration of thrust, fine-grained control of fuel consumption is simultaneously achieved during beam calibration and orbit adjustment, avoiding fuel waste and enabling coordinated execution of link reconfiguration and fuel optimization.
[0103] This embodiment achieves rapid link recovery under dynamic blockage by coordinating beam calibration, orbit adjustment, and optimal fuel control, reducing the probability of communication interruption, while also reducing satellite fuel consumption, extending satellite on-orbit lifespan, and improving the reliability and economy of the 6GNTN link.
[0104] In large-scale multi-satellite collaborative scenarios, independent training of each satellite node leads to problems such as data silos, privacy leaks, and weak model generalization capabilities. To address this issue, this application employs context-aware differential privacy federated learning to achieve distributed collaborative training and data privacy protection. Low-Earth orbit satellites complete model training based on local data, while high-Earth orbit satellites aggregate and distribute global models based on differential privacy, thus realizing context-aware differential privacy federated learning.
[0105] In some embodiments, Figure 7 A flowchart illustrating the context-aware differential privacy federated learning training process provided in an embodiment of this application is shown. Figure 7 As shown, the process includes the following steps (S601-S604): S601, low-Earth orbit satellites adjust local model parameters based on transmission parameters, orbit adjustment parameters, and optimal fuel control strategy; In this embodiment, the low-orbit satellite is based on the transmittance matrix obtained in step S101. Step S103 obtains the orbit adjustment parameters (surface beam pointing). and orbital parameters ) and optimal fuel control strategy (fuel consumption) (or fuel consumption rate) ), build a context-enhanced dataset and tune local model parameters; Each low-Earth orbit satellite node Above, construct a context-enhanced local training dataset. Unlike traditional federated learning, the local dataset here contains not only raw spatiotemporal data samples. It also incorporates contextual information from preceding steps. (or Therefore, a context-enhanced local dataset can be represented as .
[0106] For example, contextual information Integrating this into the calculation of ReconLoss and KL Divergence allows the loss function to be aware of the current environmental state and system operating parameters. One implementation method is to use contextual information. Dynamically adjust the weighting coefficients of reconstruction loss and KL divergence, or directly... As a moderating factor, it affects the specific calculation method of the loss function. The calculation formula is as follows:
[0107]
[0108] in, It is an adjustable hyperparameter. For contextual information, The average transmittance of the target area. For fuel consumption. The local training loss function becomes:
[0109] S602, low-Earth orbit satellites upload the trained local model parameters to high-Earth orbit satellites; Each low-Earth orbit satellite node uploads its own model parameters. .
[0110] S603: The high-orbit satellite aggregates the local model parameters uploaded by all low-orbit satellites and adds Gaussian noise to generate global model parameters. In each iteration of federated learning, the high-orbit satellite aggregates the model parameters uploaded by each low-orbit satellite node. To achieve differential privacy protection, a Gaussian noise mechanism is introduced during parameter aggregation.
[0111]
[0112] in, For local model parameters, (0,σ²) represents Gaussian noise, used to protect data privacy.
[0113] S604: The high-orbit satellites distribute global model parameters to each low-orbit satellite, controlling the low-orbit satellites to continue local training based on the global model parameters.
[0114] The high-orbit satellites will aggregate the global model parameters. The data is distributed to each low-Earth orbit (LEO) satellite node, and each LEO satellite performs local training based on the global model parameters in the next round. Conduct training.
[0115] In this embodiment, a multi-satellite node collaborative optimization model is achieved, while differential privacy is used to prevent the leakage of sensitive data, thereby improving the model's generalization ability and system security.
[0116] Traditional Lyapunov stability analysis requires manually constructing functions, which is difficult to adapt to complex nonlinear satellite systems, and is computationally intensive and has poor real-time performance. To address this issue, in some embodiments, before step S104, i.e. before the low-Earth orbit satellite performs beam calibration and orbit adjustment, the method provided in this application also includes a system stability determination and rollback control process.
[0117] Figure 8 The flowchart illustrating the adaptive Lyapunov neural differential stability verification and rollback control provided in an embodiment of this application is shown. Figure 8As shown, the process includes the following steps (S701–S705): S701. The transmission parameters, orbit adjustment parameters, optimal fuel control strategy and global model parameters corresponding to the target area are used as a high-dimensional state vector to characterize the satellite's operating status. For example, a Lyapunov Neural Network (LNN) is constructed, denoted as... The high-dimensional state vector is represented as:
[0118] Each dimension is derived from the output data of the preceding steps. These are the parameters of the neural network.
[0119] S702. Input the high-dimensional state vector into the pre-trained Lyapunov neural network to obtain the corresponding Lyapunov function values. For example, the training objective is to make the Lyapunov neural network satisfy the stability condition: when ; ; and (or ).
[0120] The corresponding loss function is:
[0121] in, It is the data distribution in the state space. It is a positive constant. It is the derivative of the Lyapunov function along the system trajectory, obtained through differential calculation, and the gradient descent algorithm is used to optimize the neural network parameters. Minimize the loss function .
[0122] S703. The rate of change of the Lyapunov function along the system's trajectory is calculated using a neural differential equation model. For example, constructing a neural differential equation to validate a model describing the dynamic evolution of the system state:
[0123] in, It is a system dynamics model learned by a neural network. These are known physical constraints. It is random noise.
[0124] S704. Compare the rate of change with a preset stability threshold. If the rate of change is greater than the preset stability threshold, determine that the system is unstable and proceed to step S705. Otherwise, determine that the system is stable and proceed to step S104. For example, define a stability boundary threshold. Real-time calculation of the rate of change of the Lyapunov function When satisfied When this happens, the system is determined to be in an unstable state.
[0125] In some embodiments, this application further employs a dynamic adaptive rollback threshold to further reduce the false trigger rate. The dynamic rollback threshold is: Parameters are adaptively optimized through online reinforcement learning. , This makes stability assessment more accurate and reduces error rollbacks.
[0126] S705, trigger the rollback control strategy to adjust the orbit adjustment parameters of the low-orbit satellite to the preset safe operating parameters.
[0127] For example, when the system is determined to be unstable, the autonomous rollback control strategy is immediately triggered to restore the satellite's orbital parameters, beam pointing parameters, and thrust parameters to a safe operating state, ensuring that the system does not experience instability or loss of control.
[0128] In this embodiment, there is no need to manually construct Lyapunov functions, which enables real-time, efficient, and automatic stability assessment and fault rollback for complex satellite systems, significantly improving the reliability of system operation.
[0129] The multi-track satellite cooperative communication link optimization method provided in this application improves the accuracy of terahertz band obstruction detection and reduces the signal attenuation false positive rate by adopting multimodal frequency domain attention sensing; it achieves ultra-high-speed orbit optimization by using quantum reinforcement learning to meet the sub-millisecond handover and URLLC requirements of 6G NTN; and it realizes beam, orbit, and thrust joint optimization to reduce link handover latency and improve dynamic environment adaptability. Thus, it can effectively solve problems such as dynamic cloud obstruction, terahertz signal attenuation, excessive handover latency, excessive fuel consumption, and insufficient system stability, and significantly improve the reliability and real-time performance of 6G NTN communication links.
[0130] like Figure 1The multi-track satellite cooperative communication link optimization system shown can implement all or part of the steps in the above method embodiments. The system structure and functions are described below; for any matters not covered herein, please refer to the foregoing method embodiments. This embodiment of the multi-track satellite cooperative communication link optimization system corresponds to the above-described multi-track satellite cooperative communication link optimization method embodiments. All implementation processes and methods of the above method embodiments can be applied to this embodiment of the multi-track satellite cooperative communication link optimization system and can achieve the same technical effects.
[0131] In some embodiments, such as Figure 1 As shown, the system includes at least a high-orbit satellite and multiple low-orbit satellites. In this embodiment, the high-orbit satellite is used to acquire multimodal sensing data of the target area and obtain the transmission parameters corresponding to the target area based on the multimodal sensing data; wherein, the multimodal sensing data includes at least one of the following: frequency domain data of communication signals and spatial domain imaging data of physical fields; the transmission parameters are used to characterize the degree of attenuation of communication signals caused by obstructions; the high-orbit satellite is also used to select the optimal target orbit from multiple satellite orbits covering the target area based on the current environmental state using a quantum reinforcement learning strategy, and to distribute the optimal target orbit to the low-orbit satellites serving the target area; wherein, the current environmental state is obtained based on the transmission parameters corresponding to the target area; each low-orbit satellite serving the target area is used to determine the orbit adjustment parameters based on the coupling relationship between orbit change and beam pointing, the modulation relationship between metasurface phase and beam pointing, and the dynamic relationship between thrust and orbit change, using the optimal target orbit as the initial orbit; wherein, the orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters. Low-Earth orbit satellites are also used to perform beam calibration and orbit adjustment for communication links targeting specific regions, based on orbit adjustment parameters and initial orbits.
[0132] In some embodiments, a high-orbit satellite is used to obtain transmission parameters corresponding to a target area based on multimodal sensing data, including: performing frequency domain transformation on the multimodal sensing data to obtain corresponding frequency domain features; learning attention weights for different frequency domain features using a frequency domain attention mechanism; performing weighted fusion based on the attention weights and frequency domain features to obtain fused frequency domain features; and determining the transmission parameters corresponding to the target area based on the fused frequency domain features.
[0133] In some embodiments, the multiple satellite orbits covering the target area are pre-defined candidate orbits, and each candidate orbit can achieve signal coverage of the target area.
[0134] In some embodiments, the high-orbit satellite employs a quantum reinforcement learning strategy to select the optimal target orbit from multiple satellite orbits covering the target area, including: encoding multiple satellite orbits as quantum superposition states, iteratively updating the quantum superposition states through quantum circuits; performing quantum measurements based on the updated quantum superposition states to obtain the orbit probability distribution of each satellite orbit; and selecting the optimal target orbit from multiple satellite orbits according to the orbit probability distribution.
[0135] In some embodiments, the high-orbit satellite is also used to calculate the reward value corresponding to the optimal target orbit based on link communication quality, switching delay, orbit adjustment speed and transmission parameters; determine the fidelity and quantum importance of the updated quantum superposition state; and, if both the fidelity and quantum importance are greater than a preset threshold, store the updated quantum superposition state, the optimal target orbit and the reward value as the current interaction data in the experience set.
[0136] In some embodiments, the high-orbit satellite is also used to extract historical interaction data from an experience set and use the extracted historical interaction data to update the quantum gate parameters in the quantum circuit, so as to iteratively update the quantum superposition state through the quantum circuit.
[0137] In some embodiments, the low-Earth orbit satellite is further configured to generate disturbance samples corresponding to different disturbance scenarios based on a disturbance scenario generation model after the low-Earth orbit satellite determines its orbit adjustment parameters and before performing beam calibration and orbit adjustment; and to determine an optimal fuel control strategy based on each disturbance sample and the optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters. The optimal fuel control strategy includes at least the minimum fuel consumption, and the orbit adjustment parameters also include the optimal fuel control strategy.
[0138] In some embodiments, the low-Earth orbit satellite is further configured to determine the optimal fuel control strategy based on each disturbance sample and the optimal beam pointing parameter, optimal orbit parameter, and optimal thrust parameter. This includes: inputting each disturbance sample and the optimal beam pointing parameter, optimal orbit parameter, and optimal thrust parameter into a trained temporal feature learning network to predict the corresponding fuel consumption distribution features under different disturbance scenarios; performing an expected approximation solution on the corresponding fuel consumption distribution features under different disturbance scenarios, and iteratively optimizing based on the value function and the system state transition relationship, ultimately converging to obtain the optimal fuel control strategy.
[0139] In some embodiments, low-Earth orbit (LEO) satellites and high-Earth orbit (HEO) satellites are also used to perform federated learning: LEO satellites adjust local model parameters based on transmission parameters, orbit adjustment parameters, and optimal fuel control strategies, and upload the local model parameters to HEO satellites; HEO satellites aggregate the local model parameters uploaded by all LEO satellites, add Gaussian noise, generate global model parameters, and distribute the global model parameters to LEO satellites to control the LEO satellites to perform local training based on the global model parameters.
[0140] In some embodiments, before performing beam calibration and orbit adjustment, the low-Earth orbit satellite is further configured to use the transmission parameters, orbit adjustment parameters, optimal fuel control strategy, and global model parameters corresponding to the target area as a high-dimensional state vector characterizing the satellite's operating state; input the high-dimensional state vector into a pre-trained Lyapunov neural network to obtain the corresponding Lyapunov function values; calculate the rate of change of the Lyapunov function along the system's operating trajectory using a neural differential equation model; if the rate of change is greater than a preset stability threshold, determine that the system is unstable, trigger a rollback control strategy, and adjust the low-Earth orbit satellite's orbit adjustment parameters to preset safe operating parameters.
[0141] The multi-orbit satellite cooperative communication link optimization system provided in this application improves the accuracy of terahertz band obstruction detection and reduces the signal attenuation false positive rate by adopting multimodal frequency domain attention sensing; it uses quantum reinforcement learning to achieve ultra-high-speed orbit optimization to meet the sub-millisecond switching and URLLC requirements of 6GNTN; and it achieves joint optimization of beam, orbit, and thrust to reduce link switching latency and improve dynamic environment adaptability. Thus, it can effectively solve problems such as dynamic cloud obstruction, terahertz signal attenuation, excessive switching latency, excessive fuel consumption, and insufficient system stability, and significantly improve the reliability and real-time performance of 6GNTN communication links.
[0142] In some embodiments, the multi-orbit satellite cooperative communication link optimization system provided in this application has an innovative hardware and software architecture, as detailed below: (a) Hardware Architecture The high-orbit satellite serves as the system's global perception center, quantum decision-making center, and federated learning aggregation center. For example, the high-orbit satellite carries the following core hardware components: (1) A terahertz holographic imager is used for global terahertz and infrared imaging perception, acquiring high-resolution terahertz holographic image data and infrared data in real time. Its operating frequency band covers terahertz, enabling it to accurately capture the influence of atmospheric composition, clouds, etc., on terahertz signals, providing the most original and accurate data input for subsequent obstruction perception modeling. This imager has high sensitivity and high resolution, ensuring that the perception layer acquires high-quality environmental information.
[0143] (2) Quantum computing coprocessor, used to accelerate quantum reinforcement learning orbit decision-making; this coprocessor is not a general quantum computer, but a hardware acceleration optimization for quantum reinforcement learning algorithms. It utilizes the parallelism of quantum computing to greatly accelerate the search and decision-making process of the optimal orbit scheme, and meets the real-time scheduling requirements in ultra-dense constellation scenarios.
[0144] (3) Large-capacity onboard storage and high-speed processing unit for data storage, model aggregation and command issuance. High-orbit satellites need to store a large amount of sensing data, orbital state data and model parameters, and perform complex algorithm calculations. Therefore, it is necessary to equip them with large-capacity onboard storage and high-performance multi-core processors to support onboard data storage, preprocessing and collaborative computing with quantum computing coprocessors.
[0145] Medium-Earth orbit (MEO) satellites serve as regional coverage nodes, service relay nodes, and edge inference nodes for this system. For example, MEO satellites may carry the following core hardware components: (1) Phased array antenna, which can support ±60° multi-beam scanning. MEO satellites no longer use metasurfaces, but instead use the mature phased array antenna. The phased array antenna has multi-beam shaping capability, which can simultaneously serve multiple ground users or LEO satellites, achieving enhanced coverage and capacity improvement within the region. The ±60° scanning range allows it to flexibly adjust the beam direction to adapt to the dynamic changes of users in the region.
[0146] (2) Onboard service switching and routing unit, used for regional service aggregation and forwarding. MEO satellites undertake the important function of regional service aggregation and forwarding. For this purpose, it is necessary to equip high-performance onboard service switching and routing units with Tbps-level switching capacity, which can efficiently process service data of user terminals and LEO satellites in the region, and then aggregate them and transmit them to GEO satellites or ground core networks through inter-satellite links.
[0147] (3) Medium-power edge AI inference module for regional data analysis and QoS assurance. MEO satellites also have certain edge AI inference capabilities, but their computing power is between GEO and LEO, and they are positioned as "medium-power". This module can support regional data analysis and business optimization, such as regional user behavior prediction, localized QoS assurance, and coordination of regional federated learning.
[0148] Low-Earth orbit (LEO) satellites serve as link execution nodes, beam control, and attitude adjustment nodes for this system. For example, an LEO satellite may carry the following core hardware components: (1) Reconfigurable metasurface arrays are used for rapid and precise beamforming. The array can achieve rapid and precise scanning and shaping of beams, meeting the requirements of high-frequency communication for beam flexibility. Its "reconfigurable" characteristic means that beam parameters can be dynamically adjusted according to software instructions, realizing coordinated control of beam pointing and track parameters.
[0149] (2) Hall thruster system, used for high-precision orbit adjustment and thrust output. This system has the characteristics of high specific impulse, which can achieve a large speed increment with a small amount of fuel consumption, meeting the needs of precise adjustment and maintenance of satellite orbit. By receiving commands from ground or high-orbit satellites, the Hall thruster system can precisely adjust the satellite's orbital semi-major axis, inclination and other parameters, and in conjunction with metasurface beamforming, realize rapid switching and optimization of the link.
[0150] (3) Lightweight edge AI inference module for local model training and data preprocessing. This module has computing power units sufficient to support local inference of lightweight AI models such as spatiotemporal capsule networks, participate in the distributed training process of federated learning, and perform preliminary analysis and feature extraction of local data to support global decision-making for high-orbit satellites.
[0151] like Figure 1 As shown in the embodiments of this application, the system for optimizing multi-track satellite cooperative communication links further includes: a ground terminal, which serves as the entry point for users to access the network and as the interaction interface between the ground control center and the satellite network. Its hardware configuration needs to meet the following requirements: (1) 6GNTN-compatible radio frequency front-end (Sub-6GHz / terahertz dual-mode): In order to achieve seamless connection with satellite networks, ground terminals need to be equipped with radio frequency front-ends compatible with the 6GNTN protocol. More importantly, in order to support high-frequency communication, the radio frequency front-end needs to have the capability of both Sub-6GHz and terahertz dual-mode operation, and can flexibly switch the operating frequency band according to the actual channel conditions and service requirements to make full use of the spectrum resources of 6GNTN.
[0152] (2) Edge AI inference module: Some ground terminals, especially edge computing nodes, may also be equipped with an edge AI inference module. This module can assist in completing some data preprocessing, user behavior analysis, and localized QoS assurance functions, thereby reducing the computing pressure on the satellite and improving the overall efficiency of the system.
[0153] To ensure efficient collaboration and information exchange among the various hardware components, the multi-track satellite cooperative communication link optimization system provided in this application, for example, employs the following key connection technologies: (1) Laser inter-satellite link between high-orbit and low-orbit satellites: High-orbit and low-orbit satellites are connected by a high-speed laser inter-satellite link. The laser inter-satellite link has the advantages of high transmission rate, low latency and strong anti-interference capability, which can meet the needs of high-orbit satellites to transmit massive amounts of sensing data, control commands and model parameters to low-orbit satellites, and support the efficient realization of key functions such as joint beam-orbit control.
[0154] (2) Laser inter-satellite link between high-orbit satellites and medium-orbit satellites: High-speed laser inter-satellite link continues to be used between high-orbit satellites and medium-orbit satellites to transmit global perception data, macro scheduling instructions, and aggregate business data at the MEO level.
[0155] (3) Ka / V band inter-satellite links between medium-Earth orbit (MEO) satellites and low-Earth orbit (LEO) satellites: Ka / V band radio frequency inter-satellite links are used between MEO and LEO satellites. Ka / V band links are low-cost, easy to deploy, and can meet the needs of MEO to transmit regional scheduling instructions, service data and conduct inter-satellite collaborative measurements to LEO.
[0156] (4) User terminals and LEO / MEO satellite access networks adopt the Open Radio Access Network (O-RAN) protocol (Sub-6GHz / Terahertz dual-mode): User terminals can flexibly choose to access the LEO or MEO satellite network according to actual conditions. For low-latency services, LEO satellite access is given priority; for services requiring enhanced coverage or higher capacity, MEO satellite access can be used. The terrestrial access protocol still adopts the O-RAN protocol to maintain the flexibility and standardization of terrestrial terminal access.
[0157] Those skilled in the art will understand that the embodiments of this application are based on Figure 1 The illustrated method is provided as an example to introduce the multi-track satellite cooperative communication link optimization system, and is merely an illustrative example and does not constitute a limitation on this application. This application is applicable to 6G NTN integrated air-space-ground networks and can be widely used in scenarios such as remote area communication, marine communication, aviation communication, emergency communication, vehicle networking, and the Internet of Things, possessing extremely high practical value and commercial prospects. Through the above system, a closed-loop process of global perception, hierarchical decision-making, joint control, and intelligent verification can be achieved, effectively solving problems such as dynamic cloud obstruction, terahertz signal attenuation, excessively long switching delays, excessive fuel consumption, and insufficient system stability, significantly improving the reliability and real-time performance of 6G NTN communication links.
[0158] (II) Software Architecture To support the hardware architecture and achieve end-to-end coordination of perception, decision-making, control, and verification, this application also constructs a layered, closed-loop system software architecture, including a perception layer, a decision-making layer, a control layer, and a verification layer. Data interaction and command flow between these layers are achieved through inter-satellite links. The four-layer closed-loop software architecture of this multi-orbit satellite collaborative communication link optimization system is as follows: (1) Perception layer: global macro perception module (DM-CapsNet-Global) and regional fine perception module (DM-CapsNet-Regional) to realize multi-physics field coupled occlusion perception; The Global Macroscopic Sensing Module is the sensing layer software on GEO satellites, focusing on global macroscopic environmental perception. It employs a global dual-modal capsule network, utilizing terahertz holographic imager data and a multiphysics coupled database to generate macroscopic environmental situation maps, such as global cloud distribution and atmospheric composition concentration distribution, providing global background information for regional perception by MEO and LEO satellites.
[0159] The regional fine-sensing module is the sensing layer software on MEO and LEO satellites, focusing on fine-grained regional environmental perception. It employs a regional dual-modal capsule network, utilizing its onboard sensors (LEO primarily relies on ground terminal feedback, while MEO may carry miniaturized meteorological sensors), combined with global environmental information provided by GEO satellites, to perform fine-grained environmental perception within a specific region. This includes details such as cloud cover and local atmospheric turbulence intensity, providing more accurate environmental information for local link reconstruction and resource scheduling.
[0160] (2) Decision-making layer: global quantum decision-making center, regional resource scheduler, local link selector, to realize track and beam collaborative decision-making. The Global Quantum Decision Center (Quantum-RL-Global-Scheduler) is deployed on GEO high-orbit satellites and employs a quantum reinforcement learning scheduler to be responsible for macro-level resource scheduling and task allocation across the three-layer satellite constellation. For example, based on global user service requirements, satellite load status, and environmental conditions, it makes macro-level strategic decisions such as task allocation between MEO and LEO satellites, orbital collaborative planning, and cross-layer resource reservation, leveraging the parallelism of quantum computing to achieve ultra-high-speed orbit matching and decision output.
[0161] The Classical-RL Regional Scheduler is deployed on MEO (Medium Earth Orbit) satellites and employs a classic reinforcement learning-based regional scheduler to optimize resource scheduling within its designated region. For example, based on regional user service needs and the distribution of LEO satellites, it makes regional decisions such as beamforming optimization for MEO satellites, user access control, and collaborative resource scheduling with LEO satellites to ensure regional coverage and service carrying capacity.
[0162] The Greedy-Link-Selector, deployed on LEO (Low Earth Orbit) satellites, employs a greedy link selection algorithm and is primarily responsible for access control and link selection for local user terminals. For example, it selects the optimal beam and frequency band for communication based on the local user terminal's channel quality, service type, and QoS requirements, or adjusts local resource configuration according to MEO satellite scheduling instructions to ensure low-latency, high-reliability link access.
[0163] (3) Control layer: macro-instruction controller and joint control middleware to realize beamforming track thrust Fuel Joint Control The Global Command Controller is deployed on GEO high-orbit satellites and is responsible for converting the macro-scheduling commands output by the global quantum decision center into specific control commands, which are then sent to the joint control middleware of MEO and LEO satellites.
[0164] The Regional-Controller (DDS) middleware is deployed on MEO (Medium Earth Orbit) and LEO (Low Earth Orbit) satellites, enabling cross-layer and cross-satellite collaborative control based on the Data Distribution Service (DDS) protocol. Specifically, the DDS middleware for MEO satellites executes control commands generated by the regional resource scheduler and coordinates the control of the MEO satellite's own hardware resources. The DDS middleware for LEO satellites primarily receives commands from GEO and MEO satellites, executes control strategies generated by the local link selector, and coordinates the control of LEO satellite's reconfigurable metasurface array, Hall thruster system, and other actuators, achieving unified and coordinated control of beam, orbit, thrust, and fuel.
[0165] (4) Verification layer: global stability monitor and regional stability verifier to realize real-time stability judgment and rollback control. The Global-Stability-Monitor is deployed on GEO satellites. Based on a neural differential equation solver and a digital twin verification sandbox, it performs global monitoring and stability assessment of the overall operation status of the three-layer satellite constellation. It is responsible for global stability determination, risk warning, and triggering global rollback strategies when necessary.
[0166] The Regional Stability Verifier is deployed on MEO medium-Earth orbit satellites and LEO low-Earth orbit satellites. It uses a lightweight neural differential equation solver to perform regional real-time monitoring and evaluation of the operational status of satellites in its region. It is responsible for regional stability determination and risk warning, and executes regional rollback control according to the global commands of GEO satellites to quickly restore satellite attitude, orbit, beam and other parameters to a safe operating range.
[0167] Software call relationships and data flow: Software implementation of four-layer closed-loop control.
[0168] The modules in the software architecture do not operate in isolation, but rather collaborate according to a closed-loop control process of "perception-decision-execution-verification." Data flows between layers, and instructions are transmitted between layers, forming an organic whole. (I) Layered Sensing Data Flow: GEO satellites perceive the global macroscopic environment, while MEO and LEO satellites perceive the fine-grained regional environment. Sensing data flows between satellites at different levels, forming a multi-scale, multi-precision, three-dimensional environmental sensing system. GEO's global sensing data provides background information for MEO and LEO; while MEO and LEO's regional sensing data provides local detail supplements for GEO.
[0169] (II) Hierarchical Decision-Making Command Flow: Global policy decision-making commands from GEO satellites are issued to MEO and LEO satellites. MEO satellites, guided by global commands, perform regional resource scheduling; LEO satellites, guided by both global and regional commands, select local links. Decision-making commands are transmitted between satellites at different levels, forming a hierarchical decision-making system that combines macro and micro perspectives.
[0170] (III) Layered Execution Control Flow: The macro-command controller of GEO satellites, the joint control middleware of MEO and LEO satellites, and the satellite hardware execution units together constitute a multi-layered collaborative execution control system. Control commands are transmitted layer by layer between the software and hardware layers, achieving precise control of the three-layer satellite constellation.
[0171] (IV) Layered verification feedback flow: The global stability monitor of GEO satellites and the regional stability verifier of MEO / LEO satellites monitor and evaluate the global and regional system stability respectively. The verification results are fed back between satellites at different levels, forming a risk monitoring and rollback control system that combines global and regional aspects.
[0172] (V) Cross-layer collaboration: The federated learning process is carried out collaboratively in a three-layer satellite constellation. GEO satellites can serve as the central nodes of federated learning, aggregating model parameters from MEO and LEO satellites for global model updates and optimization. Cross-layer collaboration is also reflected in resource scheduling, task allocation, stability assurance, and other aspects. The three layers of satellites cooperate with each other to complete complex communication tasks.
[0173] Through the organic combination and collaborative work of the above hardware and software architectures, this application embodiment constructs a highly intelligent, highly adaptive, safe and reliable 6G NTN multi-track satellite collaborative communication link optimization system, laying a solid technical foundation for the future development of integrated air-space-ground communication.
[0174] Figure 9 A structural block diagram of an electronic device 1000 illustrating an exemplary embodiment of this application is shown. This electronic device 1000 can be implemented as the multi-track satellite cooperative communication link optimization method described above.
[0175] Typically, electronic device 1000 includes a processor 1001 and a memory 1002.
[0176] Processor 1001 may include one or more processing cores, such as a quad-core processor, a deca-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0177] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 is used to store at least one instruction, which is executed by the processor 1001 to implement all or part of the steps in the multi-orbit satellite cooperative communication link optimization method shown in the method embodiments of this application.
[0178] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the electronic device 1000, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0179] In one exemplary embodiment, a readable storage medium is also provided, which stores a program or instructions that, when executed by a processor, implement all or part of the steps in the multi-orbit satellite cooperative communication link optimization method described above. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.
[0180] In one exemplary embodiment, a computer program product is also provided, comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform all or part of the steps of the above-described multi-track satellite cooperative communication link optimization method.
[0181] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0182] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for optimizing multi-track satellite cooperative communication links, characterized in that, include: A high-orbit satellite acquires multimodal sensing data of a target area and obtains transmission parameters corresponding to the target area based on the multimodal sensing data; wherein the multimodal sensing data includes at least one of the following: frequency domain data of communication signals and spatial domain imaging data of physical fields; the transmission parameters are used to characterize the degree of attenuation of communication signals caused by obstructions; Based on the current environmental state, the high-orbit satellite uses a quantum reinforcement learning strategy to select the optimal target orbit from multiple satellite orbits covering the target area, and then distributes the optimal target orbit to the low-orbit satellites serving the target area; wherein, the current environmental state is obtained based on the transmission parameters corresponding to the target area; The low-Earth orbit satellite uses the optimal target orbit as its initial orbit. Based on the coupling relationship between orbit change and beam pointing, the relationship between metasurface phase and beam pointing adjustment, and the dynamic relationship between thrust and orbit change, the orbit adjustment parameters are determined. The orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters. The low-orbit satellite performs beam calibration and orbit adjustment on the communication link facing the target area based on the orbit adjustment parameters.
2. The method according to claim 1, characterized in that, The step of obtaining the transmission parameters corresponding to the target region based on the multimodal sensing data includes: The corresponding frequency domain features are obtained by performing frequency domain transformation on the multimodal sensing data; A frequency domain attention mechanism is used to learn the attention weights for different frequency domain features; The fused frequency domain features are obtained by weighting and fusing the attention weights and the frequency domain features. The transmission parameters corresponding to the target region are determined based on the fused frequency domain characteristics.
3. The method according to claim 1, characterized in that, The process of using a quantum reinforcement learning strategy to select the optimal target orbit from multiple satellite orbits covering the target region includes: The orbits of multiple satellites are encoded as quantum superposition states, and the quantum superposition states are iteratively updated through quantum circuits. The orbital probability distribution of each satellite orbit is obtained by quantum measurement based on the updated quantum superposition state; The optimal target orbit is selected from multiple satellite orbits based on the orbit probability distribution.
4. The method according to claim 3, characterized in that, The method further includes: The reward value corresponding to the optimal target orbit is calculated based on the link communication quality, switching delay, orbit adjustment speed, and the transmission parameters. Determine the fidelity and quantum importance of the updated quantum superposition state; When both the fidelity and the quantum importance are greater than a preset threshold, the updated quantum superposition state, the optimal target orbit, and the reward value are used as the current interaction data and stored in the experience set.
5. The method according to claim 4, characterized in that, The method further includes: Historical interaction data is extracted from the experience set, and the extracted historical interaction data is used to update the quantum gate parameters in the quantum circuit, so as to iteratively update the quantum superposition state through the quantum circuit.
6. The method according to claim 1, characterized in that, After the low-Earth orbit satellite determines its orbit adjustment parameters but before performing beam calibration and orbit adjustment, the method further includes: Based on the disturbance scenario generation model, disturbance samples corresponding to different disturbance scenarios are generated; Based on each of the disturbance samples and the optimal beam pointing parameter, the optimal orbit parameter, and the optimal thrust parameter, an optimal fuel control strategy is determined. The optimal fuel control strategy includes at least the minimum fuel consumption. The orbit adjustment parameter also includes the optimal fuel control strategy.
7. The method according to claim 6, characterized in that, The step of determining the optimal fuel control strategy based on each of the disturbance samples, the optimal beam pointing parameter, the optimal orbit parameter, and the optimal thrust parameter includes: Each disturbance sample, along with the optimal beam pointing parameter, the optimal orbit parameter, and the optimal thrust parameter, is input into a pre-trained temporal feature learning network to predict the fuel consumption distribution features corresponding to different disturbance scenarios. The expected approximation of fuel consumption distribution characteristics under different disturbance scenarios is obtained, and the optimal fuel control strategy is finally converged based on the value function and the relationship between system state transition.
8. The method according to claim 6, characterized in that, The method further includes: The low-orbit satellite adjusts its local model parameters based on the transmission parameters, the orbit adjustment parameters, and the optimal fuel control strategy, and uploads the local model parameters to the high-orbit satellite. The high-orbit satellite aggregates the local model parameters uploaded by all low-orbit satellites, adds Gaussian noise to generate global model parameters, and then sends the global model parameters to the low-orbit satellites to control the low-orbit satellites to perform local training based on the global model parameters.
9. The method according to claim 8, characterized in that, Before the low-Earth orbit satellite performs beam calibration and orbit adjustment, the method further includes: The transmission parameters corresponding to the target area, the orbit adjustment parameters, the optimal fuel control strategy, and the global model parameters are used as a high-dimensional state vector characterizing the satellite's operational status. The high-dimensional state vector is input into a pre-trained Lyapunov neural network to obtain the corresponding Lyapunov function values; and the rate of change of the Lyapunov function along the system's trajectory is calculated using a neural differential equation model. If the rate of change is greater than a preset stability threshold, the system is determined to be unstable, and a rollback control strategy is triggered to adjust the orbit adjustment parameters of the low-orbit satellite to preset safe operating parameters.
10. A multi-track satellite cooperative communication link optimization system, characterized in that, At least including: High-orbit satellites and multiple low-orbit satellites; The high-orbit satellite is used to acquire multimodal sensing data of the target area and obtain the transmission parameters corresponding to the target area based on the multimodal sensing data; wherein, the multimodal sensing data includes at least one of the following: communication signal frequency domain data and physical field spatial domain imaging data; the transmission parameters are used to characterize the degree of communication signal attenuation caused by obstructions; The high-orbit satellite is also used to select the optimal target orbit from multiple satellite orbits covering the target area based on the current environmental state using a quantum reinforcement learning strategy, and to distribute the optimal target orbit to the low-orbit satellites serving the target area; wherein, the current environmental state is obtained based on the transmission parameters corresponding to the target area; Each of the low-Earth orbit satellites serving the target area is used to determine orbit adjustment parameters based on the optimal target orbit as the initial orbit, the coupling relationship between orbit change and beam pointing, the control relationship between metasurface phase and beam pointing, and the dynamic relationship between thrust and orbit change, wherein the orbit adjustment parameters include at least one of the following: optimal beam pointing parameters, optimal orbit parameters, and optimal thrust parameters. The low-orbit satellite is also used to perform beam calibration and orbit adjustment on the communication link facing the target area based on the orbit adjustment parameters and the initial orbit.