Aviation heterogeneous network airborne dynamic switching method based on Q learning
By employing a Q-learning-based airborne dynamic handover method for heterogeneous aviation networks, and utilizing GRU neural networks to predict aircraft trajectories and hierarchical analysis for link quality assessment, the problem of frequent handover in aviation communication systems within heterogeneous networks is solved, achieving efficient and reliable communication link management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional aviation communication systems have bottlenecks in terms of capacity, speed, and coverage, which cannot meet the future development needs of civil aviation. Furthermore, frequent switching in heterogeneous aviation networks leads to communication interruptions and a decline in service quality.
A Q-learning-based airborne dynamic handover method for heterogeneous aviation networks is adopted. By predicting the aircraft trajectory through a GRU neural network and combining the analytic hierarchy process (AHP) and the Q-learning framework, link quality prediction and comprehensive evaluation are achieved, enabling intelligent and seamless link handover decisions.
It significantly reduces the risks of erroneous handover and handover lag, improves decision reliability and communication continuity, adapts to complex and ever-changing network environments, and reduces the probability of communication interruption and handover failure.
Smart Images

Figure CN121842785A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aviation communication technology, specifically to an airborne dynamic handover method for heterogeneous aviation networks based on Q-learning. Background Technology
[0002] With the continuous growth of global air transport demand, the modern civil aviation industry has increasingly higher requirements for communication systems, expecting to achieve seamless connectivity with high speed, large bandwidth, low latency, high reliability and wide coverage across the entire air route. However, the existing aviation communication system faces many challenges.
[0003] Traditional ground-based aviation communications have long relied on Very High Frequency (VHF) base stations. However, limited by spectrum resources and technical systems, VHF suffers from narrow bandwidth and low speeds, making it difficult to support high-bandwidth services such as Internet protocols and multimedia. For ocean areas, polar regions, and remote airspace, communications primarily rely on Geostationary Earth Orbit (GEO) satellites. While GEO satellites offer wide-area coverage, their high orbit of approximately 36,000 kilometers results in significant propagation delays, high path losses, and limited effective throughput, failing to meet the demands of new aviation communication services requiring high real-time data transmission, such as high-quality video transmission and real-time aircraft status monitoring. Therefore, the inherent technical bottlenecks in capacity, speed, and coverage of traditional aviation communication systems can no longer meet the future development needs of civil aviation.
[0004] To overcome the aforementioned bottlenecks, modern aeronautical communication systems are evolving towards a multi-layered heterogeneous network architecture that deeply integrates terrestrial and space communication networks. This architecture, while retaining existing VHF and GEO satellite systems, introduces the L-band Digital Aeronautical Communications System (LDACS) and Low Earth Orbit (LEO) satellite communication systems. LDACS, as a next-generation high-speed terrestrial data link, complements the data transmission limitations of VHF; while the LEO satellite system, with its low-orbit characteristics, provides lower latency and higher data rates for space-based communication than GEO satellites. By integrating VHF, LDACS, GEO, and LEO communication nodes, a comprehensive, three-dimensional, wide-coverage, and high-capacity aeronautical communication service system can be constructed.
[0005] However, this multi-layered heterogeneous network architecture also brings new technical challenges. The high-speed mobility of aircraft, the rapid movement of LEO satellites, the relative stationary nature of GEO satellites, and the fixed deployment of ground base stations together constitute an extremely complex spatiotemporal dynamic network environment, resulting in continuous changes in the end-to-end network topology. In this environment, aircraft need to frequently perform complex switching operations between different networks, different satellites, and between terrestrial and ground networks to maintain communication continuity.
[0006] Traditional handover decision-making methods, such as simple threshold strategies based on received signal strength, typically rely solely on instantaneous network state parameters. In the highly dynamic environment of heterogeneous aviation networks, these methods ignore future trends in link state and the overall network resource status, easily leading to problems such as frequent handovers, handover failures, communication interruptions, or degraded service quality. Therefore, there is an urgent need for an intelligent, efficient, and seamless link handover method that can predict link quality, comprehensively assess network state, and make optimal decisions to ensure the reliability and service quality of future aviation communications. Summary of the Invention
[0007] To address the problems existing in the background technology, this invention proposes an airborne dynamic handover method for heterogeneous aviation networks based on Q-learning. This method aims to solve the problems of frequent handover, unstable links, and degraded service quality caused by high-speed aircraft operation and differences in node characteristics in highly dynamic heterogeneous aviation networks, thereby enabling intelligent, seamless, and reliable handover of communication links for civil aircraft in all air routes.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: A Q-learning-based airborne dynamic handover method for heterogeneous aerospace networks, applied to heterogeneous aerospace networks including VHF base stations, L-band digital aerospace communication system base stations, geostationary orbit satellite systems, and low Earth orbit satellite systems, includes the following steps: S1. Based on the aircraft's historical position, speed, and heading data, use a GRU neural network to predict the aircraft's flight trajectory for future periods. S2. Combining the geographical distribution data of VHF base stations and L-band digital aviation communication system base stations with the real-time ephemeris data of geostationary orbit satellites and low Earth orbit satellites, a set of candidate links within the flight trajectory coverage area is selected. S3. Real-time monitoring and collection of communication quality indicators of all candidate links in the candidate link set, and calculation of predictive indicators for each candidate link based on the flight trajectory, including maximum sustainable service time and average link quality in future periods. S4. Divide the communication quality indicators and predictive indicators of each candidate link into benefit indicators and cost indicators, and map them into dimensionless utility values through different utility functions. S5. Use the analytic hierarchy process (AHP) to construct a judgment matrix, calculate the weight vector of each indicator for each candidate link, and linearly weight and sum the dimensionless utility value of each indicator with its corresponding weight to obtain the comprehensive utility value of each candidate link. S6. Execute link switching decisions based on the comprehensive utility value of candidate links. Link switching decisions include centralized switching decisions and distributed switching decisions, which are used to determine the optimal switching link.
[0009] Specifically, the calculation process of the GRU neural network in step S1 includes: S11, Reset door calculation: ; in, It is the sigmoid activation function. , This is the weight coefficient matrix for the reset gate; For a moment The input vector includes the aircraft's historical position, speed, and heading data. For a moment The hidden state vector; S12, Update gate calculation: ; in, , To update the weight coefficient matrix of the gate; S13, Update hidden status, always Hidden state vector The calculation is as follows: ; in, is the candidate hidden state vector.
[0010] Specifically, the selection criteria for the candidate link set in step S2 are: the intersection of the aircraft's flight trajectory with the coverage of VHF base stations and L-band digital aviation communication system base stations, or the intersection with the coverage of geostationary orbit satellites and low Earth orbit satellites.
[0011] Specifically, the maximum sustainable service time in step S3 is determined by analyzing the spatiotemporal geometric relationship between the predicted flight trajectory and the coverage of the candidate links; The method for calculating the average link quality in the future period is as follows: within the time window of the maximum sustainable service time of the candidate link, discrete sampling is performed along the predicted flight trajectory, the signal quality of each sampling point is calculated using the channel propagation model, and finally the average link quality in the future period is obtained by averaging the signal quality of all sampling points.
[0012] Specifically, in step S4, a larger value for the benefit-type indicator indicates better performance of the candidate link, while a larger value for the cost-type indicator indicates worse performance of the candidate link. The utility function for the benefit-type indicator is: ; The utility function of cost-based indicators is: ; in, This represents the set of candidate links within the current decision-making cycle. For link index; This represents a set of communication quality metrics and predictive metrics. For indexing indicators; This represents a subset of benefit-type indicators. Represents a subset of cost-type indicators, and satisfies , ; Indicate candidate link In terms of indicators The original value on, Indicates link In terms of indicators The dimensionless utility value on.
[0013] Specifically, step S5 includes: S51. Use the Saaty 1-9 scaling method to perform pairwise comparisons of several indicators for each candidate link, and construct a switching factor judgment matrix: ; It satisfies: , , ;in, Indicators relative to indicators Importance scale, the larger the value, the more important the indicator. Relative indicators The more important; S52. Calculate the switching factor judgment matrix. Maximum eigenvalue The corresponding feature vectors are then normalized to obtain the weight vectors for each indicator. ,in These represent the weights of different metrics in each candidate link, and satisfy the following conditions: ; S53. Calculate the consistency ratio ,in , For switching factor judgment matrix The order of The average random consistency index, if If the consistency of the switching factor judgment matrix is accepted, then the switching factor judgment matrix must be adjusted; otherwise, the switching factor judgment matrix must be adjusted. S54. Calculate the dimensionless utility value of each index in each candidate link obtained through utility function mapping in step S4. Weight vectors corresponding to each metric of the candidate link Perform a linear weighted summation to calculate the overall utility value of each candidate link. : ; Overall utility value A higher value indicates better overall performance of the candidate link.
[0014] Specifically, in step S6, the centralized handover decision-making first involves constructing a centralized decision state, and then the central controller executes the centralized link handover process based on Q-learning; the process of constructing the centralized decision state is as follows: (1) Constructing the global state space: The central controller collects the global network state at each decision moment. The global network state is represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of aircraft. The total number of candidate links; (2) Define the action space: The action space includes the switching decisions of all aircraft and is represented as an action matrix. : ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time; for any aircraft... , ; (3) Design the reward function: The reward function is the average comprehensive utility value of all aircraft, and the calculation method is as follows: ; in Indicate candidate link The overall utility value.
[0015] Specifically, the centralized link switching process based on Q-learning includes: (1) The central controller collects the global network state at each decision moment. Based on global network state The central controller adopts Greedy strategy selects action matrix , The greedy strategy is specifically: The probability selection maximizes the current Q-value; that is, for each aircraft, the link with the largest Q-value among the candidate links is selected as the switching target. The probability of randomly selecting an action. For exploration rate, ; (2) Incorporate decision-making experience The tuples are stored in the experience pool, where... Current global network state , For the selected action matrix , Instantaneous reward calculated for the reward function , The next global network state after the switch; (3) Model training: Randomly sample a batch of experience quadruples from the experience pool and perform the following training operations: a. Calculate the temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training, and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the temporal difference error using stochastic gradient descent and adjust the parameters of the online Q-network accordingly; c. Periodically synchronize the parameters of the target Q network with the parameters of the online Q network; (4) The central controller sends action matrices to each aircraft. The corresponding optimal switching command is executed by each aircraft, which then performs a link switching operation and switches to the candidate link specified in the action matrix. (5) The central controller collects the next global state after the switch. This will be used as the input state for the next round of decision-making.
[0016] Specifically, in step S6, the distributed handover decision-making process first involves constructing a distributed decision state, and then each aircraft, acting as an independent agent, executes the distributed link handover process based on Q-learning. The process of constructing the distributed decision state is as follows: (1) Constructing the state space: Each aircraft constructs its own state space based on locally acquired observation information. Network status Represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of candidate links; (2) Define the action space: The action space includes only the switching selection of a single aircraft, represented as: ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time, i.e., from Select a single action to execute; (3) Design the reward function: Airplane The reward function based on its own communication quality is expressed as: ; in Indicate candidate link The overall utility value.
[0017] Specifically, the distributed link switching process based on Q-learning includes: (1) Each aircraft, as an intelligent agent, independently collects local state information at each decision-making moment. Based on local network status information Each aircraft adopts Greedy strategy selects switching action , The greedy strategy is specifically: The probability of selecting the candidate link with the largest local Q value is used to... Candidate links are randomly selected with a probability. For exploration rate, ; (2) The airplane Decision-making experience The tuples are stored in the local experience pool, where For airplane Current local network status, For airplane Selected switching action, For airplane Instantaneous reward calculated using a reward function For airplane The next local network status after the switch; (3) Distributed model training, for each aircraft Randomly sample a batch of experience quadruples from the local experience pool and perform the following training operations: a. Calculating the aircraft The temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the time-series difference error using stochastic gradient descent and adjust the updated aircraft. Parameters of the online Q network; c. Regularly transport aircraft The parameters of the target Q-network are synchronized with the parameters of its online Q-network; (4) Each aircraft Based on the locally trained online Q-network, it autonomously selects the optimal switching link and performs the switching operation; (5) Each aircraft Collect the next local state after the switch This will be used as the input state for the next round of decision-making.
[0018] In summary, the beneficial technical effects of the present invention are as follows: 1. By using GRU trajectory prediction and link quality estimation, the handover decision-making process has been transformed from a passive response to an active prediction, significantly reducing the risks of erroneous handovers and handover delays; 2. The comprehensive utility function constructed by the analytic hierarchy process (AHP) enables intelligent trade-offs among multi-dimensional indicators of candidate switching links, thereby improving the decision-making reliability of highly dynamic hybrid networks. 3. Introduce the Q-learning framework to enable the switching system to learn autonomously through interaction with the environment, dynamically adjust strategies to maximize long-term communication benefits, and adapt to complex and ever-changing network environments; 4. Through intelligent and precise switching, the probability of communication interruption and switching failure is significantly reduced, effectively ensuring the continuity and stability of data transmission; 5. By employing both centralized and distributed decision-making models, focusing on global optimization and local decision-making efficiency respectively, it can adapt to different aviation communication management needs and network architectures. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a civil aircraft switching between heterogeneous aviation networks. Figure 2 This is the switching decision flowchart of the present invention. Detailed Implementation
[0020] To make the technical means, creative features, objectives and effects of this invention clearer and easier to understand, the invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0021] Example This invention provides a Q-learning-based airborne dynamic handover method for heterogeneous aviation networks, applicable to applications such as... Figure 1 The network scenario shown is as follows: Figure 2 As shown, the specific steps include: Step S1: Based on the aircraft's historical position, speed, and heading data, use a GRU neural network to predict the aircraft's flight trajectory for future periods. The core of the GRU neural network includes two types of gated variables: reset gate and update gate. These are used to selectively extract historical information and control the state update magnitude, thereby outputting the position information within the future time window. The calculation process includes: S11. The reset gate is used to determine the degree of retention of historical information, and is calculated as follows: ; in, It is the sigmoid activation function. , This is the weight coefficient matrix for the reset gate; For a moment The input vector includes the aircraft's historical position, speed, and heading data. For a moment The hidden state vector; S12, The update gate controls the fusion ratio of current input information and historical information, and is calculated as follows: ; in, , To update the weight coefficient matrix of the gate; S13, Update hidden status, always Hidden state vector The calculation is as follows: ; in, The candidate hidden state vectors are computed based on the reset gate. .
[0022] Step S2: Based on the civil aircraft trajectory prediction results output by GRU, and combined with the known geographical distribution data of VHF base stations and LDACS base stations and the real-time ephemeris data of GEO satellites and LEO satellites, through geometric calculation and coverage analysis, dynamically determine all VHF base stations, LDACS base stations and GEO satellites that may provide communication services during the predicted trajectory period, and form a set of handover candidate links accordingly.
[0023] Step S3: Real-time monitoring and collection of communication quality indicators for all candidate links within the candidate link set. Communication quality indicators include multiple dimensions such as signal received strength, available bandwidth, transmission delay, and available channel resources of nodes. Then, based on the flight trajectory, predictive indicators for each candidate link are calculated, including maximum sustainable service time and average link quality for the future period. The maximum sustainable service time is determined by analyzing the spatiotemporal geometric relationship between the predicted flight trajectory and the coverage area of the candidate link. The average link quality for the future period is calculated by discretely sampling along the predicted flight trajectory within the time window of the maximum sustainable service time of the candidate link, calculating the signal quality of each sampling point using the channel propagation model, and finally averaging the signal quality of all sampling points to obtain the average link quality for the future period. This indicator effectively overcomes the impact of instantaneous channel fluctuations and achieves robust evaluation of the overall performance of the future link.
[0024] Step S4: After completing the collection and prediction of link quality indicators, this invention transforms heterogeneous performance indicators from different candidate handover base stations, which have different physical meanings and dimensions, into directly comparable dimensionless utility values using mathematical methods. This method is used to quantify the importance of each attribute and its impact on the decision-making results. The communication quality and predictive metrics of the acquired candidate links are categorized into benefit-based metrics and cost-based metrics according to their physical characteristics. Higher values for benefit-based metrics indicate better link performance, such as signal strength, available bandwidth, and average link quality over future periods. Higher values for cost-based metrics indicate worse link performance, such as node load (or the reciprocal of available channel resources) and transmission delay. For each type of metric, a monotonic utility mapping function is designed. The utility function of the benefit-type indicator is: ; The utility function of cost-based indicators is: ; in, This represents the set of candidate links within the current decision-making cycle. For link index; This represents a set of communication quality metrics and predictive metrics. For indexing indicators; This represents a subset of benefit-type indicators. Represents a subset of cost-type indicators, and satisfies , ; Indicate candidate link In terms of indicators The original value on, Indicates link In terms of indicators The dimensionless utility value on.
[0025] Step S5: This invention integrates multiple heterogeneous link performance indicators into a unified comprehensive utility value using the analytic hierarchy process (AHP), providing a quantitative basis for handover decisions. Specifically, this invention comprehensively considers the service characteristics and service quality requirements of aviation communications, selecting four key decision factors to construct an evaluation system: 1. Average link quality over future periods, which, as a core factor, directly determines the transmission performance and reliability of the communication link; 2. Available channel resources, as an important factor, reflect the real-time load status of network nodes and directly affect the success rate of handover requests; 3. Maximum sustainable service time, as a key factor, determines the sustainable service capability of the link and affects the system handover frequency and stability; 4. Transmission latency, as a fundamental factor, reflects the real-time characteristics of data transmission and ensures the timeliness requirements of critical services. Based on this importance ranking, a judgment matrix is constructed and subsequent weight calculations are performed. The specific steps are as follows: S51. Use the Saaty 1-9 scaling method to perform pairwise comparisons of selected indicators for each candidate link, and construct a switching factor judgment matrix: ; It satisfies: , , ;in, Indicators relative to indicators Importance scale, the larger the value, the more important the indicator. Relative indicators The more important; S52. Calculate the switching factor judgment matrix. Maximum eigenvalue The corresponding feature vectors are then normalized to obtain the weight vectors for each indicator. ,in These represent the weights of different metrics in each candidate link, and satisfy the following conditions: ; S53. Calculate the consistency ratio ,in , For switching factor judgment matrix The order of The average random consistency index, if If the consistency of the switching factor judgment matrix is accepted, then the switching factor judgment matrix must be adjusted; otherwise, the switching factor judgment matrix must be adjusted. S54. Calculate the dimensionless utility value of each index in each candidate link obtained through utility function mapping in step S4. Weight vectors corresponding to each metric of the candidate link Perform a linear weighted summation to calculate the overall utility value of each candidate link. : ; Overall utility value A higher value indicates better overall performance of the candidate link, providing direct input for subsequent reinforcement learning-based switching decisions.
[0026] Step S6: Perform link switching decisions based on the comprehensive utility value of candidate links. Link switching decisions include centralized switching decisions and distributed switching decisions, which are used to determine the optimal switching link.
[0027] For centralized switching decisions, the switching decision process is modeled as a Markov decision process, where the utility function of the switching decision is regarded as the action reward. The environmental state, action space, and reward function are defined as follows: (1) Constructing the global state space: In the centralized architecture, the central controller establishes a single-agent MDP framework to describe the switching process of the entire network. The central controller collects the global network state at each decision moment, representing the complete feature information of all available candidate links at each decision moment. The global network state is represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of aircraft. The total number of candidate links; (2) Define the action space: The action space includes the switching decisions of all aircraft and is represented as an action matrix. : ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time; for any aircraft... , ; (3) Design the reward function: The reward function is the average comprehensive utility value of all aircraft, and the calculation method is as follows: ; in Indicate candidate link The overall utility value.
[0028] Based on the completed state space and reward function design, a centralized link switching process is executed using Q-learning, specifically as follows: (1) The central controller collects the global network state at each decision moment. Based on global network state The central controller adopts Greedy strategy selects action matrix , The greedy strategy is specifically: The probability selection maximizes the current Q-value; that is, for each aircraft, the link with the largest Q-value among the candidate links is selected as the switching target. The probability of randomly selecting an action. For exploration rate, ; (2) Incorporate decision-making experience The tuples are stored in the experience pool, where... Current global network state , For the selected action matrix , Instantaneous reward calculated for the reward function , The next global network state after the switch; (3) Model training: Randomly sample a batch of experience quadruples from the experience pool and perform the following training operations: a. Calculate the temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training, and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the temporal difference error using stochastic gradient descent and adjust the parameters of the online Q-network accordingly; c. Periodically synchronize the parameters of the target Q network with the parameters of the online Q network; (4) The central controller sends action matrices to each aircraft based on the online Q network. The corresponding optimal switching command is executed by each aircraft, which then performs a link switching operation and switches to the candidate link specified in the action matrix. (5) The central controller collects the next global state after the switch. This state is used as the input state for the next round of decision-making, realizing a closed-loop optimization of "global state collection → action selection → experience storage → model training → switching execution → state feedback".
[0029] In response to the potential significant changes in the dynamic characteristics of the environment, the system supports pre-training a set of policies covering a variety of typical scenarios. In actual deployment, the system can quickly match and load the optimal policy based on the environmental characteristics perceived in real time, thereby effectively avoiding the time overhead and computational burden caused by retraining.
[0030] To address the issues of high signaling overhead and heavy computational burden on the control center that may arise from centralized handover decision-making schemes when the number of satellites, ground base stations, and aircraft increases significantly, this invention further proposes a distributed intelligent handover decision-making scheme. In this scheme, each aircraft acts as an independent intelligent agent, autonomously making handover decisions based solely on local observation information, effectively achieving decentralized handover control. The first step is to construct a distributed decision-making state: (1) Constructing the state space: In a distributed architecture, each aircraft constructs its own state space based on locally acquired observation information. Network status Represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of candidate links; (2) Define the action space: The action space in the distributed scheme only includes the handover selection of a single aircraft, represented as: ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time, i.e., from Select a single action to execute; (3) Design the reward function: Airplane The reward function based on its own communication quality is expressed as: ; in Indicate candidate link The overall utility value is calculated. Through the above design, a complete decentralized handover decision-making framework is constructed. Each aircraft independently runs a reinforcement learning algorithm based on its local state information, achieving rapid and adaptive handover decisions. While ensuring individual service quality, this significantly reduces system signaling overhead and computational burden, and exhibits good scalability.
[0031] Based on the completed design of the distributed state space and reward function, a distributed link switching process is executed based on Q-learning, specifically as follows: (1) Each aircraft, as an intelligent agent, independently collects local state information at each decision-making moment. Based on local network status information Each aircraft adopts Greedy strategy selects switching action , The greedy strategy is specifically: The probability of selecting the candidate link with the largest local Q value is used to... Candidate links are randomly selected with a probability. For exploration rate, ; (2) The airplane Decision-making experience The tuples are stored in the local experience pool, where For airplane Current local network status, For airplane Selected switching action, For airplane Instantaneous reward calculated using a reward function For airplane The next local network status after the switch; (3) Distributed model training, for each aircraft Randomly sample a batch of experience quadruples from the local experience pool and perform the following training operations: a. Calculating the aircraft The temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the time-series difference error using stochastic gradient descent and adjust the updated aircraft. Parameters of the online Q network; c. Regularly transport aircraft The parameters of the target Q-network are synchronized with the parameters of its online Q-network; (4) Each aircraft Based on the locally trained online Q-network, it autonomously selects the optimal switching link and performs the switching operation; (5) Each aircraft Collect the next local state after the switch This state is used as the input state for the next round of decision-making, realizing a closed-loop optimization of "local state collection → action selection → experience storage → model training → switching execution → state feedback".
[0032] Through the above design, a fully distributed multi-agent switching decision framework was realized. While ensuring the quality of individual services for each aircraft, the system response speed and scalability were significantly improved, and signaling overhead and computational complexity were reduced.
[0033] Therefore, this invention provides an airborne dynamic handover method for heterogeneous aviation networks based on Q-learning. This method employs a core process of using a GRU neural network to predict aircraft trajectories to screen candidate links, applying the Analytic Hierarchy Process (AHP) to construct a comprehensive utility function to quantitatively evaluate multi-dimensional link indicators, and finally making adaptive handover decisions based on a Q-learning framework. This achieves a closed-loop intelligent handover process of "forward-looking prediction - comprehensive evaluation - intelligent decision-making," solving the pain points of traditional handover methods that rely solely on instantaneous, single network parameters, leading to frequent handovers, communication interruptions, and degraded service quality in highly dynamic heterogeneous networks. It improves the scientific rigor and foresight of handover decisions, significantly reduces handover failure rates and communication interruption probabilities, ensures the continuity and stability of communication links throughout the entire flight route, and ultimately provides an efficient and reliable link management solution for future integrated air-space-ground communication networks through intelligent and adaptive learning capabilities, ensuring the reliability and service quality of communication in highly dynamic aviation scenarios.
[0034] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for dynamic airborne handover in heterogeneous aviation networks based on Q-learning, characterized in that: Applied to heterogeneous aviation networks including VHF base stations, L-band digital aviation communication system base stations, geostationary orbit satellite systems, and low Earth orbit satellite systems, the following steps are included: S1. Based on the aircraft's historical position, speed, and heading data, use a GRU neural network to predict the aircraft's flight trajectory for future periods. S2. Combining the geographical distribution data of VHF base stations and L-band digital aviation communication system base stations with the real-time ephemeris data of geostationary orbit satellites and low Earth orbit satellites, a set of candidate links within the flight trajectory coverage area is selected. S3. Real-time monitoring and collection of communication quality indicators of all candidate links in the candidate link set, and calculation of predictive indicators for each candidate link based on the flight trajectory, including maximum sustainable service time and average link quality in future periods. S4. Divide the communication quality indicators and predictive indicators of each candidate link into benefit indicators and cost indicators, and map them into dimensionless utility values through different utility functions. S5. Use the analytic hierarchy process (AHP) to construct a judgment matrix, calculate the weight vector of each indicator for each candidate link, and linearly weight and sum the dimensionless utility value of each indicator with its corresponding weight to obtain the comprehensive utility value of each candidate link. S6. Execute link switching decisions based on the comprehensive utility value of candidate links. Link switching decisions include centralized switching decisions and distributed switching decisions, which are used to determine the optimal switching link.
2. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 1, characterized in that, The computation process of the GRU neural network in step S1 includes: S11, Reset door calculation: ; in, It is the sigmoid activation function. , This is the weight coefficient matrix for the reset gate; For a moment The input vector includes the aircraft's historical position, speed, and heading data. For a moment The hidden state vector; S12, Update gate calculation: ; in, , To update the weight coefficient matrix of the gate; S13, Update hidden status, always Hidden state vector The calculation is as follows: ; in, is the candidate hidden state vector.
3. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 1, characterized in that, The selection criteria for the candidate link set in step S2 are: the intersection of the aircraft's flight trajectory with the coverage of VHF base stations and L-band digital aviation communication system base stations, or the intersection with the coverage of geostationary orbit satellites and low Earth orbit satellites.
4. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 1, characterized in that, In step S3, the maximum sustainable service time is determined by analyzing the spatiotemporal geometric relationship between the predicted flight trajectory and the coverage area of the candidate links; The method for calculating the average link quality in the future period is as follows: within the time window of the maximum sustainable service time of the candidate link, discrete sampling is performed along the predicted flight trajectory, the signal quality of each sampling point is calculated using the channel propagation model, and finally the average link quality in the future period is obtained by averaging the signal quality of all sampling points.
5. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 1, characterized in that, In step S4, a higher value for the benefit-type indicator indicates better performance of the candidate link, while a higher value for the cost-type indicator indicates worse performance. The utility function for the benefit-type indicator is: ; The utility function of cost-based indicators is: ; in, This represents the set of candidate links within the current decision-making cycle. For link index; This represents a set of communication quality metrics and predictive metrics. For indexing indicators; This represents a subset of benefit-type indicators. Represents a subset of cost-type indicators, and satisfies , ; Indicate candidate link In terms of indicators The original value on, Indicate candidate link In terms of indicators The dimensionless utility value on.
6. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 5, characterized in that, Step S5 specifically includes: S51. Use the Saaty 1-9 scaling method to perform pairwise comparisons of several indicators for each candidate link, and construct a switching factor judgment matrix: ; It satisfies: , , ;in, Indicators relative to indicators Importance scale, the larger the value, the more important the indicator. Relative indicators The more important; S52. Calculate the switching factor judgment matrix. Maximum eigenvalue The corresponding feature vectors are then normalized to obtain the weight vectors for each indicator. ,in These represent the weights of different metrics in each candidate link, and satisfy the following conditions: ; S53. Calculate the consistency ratio ,in , For switching factor judgment matrix The order of The average random consistency index, if If the consistency of the switching factor judgment matrix is accepted, then the consistency of the switching factor judgment matrix is accepted; otherwise, the switching factor judgment matrix is adjusted. S54. Calculate the dimensionless utility value of each index in each candidate link obtained through utility function mapping in step S4. Weight vectors corresponding to each metric of the candidate link Perform a linear weighted summation to calculate the overall utility value of each candidate link. : ; Overall utility value A higher value indicates better overall performance of the candidate link.
7. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 6, characterized in that, In step S6, the centralized handover decision-making process first involves constructing a centralized decision state, and then the central controller executes the centralized link handover process based on Q-learning. The specific process of constructing the centralized decision state is as follows: (1) Constructing the global state space: The central controller collects the global network state at each decision moment. The global network state is represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of aircraft. The total number of candidate links; (2) Define the action space: The action space includes the switching decisions of all aircraft and is represented as an action matrix. : ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time; for any aircraft... , ; (3) Design the reward function: The reward function is the average comprehensive utility value of all aircraft, and the calculation method is as follows: ; in Indicate candidate link The overall utility value.
8. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 7, characterized in that, The centralized link switching process based on Q-learning includes: (1) The central controller collects the global network state at each decision moment. Based on global network state Calculate the Q-value of each candidate link, using... Greedy strategy selects action matrix , The greedy strategy is specifically: The probability selection maximizes the current Q-value; that is, for each aircraft, the link with the largest Q-value among the candidate links is selected as the switching target. The probability of randomly selecting an action. For exploration rate, ; (2) Incorporate decision-making experience The tuples are stored in the experience pool, where... Current global network state , For the selected action matrix , Instantaneous reward calculated for the reward function , The next global network state after the switch; (3) Model training: Randomly sample a batch of experience quadruples from the experience pool and perform the following training operations: a. Calculate the temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training, and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the temporal difference error using stochastic gradient descent and adjust the parameters of the online Q-network accordingly; c. Periodically synchronize the parameters of the target Q network with the parameters of the online Q network; (4) The central controller sends action matrices to each aircraft. The corresponding optimal switching command is executed by each aircraft, which then performs a link switching operation and switches to the candidate link specified in the action matrix. (5) The central controller collects the next global state after the switch. This will be used as the input state for the next round of decision-making.
9. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 6, characterized in that, In step S6, the distributed handover decision-making process first involves constructing a distributed decision state. Then, each aircraft acts as an independent agent, executing the distributed link handover process based on Q-learning. The specific process of constructing the distributed decision state is as follows: (1) Constructing the state space: Each aircraft constructs its own state space based on locally acquired observation information. Network status Represented as: ; in, Indicates airplane For candidate links The state observation vector includes the link characteristics of the candidate link. The total number of candidate links; (2) Define the action space: The action space includes only the switching selection of a single aircraft, represented as: ; in, Indicates airplane Should we switch to the candidate link? Each aircraft selects only one link to connect to at any given time, i.e., from Select a single action to execute; (3) Design the reward function: Airplane The reward function based on its own communication quality is expressed as: ; in Indicate candidate link The overall utility value.
10. The airborne dynamic handover method for heterogeneous aviation networks based on Q-learning according to claim 9, characterized in that, The distributed link switching process based on Q-learning includes: (1) Each aircraft, as an intelligent agent, independently collects local state information at each decision-making moment. Based on local network status information Calculate the Q value of each link, and each aircraft adopts... Greedy strategy selects switching action , The greedy strategy is specifically: The probability of selecting the candidate link with the largest local Q value is used to... Candidate links are randomly selected with a probability. For exploration rate, ; (2) The airplane Decision-making experience Stored in local experience pool as tuples, where, For airplane Current local network status, For airplane Selected switching action, For airplane Instantaneous reward calculated using a reward function For airplane The next local network status after the switch; (3) Distributed model training, for each aircraft Randomly sample a batch of experience quadruples from the local experience pool and perform the following training operations: a. Calculating the aircraft The temporal difference error between the target Q-network and the online Q-network, where the target Q-network is a fixed network used for stable training and the online Q-network is an updatable network used for real-time prediction of Q-values; b. Minimize the time-series difference error using stochastic gradient descent and adjust the updated aircraft. Parameters of the online Q network; c. Regularly transport aircraft The parameters of the target Q-network are synchronized with the parameters of its online Q-network; (4) Each aircraft Based on the locally trained online Q-network, it autonomously selects the optimal switching link and performs the switching operation; (5) Each aircraft Collect the next local state after the switch This will be used as the input state for the next round of decision-making.