A method, system and electronic equipment for dynamic handover of satellite-ground heterogeneous networks

By collecting multi-dimensional state parameters and performing dynamic correlation processing, reinforcement learning and Bayesian execution engine are used to optimize the handover of heterogeneous satellite-ground networks. This solves the problems of frequent misjudgments and ping-pong handover in existing technologies, achieving higher handover accuracy and flexibility, and improving user experience.

CN121056959BActive Publication Date: 2026-04-03SHANGHAI JUZHIXING NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, handover in heterogeneous satellite-ground networks relies solely on the signal strength of ground base stations or satellite links, leading to frequent misjudgments, severe ping-pong handover phenomena, and impacting service continuity and network resource utilization.

Method used

By collecting multi-dimensional state parameters from the network side, user side, and business side, performing dynamic correlation processing, and using reinforcement learning models and Bayesian execution engines for dynamic decision optimization, a unified state vector and switching decision weights are generated to determine the network switching target and execute parameters.

Benefits of technology

It improves the accuracy, reliability, and flexibility of handover between satellite and ground heterogeneous networks, reduces ping-pong handover, enhances network resilience, and provides better service continuity and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056959B_ABST
    Figure CN121056959B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and electronic device for dynamic handover in heterogeneous satellite-ground networks, relating to the field of wireless communication technology. The method includes: collecting core state parameters from the network side, user side, and service side respectively to obtain state parameters related to network handover on the network side, user side, and service side; dynamically associating the state parameters corresponding to the network side, user side, and service side to generate a unified state vector; using a reinforcement learning model, performing dynamic decision optimization based on the unified state vector to obtain satellite-ground handover decision weights; generating a satellite-ground network comprehensive score difference based on the satellite-ground handover decision weights, and determining the network handover target and its decision confidence level based on the satellite-ground network comprehensive score difference; using a Bayesian execution engine, determining the execution parameters of the network handover target based on the decision confidence level, and performing network handover based on the execution parameters. This invention improves the handover performance of heterogeneous satellite-ground networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and more specifically, to a method, system, and electronic device for dynamic switching of satellite-ground heterogeneous networks. Background Technology

[0002] With the rapid development of low-Earth orbit satellite communication systems, heterogeneous satellite-ground networks have become a key path to achieve seamless global coverage and enhance network resilience. To ensure service continuity and a better user experience, it is necessary to switch between satellite-ground heterogeneous networks based on actual conditions. This switching is typically achieved by making fixed decisions based on the signal strength of the ground base station or satellite link.

[0003] In related technologies, using only the signal strength of terrestrial base stations or satellite links as the basis for judging handover in heterogeneous satellite-terrestrial networks cannot fully reflect the true state of available networks, leading to frequent misjudgments. Furthermore, since handover is ultimately executed through a fixed strategy, rapid fluctuations in the signal strength of terrestrial base stations or satellite links in high-mobility scenarios can result in ping-pong handovers, leading to prolonged service interruptions and decreased network resource utilization. Summary of the Invention

[0004] The problem addressed by this invention is how to improve the handover performance of heterogeneous satellite-ground networks.

[0005] To address the above problems, this invention provides a method, system, and electronic device for dynamic switching of satellite-ground heterogeneous networks.

[0006] In a first aspect, the present invention provides a method for dynamic handover of heterogeneous satellite-ground networks, comprising:

[0007] Core status parameters are collected from the network side, user side, and service side respectively to obtain the status parameters related to network handover on the network side, the user side, and the service side respectively;

[0008] Dynamic association processing is performed on the state parameters corresponding to the network side, the user side, and the service side respectively to generate a unified state vector;

[0009] By using a reinforcement learning model, dynamic decision optimization is performed based on the unified state vector to obtain the satellite-to-ground switching decision weights.

[0010] Based on the satellite-to-ground handover decision weights, a satellite-to-ground network comprehensive score difference is generated, and based on the satellite-to-ground network comprehensive score difference, a network handover target and the decision confidence level of the network handover target are determined;

[0011] The Bayesian execution engine determines the execution parameters of the network switching target based on the decision confidence level, and performs network switching based on the execution parameters.

[0012] Optionally, the step of collecting core status parameters from the network side, user side, and service side respectively to obtain status parameters related to network handover on the network side, user side, and service side respectively includes:

[0013] The resource status of the network side is collected to obtain the ground base station parameters of the ground network and the low-orbit satellite parameters of the satellite network, and the ground base station parameters and the low-orbit satellite parameters are used as the status parameters of the network side.

[0014] The user on the user side is located using GPS or base station positioning data from the terminal, and the user's real-time movement speed is obtained. The real-time movement speed is then used as the status parameter on the user side.

[0015] By collecting data on the differences in service types on the service side, the QoS index of the service side is obtained, and the QoS index is used as the status parameter of the service side.

[0016] Optionally, the step of dynamically associating the state parameters corresponding to the network side, the user side, and the service side respectively to generate a unified state vector includes:

[0017] Sliding window filtering is applied to the reference signal received power in the ground base station parameters and the link signal-to-noise ratio in the low-orbit satellite parameters to obtain smoothed ground parameters and smoothed satellite parameters.

[0018] When the bandwidth utilization rate in the ground base station parameters is greater than or equal to the preset load correction threshold, the smoothed ground parameters are corrected according to the preset overload ratio to obtain the final ground parameters.

[0019] When the bandwidth utilization rate in the ground base station parameters is less than the preset load correction threshold, the smoothed ground parameters are used as the final ground parameters.

[0020] The QoS indicators are normalized according to the service type to obtain standardized QoS parameters;

[0021] The final ground parameters, the smoothed satellite parameters, and the standardized QoS parameters are concatenated into a unified state vector.

[0022] Optionally, the step of dynamically optimizing the decision based on the unified state vector using a reinforcement learning model to obtain the satellite-to-ground switching decision weights includes:

[0023] The unified state vector is input into the reinforcement learning model;

[0024] The signal quality weight, resource load weight, and service QoS weight are obtained by using the Actor network of the reinforcement learning model and performing Softmax normalization based on the unified state vector.

[0025] A three-dimensional continuous weight matrix is ​​formed based on the signal quality weight, the resource load weight, and the service QoS weight; wherein the sum of the signal quality weight, the resource load weight, and the service QoS weight is 1.

[0026] The three-dimensional continuous weight matrix is ​​used as the decision weight for satellite-to-ground switching.

[0027] Optionally, generating the satellite-to-ground network comprehensive score difference based on the satellite-to-ground handover decision weights includes:

[0028] Based on the three-dimensional continuous weight matrix and the QoS rewards corresponding to the ground network and the satellite network respectively, the comprehensive score of the ground network and the comprehensive score of the satellite network in the network side are determined.

[0029] The difference between the overall score of the ground network and the overall score of the satellite network is used as the overall score difference between the satellite and ground networks.

[0030] Optionally, determining the network handover target and the decision confidence level of the network handover target based on the comprehensive score difference of the satellite-to-ground network includes:

[0031] A dynamic lag threshold is determined by linear calculation based on the real-time movement speed on the user side.

[0032] The network handover target is determined based on the relationship between the combined score difference between the satellite and ground networks and the dynamic lag threshold.

[0033] If the combined score difference between the satellite and ground networks is greater than or equal to the dynamic lag threshold, then the network switching target is determined to be the satellite network.

[0034] If the combined score difference between the satellite and ground networks is less than the dynamic lag threshold, then the network switching target is determined to be the ground network;

[0035] The combined score difference of the satellite-ground network is substituted into the Sigmoid mapping function to obtain the decision confidence level.

[0036] Optionally, determining the execution parameters of the network switching target using a Bayesian execution engine based on the decision confidence includes:

[0037] The execution parameters are determined by performing Bayesian inference based on the decision confidence and a preset confidence threshold using the Bayesian execution engine.

[0038] Wherein, if the decision confidence level is greater than or equal to the preset confidence threshold, the execution parameters are single-link signaling transmission mode and HARQ retransmission disabled;

[0039] If the decision confidence level is less than the preset confidence threshold, then the execution parameters are dual-link parallel transmission mode and enabling one HARQ retransmission.

[0040] Optionally, the dynamic switching method for satellite-ground heterogeneous networks further includes:

[0041] After the network switch is completed, the execution effect data is obtained, including the switch success rate, service interruption duration, and service packet loss rate.

[0042] The switching success rate, the service interruption duration, and the service packet loss rate are quantified into a comprehensive feedback value according to a preset period.

[0043] The reward function weights and Actor network parameters of the reinforcement learning model are dynamically updated based on the comprehensive feedback value.

[0044] Secondly, the present invention provides a dynamic handover system for heterogeneous satellite-to-ground networks, comprising:

[0045] The multi-dimensional state awareness module is used to collect core state parameters from the network side, user side, and service side to obtain state parameters related to network handover for the network side, user side, and service side respectively; and to perform dynamic correlation processing on the state parameters corresponding to the network side, user side, and service side respectively to generate a unified state vector.

[0046] The reinforcement learning decision module is used to perform dynamic decision optimization based on the unified state vector through a reinforcement learning model to obtain satellite-to-ground handover decision weights; generate a satellite-to-ground network comprehensive score difference based on the satellite-to-ground network comprehensive score difference; and determine the network handover target and the decision confidence of the network handover target based on the satellite-to-ground network comprehensive score difference.

[0047] The execution module is used to determine the execution parameters of the network switching target based on the decision confidence using a Bayesian execution engine, and to perform network switching based on the execution parameters.

[0048] Thirdly, the electronic device of the present invention includes a memory and a processor;

[0049] The memory is used to store computer programs;

[0050] The processor is used to implement the above-described method for dynamic switching of heterogeneous satellite-ground networks when executing the computer program.

[0051] The present invention discloses a dynamic handover method, system, and electronic device for heterogeneous satellite-ground networks. By collecting state parameters related to network handover from the network side, user side, and service side, the network side can include signal strength, network load, and bandwidth utilization; the user side can cover user equipment movement speed, location information, and equipment performance; and the service side involves service type (such as voice, video, and data transmission), service priority, and service traffic. This multi-dimensional collection method can more comprehensively reflect the true state of the network and avoid misjudgments caused by a single factor. Furthermore, different services have different network requirements. For example, video services have high requirements for bandwidth and latency, while voice services are more sensitive to latency and packet loss rate. By collecting state parameters from the service side, handover decisions can be made according to specific service needs, better adapting to different service scenarios and improving user experience. Moreover, there are interrelationships and influences between the state parameters from different sides. For example, the movement speed of user equipment affects signal strength, and the size of service traffic affects network load. The present invention performs dynamic correlation processing on the collected state parameters from each side, fully considering the interaction between these parameters, so that the generated unified state vector more accurately reflects the overall state of the network. Integrating parameters from different sources and of different types into a unified state vector provides a unified input format for subsequent decision optimization, avoids decision bias caused by inconsistent parameters, and improves the accuracy and consistency of decision-making.

[0052] By employing a reinforcement learning model, the handover strategy is automatically learned and adjusted based on the current network state and historical experience. Compared to traditional fixed decision-making strategies, this invention dynamically optimizes the decision-making process according to changes in the network environment, better addressing complex network scenarios and rapidly changing signal conditions. The reinforcement learning model comprehensively considers the impact of multiple factors on handover effectiveness, finding the optimal satellite-to-ground handover decision weights. This allows the generated satellite-to-ground network comprehensive score difference to more accurately reflect the differences in quality between different networks, thereby determining more reasonable network handover targets and improving the accuracy and reliability of handover.

[0053] The Bayesian method evaluates the confidence level of decision results based on prior knowledge and current observation data. In handover of heterogeneous satellite-to-ground networks, the Bayesian execution engine determines the execution parameters of the network handover target based on the decision confidence level, effectively avoiding erroneous handovers caused by uncertainties and improving the reliability of handover decisions. Determining the execution parameters of the network handover target based on the decision confidence level allows for flexible adjustment of various parameters during the handover process, such as handover timing and handover speed. In high-mobility scenarios, when signal strength fluctuates rapidly, a more cautious handover strategy can be adopted based on lower confidence levels, avoiding ping-pong handovers, thereby reducing service interruption time and improving network resource utilization.

[0054] In summary, this invention comprehensively improves the handover performance of heterogeneous satellite-ground networks by acquiring multi-dimensional state parameters, performing dynamic correlation processing, optimizing dynamic decisions using reinforcement learning models, evaluating decision confidence using a Bayesian execution engine, and finally determining execution parameters. This enhances the accuracy, reliability, and flexibility of handover, effectively solves problems such as frequent misjudgments and ping-pong handover in existing technologies, and also strengthens network resilience, providing users with better service continuity and experience. Attached Figure Description

[0055] Figure 1 This is a flowchart of the dynamic switching method for heterogeneous satellite-ground networks according to an embodiment of the present invention;

[0056] Figure 2 This is a structural block diagram of a satellite-ground heterogeneous network dynamic switching system according to an embodiment of the present invention. Detailed Implementation

[0057] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0058] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0059] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0060] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0061] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this invention are all authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0062] Combination Figure 1 As shown in the figure, an embodiment of the present invention provides a method for dynamic handover of heterogeneous satellite-ground networks, comprising:

[0063] Core status parameters are collected from the network side, user side, and service side respectively to obtain the status parameters related to network handover on the network side, user side, and service side respectively.

[0064] Specifically, key parameters are collected in layers according to "network side - user side - service side", with a uniform collection period of 5ms (to meet the real-time requirements of high mobility scenarios). Network side resource status includes: ground base station parameters: current bandwidth utilization rate ( (unit: %) reflects the degree of scarcity of ground resources; reference signal receiving power (%) (Unit: dBm, raw ground signal strength); Low Earth Orbit (LEO) satellites: ratio of current active users to maximum capacity ( The value ranges from 0 to 1, and the calculation method is "number of active users within the beam / maximum number of users supported by the beam", reflecting satellite resource load and satellite link signal-to-noise ratio. (Unit: dB, representing the original satellite signal anti-interference capability). User-side mobility status: Real-time mobility speed of the terminal is obtained through terminal GPS or base station positioning data. (Unit: km / h), used to distinguish high-mobility scenarios ( >100km / h, such as high-speed rail and aviation) and low-mobility scenarios (such as urban commuting and static office work), providing a basis for subsequent scenario adaptation. Business-side QoS status: Collect core QoS indicators differentiated according to business type to ensure strong correlation with user experience: Real-time services (such as voice calls and telemedicine): Collect end-to-end latency ( (Unit: milliseconds), reflecting the real-time requirements of the business; real-time businesses (such as file download, video on demand): collect data packet loss rate ( (unit: %), reflecting the business reliability requirements.

[0065] The state parameters corresponding to the network side, the user side, and the service side are dynamically correlated to generate a unified state vector.

[0066] Specifically, firstly, on the network side, the reference received power (RSRP) of the ground base station and the smoothed signal-to-noise ratio (SNR) of the satellite link need to be processed. Specifically, since signal strength is easily affected by environmental interference and instantaneous changes, a sliding window filtering technique is used to smooth these parameters. This filtering technique reduces noise interference caused by signal fluctuations by averaging three consecutive signal strength values, resulting in smoothed ground signal parameters and satellite signal parameters. This processing more stably reflects the actual signal quality of the network, providing a more reliable basis for subsequent network handover decisions.

[0067] Secondly, for the user side, the user's real-time movement speed is obtained through the terminal's GPS positioning data or base station positioning information. Real-time movement speed reflects the user's current movement status and is crucial for determining whether the user is in a high-mobility scenario (such as high-speed rail or air travel). In high-mobility scenarios, signal strength changes more drastically, therefore different handover strategies are needed to ensure communication continuity and stability.

[0068] Finally, on the business side, Quality of Service (QoS) metrics are normalized according to different business types. For example, for real-time services (such as voice calls and telemedicine), low latency is emphasized, while for non-real-time services (such as file downloads and video-on-demand), the reliability of data transmission and packet loss rate are more important. By standardizing these QoS metrics, their comparability in the decision-making process can be ensured, thereby providing appropriate priorities for different business types.

[0069] By combining the parameters processed from the network, user, and service sides, a unified state vector is formed. This state vector includes not only filtered signal strength parameters but also user mobility characteristics and service QoS requirements. This unified state vector provides a comprehensive, accurate, and integrated information foundation reflecting the current network and user status for subsequent network handover decisions. This process effectively integrates parameters that were originally scattered across different dimensions, enabling the decision model to analyze and process data within a unified framework, thereby achieving more accurate and efficient network handover decisions.

[0070] By using a reinforcement learning model, dynamic decision optimization is performed based on the unified state vector to obtain the satellite-to-ground switching decision weights.

[0071] Specifically, a continuous state space is constructed using a 7-dimensional state vector S as input, and a 3-dimensional continuous action space is defined. ,in, For signal quality weights, For resource load weight, For service QoS weights, Furthermore, the weights of each dimension are ∈ [0,1]. By designing a dynamic weighted reward function, a synergistic balance is achieved between the three major objectives of "ensuring business QoS, maintaining handover stability, and optimizing satellite-ground resource efficiency," avoiding decision-making biases caused by a single objective (such as ignoring network overload in pursuit of QoS, or sacrificing business experience for stability), and guiding the reinforcement learning agent to learn the optimal strategy that is more in line with the actual scenario. The reward function design logic is as follows: the reward function r adopts a dynamic weighted form, quantifying the achievement degree of different objectives through three core reward items, and adaptively adjusting the weight coefficients according to the scenario characteristics to ensure accurate matching between the reward signal and the scenario requirements.

[0072] ;

[0073] in Dynamically adjusted weighting coefficients These are QoS rewards, stability rewards, and resource efficiency rewards, with values ​​ranging from [0,1] (a higher value indicates a better achievement of the goal).

[0074] Finally, the Deep Deterministic Policy Gradient (DDPG) algorithm is used for training. After training, the Actor network outputs the optimal weights A. .

[0075] Based on the satellite-to-ground handover decision weights, a satellite-to-ground network comprehensive score difference is generated, and based on the satellite-to-ground network comprehensive score difference, a network handover target and the decision confidence level of the network handover target are determined.

[0076] Specifically, after training, the Actor network is solidified into a decision model, and in real-time scenarios, it outputs switching decisions according to the following process: Comprehensive score calculation: Based on the current state S, the agent outputs the optimal weights through the Actor network. , For optimal signal quality weights, To achieve the optimal resource load weight, To determine the optimal service QoS weight, calculate the combined scores for the terrestrial network and satellite network separately:

[0077] ;

[0078] ;

[0079] in, QoS rewards for terrestrial / satellite networks (i.e., under the current network) value), The signal-to-noise ratio of the smoothed satellite link. This is the power received as a reference signal.

[0080] Furthermore, based on the combined score difference between the output satellite and ground networks, the abstract "decision reliability" is transformed into a calculable confidence value, providing prior evidence for Bayesian execution:

[0081] Let the ground network score be the output of reinforcement learning decisions. Satellite network score Define decision confidence level The formula for calculating the range of values ​​([0,1]) is:

[0082] ;

[0083] Where k=10 is the sensitivity coefficient (verified by actual testing, this coefficient can make the confidence level ≥0.8 when the score difference >0.1 and the confidence level ≤0.6 when the score difference <0.05, with the best discrimination).

[0084] when When this occurs, it is determined to be a "high-confidence decision" (the decision logic is reliable and no redundant execution is required).

[0085] when When this occurs, it is judged as a "low-confidence decision" (the decision is uncertain and its execution reliability needs to be strengthened).

[0086] The Bayesian execution engine determines the execution parameters of the network switching target based on the decision confidence level, and performs network switching based on the execution parameters.

[0087] Specifically, when the confidence level is high (C≥0.8), the ground signaling link is prioritized; when the confidence level is low (C<0.8), parallel transmission via both satellite and ground links is used, and valid signaling is selected based on Bayesian posterior probability. Furthermore, HARQ retransmission is disabled for real-time services, with one emergency retransmission only enabled when the confidence level is low and the packet loss rate is high; for non-real-time services, the number of retransmissions is dynamically calculated based on the packet loss probability estimated by Bayes. The handover success rate R, service interruption time T, and service packet loss P are collected, and a comprehensive feedback value F is calculated to feed back into the reinforcement learning decision model.

[0088] This embodiment of the dynamic handover method for heterogeneous satellite-ground networks collects state parameters related to network handover from the network side, user side, and service side. The network side parameters include signal strength, network load, and bandwidth utilization; the user side parameters include user device movement speed, location information, and device performance; and the service side parameters include service type (e.g., voice, video, data transmission), service priority, and service traffic. This multi-dimensional collection method more comprehensively reflects the true state of the network, avoiding misjudgments caused by a single factor. Furthermore, different services have different network requirements. For example, video services have high bandwidth and latency requirements, while voice services are more sensitive to latency and packet loss. By collecting state parameters from the service side, handover decisions can be made based on specific service needs, better adapting to different service scenarios and improving user experience. Moreover, there are interrelationships and influences between the state parameters from different sides. For example, the movement speed of user devices affects signal strength, and the amount of service traffic affects network load. This invention dynamically correlates the collected state parameters from each side, fully considering the interactions between these parameters, so that the generated unified state vector more accurately reflects the overall state of the network. Integrating parameters from different sources and of different types into a unified state vector provides a unified input format for subsequent decision optimization, avoids decision bias caused by inconsistent parameters, and improves the accuracy and consistency of decision-making.

[0089] By employing a reinforcement learning model, the handover strategy is automatically learned and adjusted based on the current network state and historical experience. Compared to traditional fixed decision-making strategies, this invention dynamically optimizes the decision-making process according to changes in the network environment, better addressing complex network scenarios and rapidly changing signal conditions. The reinforcement learning model comprehensively considers the impact of multiple factors on handover effectiveness, finding the optimal satellite-to-ground handover decision weights. This allows the generated satellite-to-ground network comprehensive score difference to more accurately reflect the differences in quality between different networks, thereby determining more reasonable network handover targets and improving the accuracy and reliability of handover.

[0090] The Bayesian method evaluates the confidence level of decision results based on prior knowledge and current observation data. In handover of heterogeneous satellite-to-ground networks, the Bayesian execution engine determines the execution parameters of the network handover target based on the decision confidence level, effectively avoiding erroneous handovers caused by uncertainties and improving the reliability of handover decisions. Determining the execution parameters of the network handover target based on the decision confidence level allows for flexible adjustment of various parameters during the handover process, such as handover timing and handover speed. In high-mobility scenarios, when signal strength fluctuates rapidly, a more cautious handover strategy can be adopted based on lower confidence levels, avoiding ping-pong handovers, thereby reducing service interruption time and improving network resource utilization.

[0091] In summary, this embodiment comprehensively improves the handover effect of heterogeneous satellite-ground networks by collecting multi-dimensional state parameters, dynamically associating data, optimizing dynamic decisions using reinforcement learning models, evaluating decision confidence using a Bayesian execution engine, and finally determining execution parameters. It enhances the accuracy, reliability, and flexibility of handover, effectively solves problems such as frequent misjudgments and ping-pong handover in existing technologies, and also strengthens network resilience, providing users with better service continuity and experience.

[0092] Optionally, the step of collecting core status parameters from the network side, user side, and service side respectively to obtain status parameters related to network handover on the network side, user side, and service side respectively includes:

[0093] The resource status of the network side is collected to obtain the ground base station parameters of the ground network and the low-orbit satellite parameters of the satellite network, and the ground base station parameters and the low-orbit satellite parameters are used as the status parameters of the network side.

[0094] The user on the user side is located using GPS or base station positioning data from the terminal, and the user's real-time movement speed is obtained. The real-time movement speed is then used as the status parameter on the user side.

[0095] By collecting data on the differences in service types on the service side, the QoS index of the service side is obtained, and the QoS index is used as the status parameter of the service side.

[0096] Specifically, when collecting core status parameters from the network side, user side, and service side, the network side first collects the bandwidth utilization rate (Lgs) and reference signal received power (RSRPgs) of the ground base station, and simultaneously collects the beam active user ratio (Lsat) and satellite link signal-to-noise ratio (SNRsat) of the low-Earth orbit satellite. These parameters reflect the resource load status and signal quality of the ground network and satellite network, respectively, and are used as the status parameters of the network side. For the user side, the real-time mobile speed (vcurr) of the terminal is obtained through terminal GPS or base station positioning data to distinguish between high-mobility scenarios (such as high-speed rail and aviation, vcurr>100km / h) and low-mobility scenarios (such as urban commuting and static office work, vcurr≤50km / h), and this real-time mobile speed is used as the status parameter of the user side. On the service side, core QoS indicators are collected differently according to the service type. For real-time services (such as voice calls and telemedicine), the end-to-end latency (Dcurr) is collected; for non-real-time services (such as file downloads and video on demand), the packet loss rate (Pcurr) is collected, and these QoS indicators are used as the status parameters of the service side.

[0097] In this optional embodiment, by collecting parameters from ground base stations and low-orbit satellites on the network side, real-time mobile speed on the user side, and QoS indicators on the service side, the true state of the network can be comprehensively and accurately reflected. Compared with traditional handover methods that rely solely on signal strength, this multi-dimensional information fusion approach avoids misjudgments caused by incomplete information, providing a more reliable data foundation for subsequent handover decisions. Furthermore, it addresses the unique characteristics of network status, user behavior, and service needs in different scenarios. For example, in high-mobility scenarios (such as high-speed rail and aviation), users move quickly, and signal strength changes frequently; in resource conflict scenarios (such as high ground base station load), network resources are strained. By collecting these multi-dimensional parameters, this embodiment allows the system to better adapt to various complex scenarios and provide more accurate handover decisions.

[0098] By acquiring a user's real-time movement speed through terminal GPS or base station positioning data, it's possible to accurately determine whether the user is in a high-mobility scenario. In high-mobility scenarios, signal strength changes rapidly, and frequent handovers can lead to service interruptions. Using real-time movement speed parameters, the system can predict signal change trends in advance, select more suitable handover times, reduce ping-pong handovers, and improve handover accuracy. QoS indicators are collected differentiated according to service type, allowing for priority ranking based on the service's sensitivity to metrics such as latency and packet loss rate. For example, for real-time services (such as voice calls and telemedicine), low-latency terrestrial network connections are prioritized; for non-real-time services (such as file downloads and video-on-demand), connection reliability is more important. This service priority adaptation mechanism ensures the continuity of critical services and improves user experience.

[0099] By collecting resource status parameters from terrestrial base stations and low-Earth orbit satellites, the system can understand the real-time load of network resources. In resource conflict scenarios, the system can dynamically adjust the handover strategy based on resource load to avoid overloading a single network. For example, when the load on terrestrial base stations is too high, the system can prioritize switching services to the satellite network, and vice versa. This resource load balancing mechanism can effectively improve the overall utilization of network resources and avoid resource waste. Based on the real-time mobility speed on the user side and the QoS indicators on the service side, the system can dynamically adjust the handover strategy. In high mobility scenarios, the system can prioritize links with more stable signals; in resource-constrained scenarios, the system can optimize the handover timing and reduce unnecessary handovers. This dynamic adjustment mechanism can further improve the utilization efficiency of network resources and reduce the impact of handover on network resources.

[0100] By collecting multi-dimensional state parameters, the system can automatically learn and adapt to various new scenarios. For example, in unpredictable scenarios such as sudden link blockages or temporary satellite beam adjustments, the system can adjust its switching strategy based on real-time parameters without manual intervention. This adaptive capability allows the system to quickly adapt to changes in the network environment and maintain stable performance. With the diversification of service types and changes in user needs, the system can dynamically adjust its switching strategy based on QoS indicators from the service side. For example, when real-time service demands increase, the system can prioritize ensuring the continuity of real-time services; when non-real-time service demands increase, the system can optimize its switching strategy to improve the transmission efficiency of non-real-time services. This flexible response mechanism ensures that the system can provide optimal switching services in different service scenarios.

[0101] This embodiment collects core status parameters from the network side, user side, and service side and uses them as status parameters respectively. This can comprehensively and accurately reflect the real status of the network, improve the accuracy of handover, optimize network resource utilization, enhance the adaptability and flexibility of the system, and provide a solid data foundation for subsequent handover decisions.

[0102] Optionally, the step of dynamically associating the state parameters corresponding to the network side, the user side, and the service side respectively to generate a unified state vector includes:

[0103] Sliding window filtering is applied to the reference signal received power in the ground base station parameters and the link signal-to-noise ratio in the low-orbit satellite parameters to obtain smoothed ground parameters and smoothed satellite parameters.

[0104] When the bandwidth utilization rate in the ground base station parameters is greater than or equal to the preset load correction threshold, the smoothed ground parameters are corrected according to the preset overload ratio to obtain the final ground parameters.

[0105] When the bandwidth utilization rate in the ground base station parameters is less than the preset load correction threshold, the smoothed ground parameters are used as the final ground parameters.

[0106] The QoS indicators are normalized according to the service type to obtain standardized QoS parameters;

[0107] The final ground parameters, the smoothed satellite parameters, and the standardized QoS parameters are concatenated into a unified state vector.

[0108] Specifically, the collected raw parameters undergo lightweight processing to eliminate the influence of noise and dimensions, and parameter correlation is established. This process includes three parts:

[0109] First, signal quality smoothing: for ground... With satellite To address the issue of data fluctuations due to transient interference (such as urban obstruction and atmospheric attenuation), a sliding window filtering algorithm is used to smooth the data. The window size is set to 3 acquisition cycles (i.e., 15ms), and the calculation formula is as follows:

[0110] ;

[0111] in, represent or t is the current acquisition period. These are the original parameter values ​​for the first two cycles. This processing can reduce the instantaneous fluctuation amplitude of the signal by more than 40%, avoiding false switching triggered by sudden drops in short-term signal strength.

[0112] Secondly, load correlation correction processing: For the misjudgment scenario of "ground signal meets the standard but the load is overloaded", the smoothed ground signal is processed. Perform load correlation correction when the bandwidth utilization of ground base stations is... When the load exceeds 80% (determined as a high load condition), the signal weight is reduced proportionally to the load excess. The correction formula is as follows:

[0113] ;

[0114] in, The corrected ground signal strength. Ensure that the portion of the load exceeding 80% is calculated linearly (maximum correction range 10%) to avoid a sharp increase in service latency after switching due to insufficient resources despite signal strength meeting the requirements.

[0115] In addition, QoS parameter standardization processing: addressing latency issues in real-time services. Non-real-time services The difference in the units of packet loss rate (ms and %) is standardized to the [0,1] interval to facilitate subsequent input into the decision model in conjunction with other parameters. The specific formula is as follows:

[0116] Real-time service latency standardization:

[0117] ;

[0118] in, The maximum tolerable latency for real-time services is set to 100ms; exceeding this value will significantly degrade the user experience.

[0119] Standardization of packet loss rate for non-real-time services:

[0120] ;

[0121] in, This is the maximum tolerable packet loss rate for non-real-time services, set to 1%. Exceeding this value requires triggering retransmission or handover. The min function prevents the normalized value from exceeding the [0,1] range when the packet loss rate is too high.

[0122] Based on the core influencing factors of handover decisions in heterogeneous satellite-ground networks (resource load, link quality, user mobility, and service demand), seven strongly correlated features were selected (weakly correlated parameters such as historical handover count and base station geographical location were removed), and the state vector was constructed as follows:

[0123] ;

[0124] All dimensions are linearly normalized to the [0,1] interval. The specific processing logic and physical meaning are as follows: Current bandwidth utilization of ground base stations ( ): Reflects the resource load status of the terrestrial cellular network. Its value is the percentage of the actual bandwidth currently occupied by the terrestrial base station relative to the total available bandwidth (original range 0%-100%), standardized and mapped to the [0,1] interval. This parameter is directly related to the service carrying capacity of the terrestrial network. The closer a value is to 1, the higher the load on the terrestrial base station, and the less redundant resources it has to support new services or maintain existing services. Therefore, it is necessary to prioritize switching to the satellite network. Conversely, if the value is less than 1, then the terrestrial network resources are sufficient, and terrestrial connections can be prioritized to ensure low latency.

[0125] In this optional embodiment, a sliding window filtering technique is used to smooth the reference received power (RSRP) of the terrestrial base station and the link signal-to-noise ratio (SNR) of the low-Earth orbit satellite. This effectively reduces instantaneous fluctuations in signal strength and minimizes drastic changes in signal strength caused by environmental interference or rapid terminal movement, thereby providing a more stable signal quality assessment. When the bandwidth utilization of the terrestrial base station exceeds a preset load correction threshold (e.g., 80%), the smoothed terrestrial parameters are corrected based on load correlation. This correction mechanism dynamically reflects the actual available resources of the terrestrial base station, avoiding incorrect selection of the terrestrial network under high load conditions due to adequate signal strength. For example, even if the signal strength of the terrestrial base station is good, if its load is too high, the corrected signal parameters will decrease, prompting the system to consider switching to the satellite network to balance network load and improve resource utilization. When the bandwidth utilization is lower than the preset load correction threshold, the smoothed terrestrial parameters are directly used as the final terrestrial parameters. This flexible approach ensures that the system can fully utilize the resources of the terrestrial network under normal load conditions, while avoiding unnecessary switching and improving the overall efficiency of the system.

[0126] QoS metrics are normalized based on service type to ensure comparability of QoS metrics across different service types during the decision-making process. For example, latency is a key metric for real-time services (such as voice calls and telemedicine), while packet loss rate is more important for non-real-time services (such as file downloads and video-on-demand). Normalization allows the decision-making model to analyze and process data within a unified framework, thereby providing appropriate priorities for different service types and improving the consistency of the service experience.

[0127] The final ground parameters, smoothed satellite parameters, and standardized QoS parameters are concatenated into a unified state vector, providing a comprehensive, accurate, and integrated information foundation reflecting the current network and user status for subsequent network handover decisions. Through dynamic correlation processing, the system can automatically adjust the processing method of state parameters based on real-time network status, user behavior, and service requirements. This dynamic adjustment mechanism enables the system to better adapt to various complex scenarios, such as high mobility scenarios, resource conflict scenarios, and link mutation scenarios, improving the system's adaptability and flexibility.

[0128] In summary, by dynamically associating the state parameters of the network side, user side, and service side and generating a unified state vector, the impact of signal fluctuations can be effectively reduced, network load changes can be dynamically adapted, service experience consistency can be improved, the comprehensiveness and accuracy of decision-making can be enhanced, and the system's adaptive capability can be strengthened, thereby significantly improving the handover effect of heterogeneous satellite-ground networks.

[0129] Optionally, the step of dynamically optimizing the decision based on the unified state vector using a reinforcement learning model to obtain the satellite-to-ground switching decision weights includes:

[0130] The unified state vector is input into the reinforcement learning model;

[0131] The signal quality weight, resource load weight, and service QoS weight are obtained by using the Actor network of the reinforcement learning model and performing Softmax normalization based on the unified state vector.

[0132] A three-dimensional continuous weight matrix is ​​formed based on the signal quality weight, the resource load weight, and the service QoS weight; wherein the sum of the signal quality weight, the resource load weight, and the service QoS weight is 1.

[0133] The three-dimensional continuous weight matrix is ​​used as the decision weight for satellite-to-ground switching.

[0134] Specifically, based on the core logic of decision weights determining policy bias, the core influencing factors of satellite-to-ground handover decisions (signal quality, resource load, and service QoS) are transformed into dynamically adjustable continuous weight dimensions. A continuous action space is adopted instead of discrete actions (such as the binary decision of switching to satellite / ground), enabling refined gradient updates of weights and avoiding the "either / or" limitations of discrete decisions. The physical rationality of the decision logic is ensured through a weight sum constraint (∑=1)—an increase in the weight of any dimension requires a moderate decrease in the weights of other dimensions, consistent with the actual network scenario of "limited resources and dynamic priority trade-offs." Combining the output characteristics of the reinforcement learning Actor network, a Softmax function is used for weight normalization, ensuring that the weights of each dimension are ∈[0,1] while smoothly responding to scene changes and avoiding decision oscillations caused by sudden weight changes. Specifically, a unified state vector generated by the multi-dimensional state awareness module is input into the reinforcement learning model. This unified state vector integrates key parameters from the network side, user side, and service side, providing comprehensive input for decision-making. The unified state vector is processed through the Actor network of the reinforcement learning model. The Actor network uses the Softmax normalization method to normalize the output signal quality weight, resource load weight, and service QoS weight, ensuring that the sum of the three is 1.

[0135] In a preferred embodiment of the present invention, the action space A is a 3-dimensional continuous weight combination that directly maps the bias of the decision-making strategy. Each dimension corresponds one-to-one with the core parameters of the state space, realizing a strong correlation between "state input and action output".

[0136] ;

[0137] in, And the weights of each dimension are ∈ [0,1]. For signal quality weights, For resource load weight, For service QoS weights.

[0138] For signal quality weights, the corresponding corrected ground RSRP in the state space ( ) and smoothed satellite SNR ( ), is a core indicator for measuring the stability of physical links; in high-mobility scenarios (terminal speed) For trains traveling at speeds of km / h, such as high-speed trains and intercity trains, this weight needs to be increased (typically within the range of 0.4-0.6) – because under high mobility, the signal connection between the terminal and the ground base station is easily interrupted. Although satellite links have slightly higher latency, their coverage is continuous, thus increasing this weight is necessary. Prioritize links with more stable signals (such as satellite) to avoid ping-pong handover caused by signal fluctuations; in low-mobility scenarios ( For speeds such as urban commuting and static office work, this weight can be reduced (typically 0.2-0.3), and the focus can be shifted to load and QoS. When the signal strength is increased, the decision-making process tends to prioritize "optimal signal quality," making it suitable for scenarios sensitive to link stability (such as drone inspections and aviation communications). However, if the signal strength is too low, the risk of signal degradation may be ignored, leading to service interruption after the switchover (such as choosing a satellite link with low load but poor signal).

[0139] For resource load weights, the associated state parameter is the terrestrial base station bandwidth utilization rate in the corresponding state space. With satellite beam load factor This directly reflects the remaining resource capacity of the satellite-to-ground network; ground / satellite overload scenarios ( This weight needs to be increased (typically within the range of 0.4-0.5). When the ground is overloaded, even if the signal strength is sufficient, priority should be given to offloading to the satellite to avoid a surge in service latency due to the depletion of ground resources. When the satellite is overloaded (e.g., multiple users accessing the satellite in a hotspot area), it is necessary to switch back to the ground to protect scarce satellite resources. In load balancing scenarios (… This weight can be reduced (typically 0.2-0.3), focusing more on signal strength and QoS; When the load is increased, the decision-making process tends to favor the "lowest load," which is suitable for peak scenarios with limited resources (such as urban commuting hours in the morning and evening, or large-scale event venues); if the load is too low, it may cause a certain network to be overloaded (such as not switching to satellite even when the ground load is 100%), resulting in overall QoS degradation.

[0140] For service QoS weights, the corresponding standardized service QoS metrics in the state space are ( ) and business type coding ( ), is the core guiding principle for ensuring user experience; real-time business scenarios ( =1, such as voice calls and telemedicine) need to increase this weight (typically 0.4-0.5) – real-time services are sensitive to latency (tolerance limit ≤100ms), so increase... Prioritize low-latency links (such as terrestrial networks), and ensure latency meets standards even if terrestrial load is slightly higher; for non-real-time service scenarios ( =0, such as file download, video caching) can reduce this weight (typically 0.2-0.3), allowing moderate latency in exchange for a lower packet loss rate (such as satellite links). When the value is increased, the decision-making process tends to favor "optimal QoS," which is suitable for services that are sensitive to user experience. However, if the value is too low, it may lead to QoS degradation for high-priority services (such as emergency calls), which violates the "service priority" principle.

[0141] The Actor network output process is as follows: The reinforcement learning Actor network takes a 7-dimensional state vector S as input, passes through a 3-layer fully connected neural network (with ReLU activation function in the hidden layers), and outputs 3 original weight values ​​(a1, a2, a3). These values ​​are then normalized using the Softmax function to obtain the final action A. The specific calculation is as follows:

[0142] ;

[0143] A=[ω sig ,ω load ,ω qos ]=[ω1,ω2,ω3];

[0144] in, To prevent numerical overflow due to small constants, the Softmax function ensures that the sum of the output weights is 1 and that each dimension ∈ [0,1]; extreme value avoidance is implemented to prevent a certain weight from approaching 0, which could lead to biased decision-making (e.g., →0 Ignore business experience), set minimum weight threshold If a weight is less than 0.1 after normalization, it is forcibly adjusted to 0.1, and other weights are reduced proportionally (keeping the total to 1) to ensure that all three dimensions participate in the decision-making process and avoid the traditional problem of "single factor dominance".

[0145] In this optional embodiment, dynamic decision optimization through reinforcement learning models can significantly improve the handover performance of heterogeneous satellite-ground networks. In high-mobility scenarios (such as high-speed rail and aviation), signal quality weights can be dynamically adjusted according to mobility speed, reducing ping-pong handover phenomena and ensuring service continuity. In resource load fluctuation scenarios (such as densely populated urban areas / suburbs), resource load weights can effectively balance the load of ground base stations and satellites, avoiding service packet loss due to resource overload. Furthermore, service QoS weights can prioritize the low-latency requirements of real-time services (such as voice calls and telemedicine) while also considering the high reliability requirements of non-real-time services (such as file downloads and video-on-demand). Through this multi-dimensional dynamic weight adjustment, highly accurate handover decisions can be achieved in complex and ever-changing network environments, improving network resource utilization and user experience.

[0146] Optionally, generating the satellite-to-ground network comprehensive score difference based on the satellite-to-ground handover decision weights includes:

[0147] Based on the three-dimensional continuous weight matrix and the QoS rewards corresponding to the ground network and the satellite network respectively, the comprehensive score of the ground network and the comprehensive score of the satellite network in the network side are determined.

[0148] The difference between the overall score of the ground network and the overall score of the satellite network is used as the overall score difference between the satellite and ground networks.

[0149] Specifically, a three-dimensional continuous weight matrix is ​​used, consisting of signal quality weights, resource load weights, and service QoS weights. These weights comprehensively reflect the multi-dimensional characteristics of the current network state. Then, combining the QoS rewards corresponding to the terrestrial and satellite networks, the comprehensive scores of the two networks are calculated separately. The QoS rewards are dynamically calculated based on the service type and current network performance, accurately reflecting the service quality of services under different networks. The difference between the comprehensive scores of the terrestrial and satellite networks is obtained by subtracting the comprehensive score of the terrestrial network from the comprehensive score of the satellite network. This difference intuitively reflects the performance difference between the two networks in the current state.

[0150] Specifically, the reward function r adopts a dynamic weighted form, quantifying the achievement of different goals through three core reward items, and adaptively adjusting the weight coefficients according to the characteristics of the scenario to ensure accurate matching between the reward signal and the scenario requirements:

[0151] ;

[0152] in, , which are dynamically adjusted weighting coefficients. These are QoS rewards, stability rewards, and resource efficiency rewards, with values ​​ranging from [0,1] (a higher value indicates a better achievement of the goal).

[0153] Among them, QoS rewards ( This directly reflects the actual experience quality of the service after the switch. Rewards are defined differently based on service type to ensure a strong correlation with core service needs. For real-time services (such as voice and remote control), latency is the core indicator, and the reward formula is:

[0154] ;

[0155] in, The packet loss rate is standardized (range [0,1]). The smaller the packet loss rate (the less packet loss), the higher the reward. For example, a packet loss rate of 0.2% ( )hour, ;

[0156] Stability rewards are used to suppress ping-pong handovers (frequent handovers between satellite and ground networks in a short period of time) and avoid experience degradation caused by service interruptions during handovers. The reward formula is as follows:

[0157] ;

[0158] in, This represents the number of switches within a unit of time (1 minute), with a value range of 0, 1, 2, ... (in actual scenarios, more than 5 switches are considered ping-pong switches). For each additional switch, the reward decreases by 0.2. When the number of switches is ≥ 5, This establishes a penalty mechanism. For example, if the user switches twice within one minute, When switching 6 times, Forcefully suppress high-frequency switching.

[0159] Resource efficiency rewards are used to optimize the overall utilization of satellite-to-ground network resources, avoiding overload or idle resources in a single network. The reward formula is as follows:

[0160] ;

[0161] in, This represents the bandwidth utilization rate of ground base stations (%). This is the satellite beam load factor (%, derived from "number of active users / maximum capacity"). Design logic: the smaller the difference between satellite and ground load (the more balanced the resources), the higher the reward. For example, when the ground load is 60% and the satellite load is 50% (a difference of 10%), When the ground load is 90% and the satellite load is 30% (difference of 60%), When one network is at 100% load and the other network is at 10% load (a difference of 90%), This guides decision-making towards optimizing load balancing.

[0162] Based on the above weighting coefficients Adjust in real time according to scene characteristics to ensure that the reward function prioritizes the core needs of the current scene, including: high-mobility scenes, static scenes and default scenes;

[0163] Among them, high mobile scenarios (terminal speed) For example, the core requirement of high-speed rail and aviation is to suppress ping-pong switching, hence the setting (Stability has the highest weight) (QoS is secondary) (Least efficient), resource conflict scenarios ( or satellite payload :

[0164] The core requirement for static scenarios is to alleviate network overload, hence the setting. Efficiency has the highest weight. (QoS is secondary) (Lowest stability).

[0165] The core requirement of the default scenario (non-high mobility, non-resource conflict) is to ensure user service experience, therefore it is set as follows: (QoS weight is the highest) Stability is secondary. (Least efficient).

[0166] For Actor networks, a deep deterministic policy gradient algorithm is typically used for training. This algorithm, combined with experience replay and a target network mechanism, ensures stable convergence during training. The specific steps are as follows: Construct the Actor (policy network) and Critic (value network): The Actor network takes a 7-dimensional state vector S as input and outputs a 3-dimensional action A (decision weight); the Critic network takes S and A as input and outputs the action value. (Evaluate the merits of the current action). Initialize the target Actor network. ) and the target Critic network The parameters are the same as the initial network; the experience replay pool capacity is set to... (Balancing training efficiency and data diversity to avoid training fluctuations caused by sample correlation). The agent interacts with a satellite-ground network simulation environment (simulating scenarios such as high mobility and load fluctuations), collecting the "state-action-reward-next state" quadruple in each time slot. The data is stored in the experience replay pool. It includes three typical scenarios: high mobility (300km / h high-speed rail), resource conflict (90% ground load), and mixed scenarios (medium speed + 70% satellite load) to ensure sample coverage. Each training iteration randomly samples 32 sets of data (mini-batch) from the replay pool, and calculates the temporal difference (TD) error using the Critic network.

[0167] ;

[0168] in, This is a discount factor (balancing immediate and long-term rewards). The target Critic network outputs the action. The Critic network parameters are updated with the goal of minimizing the TD error. Then, the Actor network parameters are updated using the policy gradient ascent algorithm, guided by the action value output by the Critic network (to make the Actor tend to output higher-value actions). r is the immediate reward, returned by the environment after performing action a in state s. γ is the discount factor (0 < γ ≤ 1), balancing immediate and long-term rewards; in this paper, it is set to 0.9. s′ is the next state transitioned to after performing action a (i.e., S′). Actor(s′) is the optimal action a′ (three-dimensional weight vector) output by the target Actor network in state s′.

[0169] Qt(s′, Actor(s′)) is the target Critic network's value estimate of the "next state-action pair" (i.e., the target Q-value). Q(s, a) is the online Critic network's value estimate of the current state-action pair (i.e., the current Q-value).

[0170] The target network parameters are updated using a soft update mechanism, and the current network parameters are proportionally merged with the target network parameters after each training iteration.

[0171] ;

[0172] in, The soft update rate (small step updates to avoid sudden changes in the target network) is used until the average reward fluctuation of 100 consecutive training rounds is ≤2%, which is considered convergence.

[0173] After training, the Actor network is solidified into a decision model, and in real-time scenarios, a comprehensive score is output according to the following process:

[0174] Based on the current state S, the agent outputs the optimal weights through the Actor network:

[0175] ;

[0176] Calculate the combined scores for the terrestrial network and the satellite network separately:

[0177] ;

[0178] ;

[0179] in, QoS rewards for terrestrial / satellite networks (i.e., under the current network) value).

[0180] In this optional embodiment, by comprehensively considering multiple dimensions such as signal quality, resource load, and service QoS, the performance of terrestrial and satellite networks can be evaluated more comprehensively, thereby making more accurate handover decisions. Simultaneously, the dynamic calculation mechanism of QoS rewards allows the system to flexibly adjust handover decisions based on different service types and real-time network conditions, ensuring service continuity and user experience. Furthermore, the calculation of the comprehensive score difference between the satellite and terrestrial networks provides a clear quantitative basis for handover decisions, reducing the possibility of misjudgments and improving the accuracy and reliability of handover.

[0181] Optionally, determining the network handover target and the decision confidence level of the network handover target based on the comprehensive score difference of the satellite-to-ground network includes:

[0182] A dynamic lag threshold is determined by linear calculation based on the real-time movement speed on the user side.

[0183] The network handover target is determined based on the relationship between the combined score difference between the satellite and ground networks and the dynamic lag threshold.

[0184] If the combined score difference between the satellite and ground networks is greater than or equal to the dynamic lag threshold, then the network switching target is determined to be the satellite network.

[0185] If the combined score difference between the satellite and ground networks is less than the dynamic lag threshold, then the network switching target is determined to be the ground network;

[0186] The combined score difference of the satellite-ground network is substituted into the Sigmoid mapping function to obtain the decision confidence level.

[0187] Specifically, a hysteresis threshold is dynamically generated based on a linear calculation of the user's real-time movement speed. The dynamic hysteresis threshold is designed to consider network handover requirements in different mobile scenarios. For example, in high-speed mobile scenarios (such as high-speed rail), the hysteresis threshold is increased accordingly to reduce frequent handovers (ping-pong handovers) caused by rapid signal changes. Next, the combined score difference between the satellite and terrestrial networks is compared with the dynamic hysteresis threshold, and the network handover target is determined based on the comparison result. If the combined score difference is greater than or equal to the dynamic hysteresis threshold, the handover target is the satellite network; otherwise, the handover target is the terrestrial network. Finally, the combined score difference is substituted into the Sigmoid mapping function to calculate the decision confidence level, which is used to evaluate the reliability of the current handover decision.

[0188] This embodiment introduces a dynamic hysteresis threshold. To suppress ping-pong switching, the judgment rule is: if If the network connection is active, the decision is to "switch to satellite network"; otherwise, the decision is to "maintain terrestrial network connection". The calculation formula is:

[0189] ;

[0190] because Increases with increasing terminal speed (high mobility scenarios) (higher), therefore static scene : (Low threshold, flexible switching); High-speed rail scenario : (High threshold, strictly suppressing handover).

[0191] In this optional embodiment, by introducing a dynamic hysteresis threshold, the system can flexibly adjust the hysteresis of handover decisions according to the movement speed of different users, effectively reducing ping-pong handover phenomena, especially in high-speed movement scenarios, significantly improving service continuity. The reliability of the current handover decision is quantified using the Sigmoid mapping function, providing an important basis for subsequent execution strategy selection. Furthermore, the handover decision-making mechanism based on dynamic hysteresis threshold and decision confidence can significantly improve network resource utilization, reduce service interruption time, and enhance network adaptability and robustness.

[0192] Optionally, determining the execution parameters of the network switching target using a Bayesian execution engine based on the decision confidence includes:

[0193] The execution parameters are determined by performing Bayesian inference based on the decision confidence and a preset confidence threshold using the Bayesian execution engine.

[0194] Wherein, if the decision confidence level is greater than or equal to the preset confidence threshold, the execution parameters are single-link signaling transmission mode and HARQ retransmission disabled;

[0195] If the decision confidence level is less than the preset confidence threshold, then the execution parameters are dual-link parallel transmission mode and enabling one HARQ retransmission.

[0196] Specifically, the Bayesian execution engine uses decision confidence and a preset confidence threshold to perform Bayesian inference to determine execution parameters. Decision confidence reflects the reliability of the current handover decision, while the preset confidence threshold is a pre-defined reference value used to distinguish between high-confidence and low-confidence decisions. The preset confidence threshold can be set according to the specific application, and will not be elaborated further here. Specifically, when the decision confidence is greater than or equal to the preset confidence threshold, it indicates that the current handover decision has high reliability. Therefore, a single-link signaling transmission mode is selected and HARQ retransmission is disabled to reduce handover latency and ensure rapid service continuity. Conversely, when the decision confidence is less than the preset confidence threshold, it indicates that the current handover decision has some uncertainty. To improve handover reliability, a dual-link parallel transmission mode is selected and one HARQ retransmission is enabled to increase the success rate of signaling transmission and the reliability of service data.

[0197] First, based on the comprehensive score difference between the satellite and ground networks, the abstract decision reliability is transformed into a calculable confidence value, providing a priori basis for Bayesian execution.

[0198] Let the ground network score be the output of reinforcement learning decisions. Satellite network score Define decision confidence level The formula for calculating the range of values ​​([0,1]) is:

[0199] ;

[0200] Where k=10 is the sensitivity coefficient (verified by actual testing, this coefficient can make the confidence level ≥0.8 when the score difference >0.1 and the confidence level ≤0.6 when the score difference <0.05, with the best discrimination).

[0201] when When this occurs, it is determined to be a "high-confidence decision" (the decision logic is reliable and no redundant execution is required).

[0202] when When this occurs, it is judged as a "low-confidence decision" (the decision is uncertain and its execution reliability needs to be strengthened).

[0203] Specifically, for the two core execution stages of signaling transmission link and HARQ retransmission, the strategy is dynamically adjusted based on confidence level and service type:

[0204] For signaling link selection, the transmission link of signaling (such as handover requests and resource allocation instructions) directly affects the handover latency and needs to be selected based on confidence level differentiation: high confidence ( When the confidence level is low, the ground signaling link is preferred—the ground link transmission latency is <10ms (far lower than the 50-100ms of the satellite link), and the decision error rate is <1% under high confidence, eliminating the need for dual-link redundancy to improve reliability and minimizing the total handover latency. When a satellite is used, a "dual-link parallel transmission" strategy is adopted, with the ground link as the primary link and the satellite link as the backup link. Valid signaling is selected using Bayesian posterior probability: Let the prior probability of valid signaling reception be... Based on historical data, with an initial value set to 0.9 (representing a default link reliability), the real-time observed signaling reception strength is... (such as ground signaling) Satellite signaling At that time, the likelihood probability ,on the contrary The posterior probability of the signaling being valid is:

[0205] ;

[0206] Ultimately, the link signaling with the higher posterior probability is selected as the valid instruction to ensure a signaling delivery success rate of ≥99.5% (avoiding handover failure due to single link loss).

[0207] HARQ retransmission strategy:

[0208] The number of retransmissions in Hybrid Automatic Repeat Request (HARQ) directly impacts service QoS—real-time services are sensitive to latency, while non-real-time services are sensitive to packet loss, requiring dynamic adjustment based on service type: For real-time services (such as voice and remote control), the core requirement is to control interruption latency, and HARQ retransmission is disabled by default; it is only disabled when there is low confidence. Furthermore, when packet loss is observed (packet loss rate > 0.5%), one emergency retransmission is initiated (retransmission latency < 20ms, far below the 100ms tolerance limit for real-time services), avoiding the cumulative latency caused by multiple retransmissions. For non-real-time services (such as file downloads and video caching), the core requirement is to reduce the packet loss rate, based on the packet loss probability estimated by Bayes. Calculate the number of retransmissions N dynamically. Assume service reliability requirements are met. (i.e., the final successful reception probability is ≥99.9%), then the formula for the number of retransmissions is:

[0209] ;

[0210] in, Estimate using "historical packet loss data + current SNR" (e.g., when satellite SNR = 10dB). When the ground SNR is 15dB, Ensure that the final packet loss rate for non-real-time services is ≤0.1%.

[0211] In this optional embodiment, by dynamically adjusting execution parameters through Bayesian inference, the signaling transmission method and retransmission strategy can be flexibly selected based on the confidence level of the handover decision. This reduces handover latency and improves service continuity under high-confidence decisions, while increasing handover reliability and reducing service interruptions under low-confidence decisions. Secondly, this dynamic adjustment mechanism can significantly reduce ping-pong handover, especially in high-mobility scenarios, effectively reducing frequent handovers caused by rapid signal fluctuations. Furthermore, by enabling a HARQ retransmission, the reliability of service data can be improved to a certain extent, reducing packet loss and enhancing user experience.

[0212] Optionally, the dynamic switching method for satellite-ground heterogeneous networks further includes:

[0213] After the network switch is completed, the execution effect data is obtained, including the switch success rate, service interruption duration, and service packet loss rate.

[0214] The switching success rate, the service interruption duration, and the service packet loss rate are quantified into a comprehensive feedback value according to a preset period.

[0215] The reward function weights and Actor network parameters of the reinforcement learning model are dynamically updated based on the comprehensive feedback value.

[0216] Specifically, after the network handover is completed, to facilitate subsequent closed-loop optimization, key performance indicators (KPIs) after the handover execution need to be quantified into feedback values, i.e., execution performance data, to feed back into the preceding reinforcement learning decision model. Specifically, three types of core feedback indicators need to be collected: handover success rate R (success = 1, failure = 0; failure scenarios include signaling loss and packet loss after retransmission); service interruption time T (duration from initiating the handover to service recovery, in milliseconds); and service packet loss rate P (the percentage of service data packets lost during execution, in %). These indicators are then integrated into a comprehensive feedback value F (range [-1, 1], positive values ​​represent better-than-expected results, negative values ​​represent degradation), calculated using the following formula:

[0217] ;

[0218] in, (Maximum tolerable downtime for satellite-to-ground handover). After each preset feedback cycle handover is completed, F is input into the reinforcement learning model to update the reward function weights and Actor network parameters, thereby achieving iterative optimization of the "execution-decision" process.

[0219] Every 10 switches (or every 30 seconds, whichever comes first, balancing statistical and real-time data), three types of core feedback indicators are collected and standardized into a "comprehensive deviation value" to quantify the degree of performance degradation:

[0220] Switching success rate deviation ,definition ,in (Target success rate of satellite-to-ground switching). This represents the percentage of successful handovers out of 10 attempts; if This indicates that the success rate has not met the target, and the decision-making or execution strategy needs to be revised first.

[0221] Service QoS degradation deviation is calculated differently based on service type: Real-time services: ,in (Maximum tolerable interruption time for real-time services); Non-real-time services: ,in (Maximum tolerable packet loss rate for non-real-time services); If If the QoS degradation is greater than 0, the weight of the QoS parameters in the perception layer needs to be adjusted or the retransmission strategy in the execution layer needs to be implemented.

[0222] Ping-Pong Switching Rate, Definition ,in Minutes (statistics window) The number of ping-pong handovers within 10 handover cycles (≥2 satellite-to-ground back-and-forth handovers in a short period are considered ping-pong), with a threshold set at 5 times / minute; if Therefore, it is necessary to strengthen the stability constraints of the decision-making level.

[0223] The above indicators are combined into a comprehensive deviation value. (Values ​​range [0,1], with larger values ​​indicating more severe performance degradation), the formula is:

[0224] ;

[0225] The weights (0.5 / 0.3 / 0.2) are set based on user experience priority—the success rate of handover has the greatest impact on user perception, followed by QoS degradation, and finally ping-pong handover.

[0226] If the overall deviation E ≤ 0.1 for three consecutive rounds (10 switches per round) (performance degradation is within an acceptable range), or the parameter correction magnitude is < 1% (parameters are close to optimal), the iteration is paused and the current parameter configuration is fixed; if the network environment undergoes a sudden change (such as a sharp drop in SNR > 15dB due to satellite beam switching, or a change in the proportion of service types > 50% (such as real-time services increasing from 20% to 70%), a new round of iteration is forcibly triggered to ensure that the system can quickly adapt to new scenarios.

[0227] In this optional embodiment, by collecting and quantifying handover execution effect data, the system can monitor and evaluate the performance of handover operations in real time, ensuring the effectiveness and reliability of handover decisions. Dynamically updating the reward function weights and Actor network parameters of the reinforcement learning model allows the system to continuously optimize its decision-making strategy based on actual execution results, thereby achieving long-term performance improvements. The closed-loop optimization mechanism significantly improves the handover success rate, reduces service interruption time and packet loss rate, and enhances the network's adaptability and robustness.

[0228] Combination Figure 2 As shown in the figure, an embodiment of the present invention provides a dynamic handover system for heterogeneous satellite-to-ground networks, comprising:

[0229] The multi-dimensional state awareness module is used to collect core state parameters from the network side, user side, and service side to obtain state parameters related to network handover for the network side, user side, and service side respectively; and to perform dynamic correlation processing on the state parameters corresponding to the network side, user side, and service side respectively to generate a unified state vector.

[0230] The reinforcement learning decision module is used to perform dynamic decision optimization based on the unified state vector through a reinforcement learning model to obtain satellite-to-ground handover decision weights; generate a satellite-to-ground network comprehensive score difference based on the satellite-to-ground network comprehensive score difference; and determine the network handover target and the decision confidence of the network handover target based on the satellite-to-ground network comprehensive score difference.

[0231] The execution module is used to determine the execution parameters of the network switching target based on the decision confidence using a Bayesian execution engine, and to perform network switching based on the execution parameters.

[0232] The advantages of the satellite-ground heterogeneous network dynamic switching system in this embodiment compared to the prior art are the same as the advantages of the satellite-ground heterogeneous network dynamic switching method compared to the prior art, and will not be repeated here.

[0233] An electronic device provided by an embodiment of the present invention includes a memory and a processor;

[0234] The memory is used to store computer programs;

[0235] The processor is used to implement the above-described method for dynamic switching of heterogeneous satellite-ground networks when executing the computer program.

[0236] The electronic device in this embodiment has the same advantages over the prior art as the above-mentioned dynamic switching method for heterogeneous satellite-ground networks, and will not be repeated here.

[0237] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A method for dynamic handover in a satellite-ground heterogeneous network, characterized in that, include: Core status parameters are collected from the network side, user side, and service side respectively to obtain the status parameters related to network handover for the network side, the user side, and the service side respectively; Dynamic association processing is performed on the state parameters corresponding to the network side, the user side, and the service side respectively to generate a unified state vector; By using a reinforcement learning model, dynamic decision optimization is performed based on the unified state vector to obtain the satellite-to-ground switching decision weights. Based on the satellite-to-ground handover decision weights, a satellite-to-ground network comprehensive score difference is generated, specifically including: determining the comprehensive score of the ground network and the comprehensive score of the satellite network on the network side based on the satellite-to-ground handover decision weights and the QoS rewards corresponding to the ground network and satellite network respectively; subtracting the comprehensive score of the ground network from the comprehensive score of the satellite network and using the difference as the satellite-to-ground network comprehensive score difference; and determining the network handover target and the decision confidence level of the network handover target based on the satellite-to-ground network comprehensive score difference, specifically including: performing linear calculation based on the real-time mobile speed on the user side to determine a dynamic lag threshold; determining the network handover target based on the relationship between the satellite-to-ground network comprehensive score difference and the dynamic lag threshold; wherein, if the satellite-to-ground network comprehensive score difference is greater than or equal to the dynamic lag threshold, the network handover target is determined to be the satellite network; if the satellite-to-ground network comprehensive score difference is less than the dynamic lag threshold, the network handover target is determined to be the ground network; and substituting the satellite-to-ground network comprehensive score difference into the Sigmoid mapping function to obtain the decision confidence level. The execution parameters for the network handover target are determined using a Bayesian execution engine based on the decision confidence level. Specifically, this includes: using the Bayesian execution engine to perform Bayesian inference based on the decision confidence level and a preset confidence threshold to determine the execution parameters; wherein, if the decision confidence level is greater than or equal to the preset confidence threshold, the execution parameters are single-link signaling transmission mode and HARQ retransmission disabled; if the decision confidence level is less than the preset confidence threshold, the execution parameters are dual-link parallel transmission mode and one HARQ retransmission enabled; and network handover is performed based on the execution parameters.

2. The dynamic handover method for heterogeneous satellite-to-ground networks according to claim 1, characterized in that, The process involves collecting core status parameters from the network side, user side, and service side respectively to obtain status parameters related to network handover on the network side, user side, and service side, including: The resource status of the network side is collected to obtain the ground base station parameters of the ground network and the low-orbit satellite parameters of the satellite network, and the ground base station parameters and the low-orbit satellite parameters are used as the status parameters of the network side. The user on the user side is located using GPS or base station positioning data from the terminal, and the user's real-time movement speed is obtained. The real-time movement speed is then used as the status parameter on the user side. By collecting data on the differences in service types on the service side, the QoS index of the service side is obtained, and the QoS index is used as the status parameter of the service side.

3. The dynamic switching method for heterogeneous satellite-ground networks according to claim 2, characterized in that, The step of dynamically associating the state parameters corresponding to the network side, the user side, and the service side respectively to generate a unified state vector includes: Sliding window filtering is applied to the reference signal received power in the ground base station parameters and the link signal-to-noise ratio in the low-orbit satellite parameters to obtain smoothed ground parameters and smoothed satellite parameters. When the bandwidth utilization rate in the ground base station parameters is greater than or equal to the preset load correction threshold, the smoothed ground parameters are corrected according to the preset overload ratio to obtain the final ground parameters. When the bandwidth utilization rate in the ground base station parameters is less than the preset load correction threshold, the smoothed ground parameters are used as the final ground parameters. The QoS indicators are normalized according to the service type to obtain standardized QoS parameters; The final ground parameters, the smoothed satellite parameters, and the standardized QoS parameters are concatenated into a unified state vector.

4. The method for dynamic handover of heterogeneous satellite-ground networks according to claim 2, characterized in that, The step of using a reinforcement learning model to perform dynamic decision optimization based on the unified state vector to obtain satellite-to-ground switching decision weights includes: The unified state vector is input into the reinforcement learning model; The signal quality weight, resource load weight, and service QoS weight are obtained by using the Actor network of the reinforcement learning model and performing Softmax normalization based on the unified state vector. A three-dimensional continuous weight matrix is ​​formed based on the signal quality weight, the resource load weight, and the service QoS weight; wherein the sum of the signal quality weight, the resource load weight, and the service QoS weight is 1. The three-dimensional continuous weight matrix is ​​used as the decision weight for satellite-to-ground switching.

5. The dynamic handover method for heterogeneous satellite-ground networks according to claim 1, characterized in that, Also includes: After the network switch is completed, the execution effect data is obtained, including the switch success rate, service interruption duration, and service packet loss rate. The switching success rate, the service interruption duration, and the service packet loss rate are quantified into a comprehensive feedback value according to a preset period. The reward function weights and Actor network parameters of the reinforcement learning model are dynamically updated based on the comprehensive feedback value.

6. A dynamic handover system for heterogeneous satellite-ground networks, characterized in that, include: The multi-dimensional state awareness module is used to collect core state parameters from the network side, user side, and service side to obtain state parameters related to network handover for the network side, user side, and service side respectively; and to perform dynamic correlation processing on the state parameters corresponding to the network side, user side, and service side respectively to generate a unified state vector. The reinforcement learning decision module is used to perform dynamic decision optimization based on the unified state vector through a reinforcement learning model to obtain satellite-to-ground handover decision weights; generate a satellite-to-ground network comprehensive score difference based on the satellite-to-ground network comprehensive score difference; and determine the network handover target and the decision confidence of the network handover target based on the satellite-to-ground network comprehensive score difference. Specifically, generating the satellite-to-ground network comprehensive score difference based on the satellite-to-ground handover decision weights includes: determining the comprehensive score of the ground network and the comprehensive score of the satellite network on the network side based on the satellite-to-ground handover decision weights and the QoS rewards corresponding to the ground network and the satellite network respectively; subtracting the comprehensive score of the ground network from the comprehensive score of the satellite network, and using the difference as the satellite-to-ground network comprehensive score difference; The process of determining the network handover target and its decision confidence level based on the combined score difference between the satellite and ground networks specifically includes: determining a dynamic lag threshold by linear calculation based on the real-time mobile speed of the user; determining the network handover target based on the relationship between the combined score difference between the satellite and ground networks and the dynamic lag threshold; wherein, if the combined score difference between the satellite and ground networks is greater than or equal to the dynamic lag threshold, the network handover target is determined to be the satellite network; if the combined score difference between the satellite and ground networks is less than the dynamic lag threshold, the network handover target is determined to be the terrestrial network; and substituting the combined score difference between the satellite and ground networks into a Sigmoid mapping function to obtain the decision confidence level. The execution module is used to determine the execution parameters of the network switching target based on the decision confidence using a Bayesian execution engine, and to perform network switching based on the execution parameters; Specifically, the execution parameters for the network handover target are determined by the Bayesian execution engine based on the decision confidence level. This includes: using the Bayesian execution engine to perform Bayesian inference based on the decision confidence level and a preset confidence threshold to determine the execution parameters; wherein, if the decision confidence level is greater than or equal to the preset confidence threshold, the execution parameters are single-link signaling transmission mode and HARQ retransmission disabled; if the decision confidence level is less than the preset confidence threshold, the execution parameters are dual-link parallel transmission mode and HARQ retransmission enabled once.

7. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the dynamic switching method for heterogeneous satellite-ground networks as described in any one of claims 1-5 when executing the computer program.

Citation Information

Patent Citations

  • Full-time communication guarantee method and system based on 5G, satellite flash and satellite communication

    CN118826836A

  • Network switching method and device, terminal and storage medium

    CN119109784A