Satellite-ground fusion network multi-target space-time switching decision-making method, system and device and storage medium

By adopting the dual mask dual-deep Q network (BiMasQ) method in the satellite-ground fusion network, the problem of ignoring the independent operation system and future network status in the prior art is solved, and a more stable network connection and higher service quality and experience quality are achieved.

CN120223155APending Publication Date: 2025-06-27HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510337180.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The connection switching decision-making method of the existing satellite-ground converged network ignores the impact of the independent operation system of the LEO satellite network and the ground network and the future network status on the current decision, resulting in unstable network connections and unable to meet high-quality service (QoS) and quality of experience (QoE) requirements.

Method used

A multi-objective space-time switching decision-making method based on dual mask dual-deep Q network (BiMasQ) is proposed. By establishing a system model, defining performance indicators and objective functions, modeling the connection switching decision process, and using the continuous interaction between the agent and the environment, we learn the optimal connection switching strategy.

Benefits of technology

This method reduces the total number of handovers during transmission, extends the average connection time, improves the average throughput, reduces latency, and optimizes QoS and QoE from the two dimensions of time and space, significantly improving the stability and performance of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005322079250000033
    Figure BDA0005322079250000033
  • Figure BDA0005322079250000034
    Figure BDA0005322079250000034
  • Figure BDA0005322079250000035
    Figure BDA0005322079250000035
Patent Text Reader

Abstract

The invention relates to a satellite-ground fusion network multi-target space-time switching decision-making method. The method comprises the following steps: S1, establishing system models based on a real environment and a Starlink satellite constellation, wherein the system models comprise a terminal mobile model, a network model and a communication model; s2, defining performance indexes of service quality and experience quality, and constructing a target function; s3, modeling a connection switching decision process: modeling the connection switching decision process as a Markov decision process; s4, proposing a space-time decision method for connection switching based on the double-mask double-depth Q network: obtaining an optimal switching strategy by utilizing continuous interaction between an intelligent agent and an environment; in the exploration stage and the utilization stage, action masks are used respectively to improve the learning efficiency of the intelligent agent in the high-dimensional action space. According to the method provided by the invention, the total switching frequency in the transmission process is reduced, the average connection duration is prolonged, the average throughput is improved, the delay is reduced, and meanwhile, dense switching in the transmission process is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer networks, and in particular to a method, system, device and storage medium for multi-target spatiotemporal switching decision-making in a satellite-ground fusion network. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and massive machine-type communications (mMTC), the global demand for high-quality and wide-coverage networks is increasing. In particular, with the increasing number of IoT terminals, traditional terrestrial networks face limitations in geographical coverage. Terrestrial networks mainly cover densely populated cities and areas, while remote rural and oceanic areas often cannot obtain adequate network services. According to statistics, terrestrial networks can only cover about 20% of the earth's land area, and the total global coverage is even less than 6%. In order to make up for the shortcomings of terrestrial networks, satellite-terrestrial integrated networks (STINs) have gradually become a solution with broad prospects. Satellite-terrestrial integrated networks combine low Earth orbit (LEO) satellites with terrestrial networks to provide high-quality network services worldwide. Compared with traditional single terrestrial networks, satellite networks can provide wider coverage, especially in remote areas, providing valuable spectrum resources and communication capabilities, thereby effectively supporting the access and data transmission of large-scale IoT terminals.

[0003] However, the high dynamics of LEO satellites and some IoT terminals also bring new challenges. Due to the relatively low orbital altitude of LEO satellites, their orbital movement speed is relatively fast, which requires terminals in the network to frequently switch with LEO satellites and base stations to ensure continuous network connection. To solve this problem, the satellite-ground fusion network architecture requires a complete connection switching strategy to ensure efficient and stable network connection to meet the growing quality of service (QoS) and quality of experience (QoE) requirements. There are two major problems with the existing connection switching strategy. The first is that it ignores the current independence of the LEO satellite network and the ground network operation system, which makes the existing network-based mobility (NBM) management used in homogeneous network switching difficult to implement; the second is that it ignores the impact of future network status on current decisions. Therefore, designing a complete connection switching decision method for the satellite-ground fusion network is a key issue that needs to be solved in the current technical field. Summary of the invention

[0004] To solve the technical problem that the existing connection switching decision method ignores the independent operating systems of LEO satellite networks and terrestrial networks and the impact of future network states on current decisions, the present invention provides a multi-objective spatio-temporal switching decision method and system for satellite-terrestrial integrated networks.

[0005] A multi-objective spatio-temporal switching decision method for a satellite-terrestrial integrated network according to the present invention, the method comprising the following steps:

[0006] S1. Establish a system model based on the real environment and the Starlink satellite constellation: including a terminal movement model, a network model, and a communication model;

[0007] S2. Define performance indicators for quality of service and quality of experience, and construct an objective function;

[0008] S3. Model the connection switching decision process: model the connection switching decision process as a Markov decision process;

[0009] S4. Propose a spatio-temporal decision method for connection switching based on a double-mask double deep Q-network: use an agent to continuously interact with the environment to obtain an optimal switching strategy; use action masks in the exploration phase and the exploitation phase respectively to improve the learning efficiency of the agent in the high-dimensional action space.

[0010] Further, S1 specifically includes the following steps:

[0011] S11. Establish a terminal movement model

[0012] Establish a terminal movement model based on the real environment, and the movement speed of the terminal is v;

[0013] S12. Establish a network model

[0014] Establish a discrete-time satellite-terrestrial integrated network model based on the two-line orbital data of Starlink, the terminal movement model, and the deployment locations of mobile communication base stations, with a sampling period of δ, in seconds;

[0015] S13. Establish a communication model

[0016] Establish satellite-terrestrial and terrestrial communication models based on 3GPP and ITU-R technical reports:

[0017] The satellite-terrestrial link communication model is:

[0018] where SNR sa represents the signal-to-noise ratio, in dB; represents the transmit power, in dBm; represents the transmit gain, in dB; Represents the G / T value of the satellite antenna, with the unit of dB·K -1 ; L sa Represents the total path loss of the satellite - to - ground link, with the unit of dB; k B Represents the Boltzmann constant, with the unit of dBm·K -1 ·Hz -1 ; B represents the bandwidth, with the unit of Hz;

[0019] Total path loss of the satellite - to - ground link: L sa = L F + L P + L S + L A #(2)

[0020] Among them, L F Represents the free - space loss; L P Represents the polarization loss; L S Represents the shadow margin; L A Represents the total atmospheric loss, and the units are all dB;

[0021] Ground - link communication model:

[0022] Among them, SNR te Represents the signal - to - noise ratio, with the unit of dB; Represents the transmit power, with the unit of dBm; Represents the transmit gain, with the unit of dB; G r Represents the receive gain of the base station, with the unit of dBi; L te Represents the total path loss of the ground link, with the unit of dB; N0 represents the noise figure, with the unit of dBm·Hz -1 ; B represents the bandwidth, with the unit of Hz;

[0023] Total path loss of the ground link: L te = L RMa + L P + L S + L B + L V #(4)

[0024] Among them, L RMa Represents the propagation loss under the rural macro - base - station line - of - sight model; L P Represents the polarization loss; L S Represents the shadow margin; L B Represents the human body loss; L V Represents the vegetation loss, and the units are all dB;

[0025] According to the Shannon formula, the satellite - to - ground and ground - channel capacities:

[0026] Among them, C is the channel capacity, with the unit of bps; is the dimensionless form of the signal-to-noise ratio.

[0027] Furthermore, S2 specifically includes the following steps:

[0028] S21. Define the performance indicators of QoS and QoE

[0029] Number of handovers: The total number of handovers during the entire transmission process, denoted as H;

[0030] Average connection duration: The mean value of the time of each connection (without experiencing handovers), denoted as μ d ;

[0031]

[0032] Among them, d i represents the duration of the i-th connection, with the unit of discrete time sampling period; for example, d8 = 3 means that the duration of the 8th connection is 3 sampling periods, that is, 3δ, with the unit of s;

[0033] Throughput: The average data rate during the entire transmission process, denoted as Θ, with the unit of bps;

[0034] LEO satellite delay: The propagation time of the data packet from the sender through the LEO satellite to the gateway station, denoted as L;

[0035]

[0036] Among them, l represents the delay of a single LEO satellite; R uplink represents the propagation time of the data packet from the sender to the LEO satellite; R backhaul represents the propagation time from the LEO satellite to the gateway station, and the units are all ms;

[0037] Badness index: An index that quantifies the density of handover events during the entire transmission process, denoted as β;

[0038]

[0039] Among them, W represents the time window size; M represents the total number of all "bad" windows; D represents the threshold for judging whether a window is "bad";

[0040] S22. Construct the objective function

[0041] Objective function and constraints:

[0042]

[0043] Among them, T represents the total number of discrete time nodes in the entire transmission process; represents the set of all available LEO satellites and base stations at time t.

[0044] Furthermore, S3 specifically includes the following steps:

[0045] S31. Define the state space

[0046] State space:

[0047] Among them, N represents the total number of nodes;

[0048] S32. Define the action space

[0049] Action space:

[0050] S33. Define the reward function

[0051] Reward function: R = w·s - C HO #(12)

[0052] Among them, R represents the immediate reward; w represents the weight vector of the performance index; s is the performance index vector; C HO is the handover cost;

[0053] Handover cost:

[0054] Among them, k represents the time weight coefficient; b represents the basic handover cost; C het represents the heterogeneous handover cost; I het is an indicator function; α represents the penalty coefficient when exceeding the maximum number of handovers;

[0055]

[0056] Furthermore, S4 specifically includes the following steps:

[0057] S41. Construct a spatio-temporal handover decision framework

[0058] Incorporate the time dimension into the handover decision framework, and DDQN is used as a variant of the Q-network in reinforcement learning:

[0059]

[0060] Among them, y represents the target Q value; represents the immediate reward for taking action a in state s; γ ∈ [0, 1] represents the discount factor, which is used to balance the immediate reward and the long-term return; Q target and Q respectively represent the target network and the online network; s' and a' represent the state and action at the next moment; θ targetθ represents the parameters of the target network and the online network respectively;

[0061] S42. Initialize the action mask

[0062] The combined network model introduces an action mask to shield the satellites that are unavailable at the current moment;

[0063] Action mask:

[0064] Among them, M(t) represents the action mask vector at time t; θ i (t) represents the elevation angle of the i-th satellite at time t, with the unit of degree;

[0065] S43. Train the BiMasQ model

[0066] The training process of BiMasQ is divided into an exploration stage and an exploitation stage: In the exploration stage, the agent randomly selects actions in the action space after applying the action mask to fully explore the unknown environment; in the exploitation stage, the agent makes decisions through the online network and uses the learned knowledge to optimize the policy performance; when the reward value of a single episode converges, the optimal connection switching policy is obtained;

[0067] Loss function (mean squared error):

[0068] Parameter update: Among them, α represents the learning rate;

[0069] S44. Apply the optimal connection switching policy

[0070] The online network with convergent reward value provides the optimal connection switching policy for the terminal. The input of the online network is the current state s, and the output is the optimal action a * .

[0071] Furthermore, S43 specifically includes the following steps:

[0072] S431. Agent exploration

[0073] In the exploration stage, the agent continuously and randomly selects actions in the action space after applying the action mask, forming a series of five-tuples composed of the current state, the current action, the immediate reward, the next state, and the done flag, that is, state transition, denoted as (s t , a t , r t , s t+1 , done). These state transitions are stored in the experience replay buffer; during training, randomly select a batch of experiences with a batch size of B from the experience replay buffer as training data; the action mask prevents the agent from exploring the unavailable nodes at the current moment;

[0074] S432. The agent utilizes

[0075] In the utilization stage, the agent adopts the action corresponding to the maximum Q value. When selecting an action, the action mask prevents the state transition of unavailable nodes from being stored in the experience replay buffer, enabling the agent to focus only on learning the switching strategy without considering the availability of nodes.

[0076] Furthermore, S44 specifically includes the following steps:

[0077] S441. State acquisition

[0078] The terminal obtains the current location of the terminal at the current moment t or through the positioning module, and the node currently connected to the terminal, to obtain the current state s;

[0079] S442. Action selection

[0080] Input the current state s into the BiMasQ network to obtain the optimal action a * , that is, the switching target node at the next moment, or maintain the connection with the current node;

[0081] S443. Online learning

[0082] BiMasQ continuously collects and records the state transitions of the terminal in the real environment and conducts online learning to adapt to the real application scenario and the satellite-terrestrial fusion network.

[0083] The present invention also relates to a satellite-terrestrial fusion network multi-objective spatio-temporal switching decision-making system, which includes a computer module that utilizes the above-mentioned satellite-terrestrial fusion network multi-objective spatio-temporal switching decision-making method.

[0084] The present invention also relates to a computer device, including a memory and a processor. The memory stores a computer program, and it is characterized in that when the processor executes the computer program, the steps of the satellite-terrestrial fusion network multi-objective spatio-temporal switching decision-making method are implemented.

[0085] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the satellite-terrestrial fusion network multi-objective spatio-temporal switching decision-making method are implemented.

[0086] Beneficial effects

[0087] The method proposed by the present invention reduces the total number of handovers during transmission, extends the average connection duration, improves the average throughput, reduces the latency, and at the same time avoids intensive handovers during transmission, optimizing QoS and QoE from both the time and space dimensions. During the training process, by applying action masks in the exploration and exploitation phases, the learning efficiency of the agent and the model performance are effectively improved, and the optimal connection handover strategy is obtained.

[0088] The present invention learns the optimal connection handover strategy with the highest long-term reward through the continuous interaction between the agent and the environment, avoiding the pre-design of complex handover mechanisms and being able to well adapt to the dynamically changing satellite-ground integrated network environment. Through the online learning method, the terminal can continuously adapt to the dynamically changing environment and update the optimal connection handover strategy to achieve adaptive adjustment. The present invention has robustness. When the optimal node given by the optimal handover strategy has too high a load or fails, a connection handover strategy with overall performance similar to that of the optimal strategy can still be found to ensure all QoS and QoE performance indicators. Description of the Drawings

[0089] Figure 1 It is a schematic diagram of the connection handover decision of the satellite-ground integrated network of the present invention;

[0090] Figure 2 It is a schematic diagram of the connection handover model of the satellite-ground integrated network of the present invention;

[0091] Figure 3 It is the architecture diagram of the dual-mask dual deep Q-network (BiMasQ) in the present invention;

[0092] Figure 4 It is a schematic diagram of the reward values during the training process of the present invention, standard DDQN, dueling DQN and DQN;

[0093] Figure 5 It is a schematic diagram of the comparison of the total number of handovers between the present invention, greedy algorithm, threshold-based method, standard DDQN, dueling DQN and DQN;

[0094] Figure 6 It is a schematic diagram of the comparison of the connection duration between the present invention, greedy algorithm, threshold-based method, standard DDQN, dueling DQN and DQN;

[0095] Figure 7 It is a schematic diagram of the comparison of the badness index between the present invention, greedy algorithm, threshold-based method, standard DDQN, dueling DQN and DQN;

[0096] Figure 8Schematic diagram of the average throughput comparison between the present invention and the greedy algorithm, threshold-based method, standard DDQN, dueling DQN, and DQN;

[0097] Figure 9 Schematic diagram of the latency comparison between the present invention and the greedy algorithm, threshold-based method, standard DDQN, dueling DQN, and DQN;

[0098] Figure 10 Schematic diagram of the robustness analysis of the multi-objective spatio-temporal handover decision method for the satellite-terrestrial integration network of the present invention in terms of the total number of handovers;

[0099] Figure 11 Schematic diagram of the robustness analysis of the multi-objective spatio-temporal handover decision method for the satellite-terrestrial integration network of the present invention in terms of the connection duration;

[0100] Figure 12 Schematic diagram of the robustness analysis of the multi-objective spatio-temporal handover decision method for the satellite-terrestrial integration network of the present invention in terms of the badness index;

[0101] Figure 13 Schematic diagram of the robustness analysis of the multi-objective spatio-temporal handover decision method for the satellite-terrestrial integration network of the present invention in terms of the average throughput;

[0102] Figure 14 Schematic diagram of the robustness analysis of the multi-objective spatio-temporal handover decision method for the satellite-terrestrial integration network of the present invention in terms of the latency;

[0103] Figure 15 Schematic diagram of the reward values during the training process of the present invention and standard DDQN, DDQN applying action masks only in the exploration phase, and DDQN applying action masks only in the exploitation phase;

[0104] Figure 16 Schematic diagram of the total number of handovers comparison between the present invention and standard DDQN, DDQN applying action masks only in the exploration phase, and DDQN applying action masks only in the exploitation phase;

[0105] Figure 17 Schematic diagram of the connection duration comparison between the present invention and standard DDQN, DDQN applying action masks only in the exploration phase, and DDQN applying action masks only in the exploitation phase;

[0106] Figure 18 Schematic diagram of the badness index comparison between the present invention and standard DDQN, DDQN applying action masks only in the exploration phase, and DDQN applying action masks only in the exploitation phase;

[0107] Figure 19Schematic diagram of the average throughput comparison between the present invention, standard DDQN, DDQN applying action masking only in the exploration stage, and DDQN applying action masking only in the exploitation stage;

[0108] Figure 20 Schematic diagram of the latency comparison between the present invention, standard DDQN, DDQN applying action masking only in the exploration stage, and DDQN applying action masking only in the exploitation stage. Detailed implementation manners

[0109] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but it is not limited to the present invention.

[0110] A multi-objective spatio-temporal switching decision method for a space-ground integrated network of the present invention includes the following steps:

[0111] Step S1: Establish a system model based on the real farm environment and the Starlink satellite constellation: including a terminal movement model, a network model, and a communication model.

[0112] In S1, it includes the following steps:

[0113] S11: Establish a terminal movement model

[0114] Based on the real environment, establish a movement model for the Internet of Things terminal when performing operations, and the movement speed of the terminal is v.

[0115] S12: Establish a network model

[0116] Based on the two-line element (TLE) data of Starlink, the movement model of the terminal, and the deployment location of the mobile communication base station, establish a discrete-time space-ground integrated network model, and the sampling period is δ, with the unit of s.

[0117] S13: Establish a communication model

[0118] Based on the 3GPP and ITU-R technical reports, establish space-ground and ground communication models.

[0119] Space-ground link communication model:

[0120]

[0121] Among them, SNR sa represents the signal-to-noise ratio, with the unit of dB; represents the transmit power, with the unit of dBm; represents the transmit gain, with the unit of dB; represents the G / T value of the satellite antenna, with the unit of dB·K -1 ; L saRepresents the total path loss of the satellite - to - ground link, with the unit of dB; k B Represents the Boltzmann constant, with the unit of dBm·K -1 ·Hz -1 ; B represents the bandwidth, with the unit of Hz.

[0122] Total path loss of the satellite - to - ground link:

[0123] L sa = L F + L P + L S + L A #(2)

[0124] Among them, L F represents the free - space loss; L P represents the polarization loss; L S represents the shadow margin; L A represents the total atmospheric loss, and the units are all dB.

[0125] Terrestrial link communication model:

[0126]

[0127] Among them, SNR te represents the signal - to - noise ratio, with the unit of dB; represents the transmit power, with the unit of dBm; represents the transmit gain, with the unit of dB; G r represents the receive gain of the base station, with the unit of dBi; L te represents the total path loss of the terrestrial link, with the unit of dB; N0 represents the noise factor, with the unit of dBm·Hz -1 ; B represents the bandwidth, with the unit of Hz.

[0128] Total path loss of the terrestrial link:

[0129] L te = L RMa + L P + L S + L B + L V #(4)

[0130] Among them, L RMa represents the propagation loss under the rural macro - base - station line - of - sight model; L P represents the polarization loss; L S represents the shadow margin; L B represents the human body loss; L V represents the vegetation loss, and the units are all dB.

[0131] According to the Shannon formula, the satellite - to - ground and terrestrial channel capacities:

[0132]

[0133] Among them, C is the channel capacity, with the unit of bps; is the dimensionless form of the signal-to-noise ratio.

[0134] Step S2: Define the performance indicators of quality of service (QoS) and quality of experience (QoE), and construct the objective function;

[0135] S2 includes the following steps:

[0136] S21: Define the performance indicators of QoS and QoE

[0137] Number of handovers:

[0138] The total number of handovers during the entire transmission process, denoted as H.

[0139] Average connection duration:

[0140] The mean value of the time of each connection (without experiencing handovers), denoted as μ d .

[0141]

[0142] Among them, d i represents the duration of the i-th connection, with the unit of discrete time sampling period. For example, d8 = 3 means that the duration of the 8th connection is 3 sampling periods, i.e., 3δ, with the unit of s.

[0143] Throughput:

[0144] The average data rate during the entire transmission process, denoted as Θ, with the unit of bps.

[0145] LEO satellite delay:

[0146] The propagation time of the data packet from the sending end through the LEO satellite to the gateway station, denoted as L.

[0147]

[0148] Among them, l represents the delay of a single LEO satellite; R uplink represents the propagation time of the data packet from the sending end to the LEO satellite; R backhaul represents the propagation time from the LEO satellite to the gateway station, and the units are all ms.

[0149] Badness index:

[0150] An indicator for quantifying the density of handover events throughout the transmission process, denoted as β. The more times the terminal initiates handovers continuously in a short period, the larger β is; conversely, the smaller β is.

[0151]

[0152] Among them, W represents the time window size; M represents the total number of all "bad" windows; D represents the threshold for determining whether a window is "bad".

[0153] S22. Construct the objective function

[0154] Objective function and constraints:

[0155]

[0156] Among them, T represents the total number of discrete time nodes throughout the transmission process; represents the set of all available LEO satellites and base stations at time t.

[0157] Step S3. Model the connection handover decision-making process: Model the connection handover decision-making process as a Markov decision-making process.

[0158] In S3, the following steps are included:

[0159] S31. Define the state space

[0160] State space:

[0161]

[0162] Among them, N represents the total number of nodes.

[0163] S32. Define the action space

[0164] Action space:

[0165]

[0166] S33. Define the reward function

[0167] Reward function:

[0168] R = w·s - C HO #(12)

[0169] Among them, R represents the immediate reward; w represents the weight vector of the performance index; s is the performance index vector; C HO is the handover cost.

[0170] Handover cost:

[0171]

[0172] Among them, k represents the time weight coefficient; b represents the basic handover cost; C het represents the heterogeneous handover cost; I het is an indicator function; α represents the penalty coefficient when the maximum number of handovers is exceeded.

[0173]

[0174] Step S4: Propose a spatio-temporal decision-making method for connection handover based on DDQN (BiMasQ): Consider the impact of future network states on the current decision-making, and use the continuous interaction between the agent and the environment to obtain the optimal handover strategy; Use action masks in the exploration and exploitation phases respectively to improve the learning efficiency of the agent in the high-dimensional action space, and at the same time improve the quality of experience replay, thereby improving the model performance.

[0175] S4 specifically includes the following steps:

[0176] S41: Construct a spatio-temporal handover decision-making framework

[0177] Incorporate the time dimension into the handover decision-making framework, that is, consider the impact of future network states on the current handover decision-making. As a variant of the Q-network in reinforcement learning, DDQN can fully consider the balance between immediate rewards and long-term returns, and at the same time well solve the problem of overestimation of Q-values in DQN.

[0178] Long-term return:

[0179]

[0180] Among them, y represents the target Q-value; represents the immediate reward for taking action a in state s; γ ∈ [0, 1] represents the discount factor, which is used to balance immediate rewards and long-term returns; Q target and Q respectively represent the target network and the online network; s' and a' represent the state and action at the next moment; θ target and θ respectively represent the parameters of the target network and the online network.

[0181] S42: Initialize the action mask

[0182] Due to the large number of satellites in the Starlink constellation (currently exceeding 5000), the dimension of the action space generated by it is extremely high, which in turn leads to low search efficiency of the agent and a decline in model performance. Therefore, by combining the network model, an action mask is introduced to shield the satellites that are unavailable at the current moment, thereby improving the learning efficiency of the agent in the high-dimensional space and further improving the performance of the model.

[0183] Action mask:

[0184]

[0185] Among them, M(t) represents the action mask vector at time t. θ i (t) represents the elevation angle of the i-th satellite at time t, with the unit of degree.

[0186] S43. Training the BiMasQ model

[0187] The training process of BiMasQ is divided into an exploration stage and a exploitation stage. In the exploration stage, the agent randomly selects actions in the action space after applying the action mask to fully explore the unknown environment; in the exploitation stage, the agent makes decisions through the online network to fully utilize the learned knowledge to optimize the policy performance. When the reward value of a single episode converges, the optimal connection switching policy is obtained.

[0188] Loss function (mean squared error):

[0189]

[0190] Parameter update:

[0191]

[0192] Among them, α represents the learning rate.

[0193] S43 specifically includes the following steps:

[0194] S431. Agent exploration

[0195] In the exploration stage, the agent continuously randomly selects actions in the action space after applying the action mask, forming a series of five-tuples composed of the current state, current action, immediate reward, next state, and completion flag, that is, state transition, denoted as (s t , a t , r t , s t+1 , done). These state transitions are stored in the experience replay buffer. During training, a batch of experiences with a batch size of B is randomly selected from the experience replay buffer as training data. The action mask can prevent the agent from exploring unavailable nodes at the current moment, effectively improving the initial reward value of the agent.

[0196] S432. Agent exploitation

[0197] In the exploitation stage, the agent adopts the action corresponding to the maximum Q value instead of randomly selecting actions. When selecting actions, the action mask can prevent the state transitions of unavailable nodes from being stored in the experience replay buffer, effectively improving the quality of the training data in the experience replay buffer. The action mask enables the agent to only focus on learning the switching strategy without considering the availability of nodes, further improving the performance of the model.

[0198] S44. Apply the optimal connection switching strategy

[0199] The online network with converging reward values can provide the optimal connection switching strategy for the terminal. The input of the online network is the current state s, and the output is the optimal action a * .

[0200] S44 specifically includes the following steps:

[0201] S441. State acquisition

[0202] The terminal obtains the current state s by the current moment t (or by the positioning module to obtain the current location of the terminal) and the node currently connected by the terminal.

[0203] S442. Action selection

[0204] Input the current state s into the BiMasQ network to obtain the optimal action a * , that is, the switching target node at the next moment (or maintain the connection with the current node).

[0205] S443. Online learning

[0206] BiMasQ continuously collects and records the state transitions of the terminal in the real environment and conducts online learning to gradually adapt to the real application scenario and the space-ground integrated network.

[0207] Embodiment

[0208] As Figure 1 shown, in this embodiment, a space-ground integrated network composed of LEO satellites, their gateway stations and mobile communication base stations is constructed. This network provides network services for a smart farm. The farm environment is an area with a length of 12 km and a width of 6 km. UAVs and autonomous driving agricultural machines are equipped with multi-modal IoT terminals to drive straight in the farmland and perform operations at a speed of 5 m / s.

[0209] During the movement of the LEO satellite and the IoT terminal, if the terminal does not switch at any time with the change of the network topology, it will lead to frequent connection interruptions, and it is difficult to guarantee QoS and QoE. However, the traditional connection switching decision method ignores the independent operation system of the LEO satellite network and the ground network and the influence of the future network state on the current decision. Therefore, based on the balance mechanism of long-term reward and immediate reward in reinforcement learning, the present invention incorporates the future network state into the decision-making framework and uses the continuous interaction between the agent and the environment to learn the optimal connection switching strategy with the highest long-term reward, avoiding the pre-design of complex switching mechanisms and being able to well adapt to the dynamically changing space-ground integrated network environment, as Figure 2 shown.

[0210] In one embodiment of the present invention, a reinforcement learning training environment is established, and the state space, action space, and reward function are defined, where the reward function is a linear weighted sum of QoS and QoE performance metrics. During the training process, by applying action masks in the exploration and exploitation phases, as Figure 3 shown, the learning efficiency of the agent and the model performance are effectively improved, and the optimal connection switching strategy is obtained.

[0211] In one embodiment of the present invention, the specific process of the double-mask double deep Q-network (BiMasQ) method is as follows.

[0212] Step 1: Initialize the online network and the target network. The online network is used to select actions and estimate Q-values; the target network is used to calculate the target Q-values.

[0213] Step 2: Initialize the experience replay buffer, which is used to store state transitions, that is, a five-tuple containing the current state, current action, immediate reward, next state, and completion flag.

[0214] Step 3: Initialize (reset) the satellite-ground fusion network environment and obtain the starting state.

[0215] Step 4: Obtain the action mask at the current time.

[0216] Step 5: Based on the ε-greedy policy, determine whether it is the exploration phase or the exploitation phase at present.

[0217] Step 6: When in the exploration phase, select the current action from the action space applied with the action mask; when in the exploitation phase, select the action with the largest Q-value as the current action.

[0218] Step 7: Execute the action and generate the corresponding state transition.

[0219] Step 8: Repeat Steps 3 to 7, and store all state transitions in the experience replay buffer.

[0220] Step 9: Randomly select a batch of state transitions from the experience replay buffer to train the online network, and update the target network parameters regularly.

[0221] Step 10: Update the exploration rate of the ε-greedy policy.

[0222] Step 11: Repeat Steps 3 to 10 until the reward value increases and finally converges, and the loss function decreases and finally converges.

[0223] The pseudocode of this embodiment is as follows:

[0224]

[0225]

[0226] In one embodiment of the present invention, a simulation platform based on a system model is adopted to evaluate the performance of the method in the present invention. A total of 8 methods including the double-mask double deep Q-network (BiMasQ), the performance evaluation baseline method, and the ablation experiment baseline method in the present invention are verified. Among them, the performance evaluation baseline methods include the greedy algorithm, the threshold-based method, the standard DDQN, the dueling DQN, and the DQN; the ablation experiment baseline methods include the standard DDQN, the DDQN that applies the action mask only in the exploration stage, and the DDQN that applies the action mask only in the exploitation stage. The simulation parameters of the communication model are shown in Table 1, and the hyperparameters of the double-mask double deep Q-network (BiMasQ) are shown in Table 2.

[0227] Table 1 Simulation parameters of the communication model

[0228]

[0229] Table 2 Hyperparameters of the double-mask double deep Q-network

[0230]

[0231]

[0232] In one embodiment of the present invention, Figure 4 Illustrate the change of the reward value during the training process of the method in the present invention, the standard DDQN, the dueling DQN, and the DQN. The reward value corresponding to the method in the present invention continuously rises and finally converges, and the reward value is higher than the other three methods, indicating its optimal performance.

[0233] In one embodiment of the present invention, Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 and Figure 9 Illustrate the total number of handovers, connection duration, badness index, average throughput, and latency performance of the method in the present invention and the performance evaluation baseline methods respectively. The method in the present invention is Pareto dominant in terms of the total number of handovers, connection duration, and badness index. Since the greedy algorithm has reached the theoretical limit of throughput and latency, and no method can be better than the greedy algorithm in terms of throughput and latency performance, the method in the present invention is better than all performance evaluation baseline methods except the greedy algorithm in terms of throughput and latency performance.

[0234] In one embodiment of the present invention, Figure 10 、 Figure 11 、 Figure 12 、 Figure 13 、 Figure 14Illustrate the total number of handovers, connection duration, badness metric, average throughput, and latency performance of the method in the present invention when the optimal node load is too high or a failure occurs. In terms of throughput performance, when all 6 optimal nodes are unavailable, the method in the present invention can still maintain the optimal average throughput; in terms of the total number of handovers, connection duration, badness metric, and latency performance, when up to 11 optimal nodes are unavailable, the method in the present invention can still maintain the optimal performance, indicating that the method in the present invention can still ensure the optimal performance when the environment changes and has robustness.

[0235] In an embodiment of the present invention, Figure 15 Illustrate the change in the reward value during the training process of the method in the present invention and the ablation experiment baseline method. Applying the action mask effectively improves the initial reward value during the exploration phase; applying the action mask effectively improves the final reward value and the stability of training during the exploitation phase, demonstrating the effectiveness of the action mask.

[0236] In an embodiment of the present invention, Figure 16 、 Figure 17 、 Figure 18 、 Figure 19 、 Figure 20 Illustrate the total number of handovers, connection duration, badness metric, average throughput, and latency performance of the method in the present invention and the ablation experiment baseline method respectively. The method in the present invention is Pareto dominant in all the above metrics, further demonstrating the effectiveness of the action mask.

[0237] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not deviate from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.

Claims

1. A multi-objective spatiotemporal switching decision method for a satellite-ground fusion network, characterized in that: The method comprises the following steps: S1. Establish a system model based on the real environment and Starlink satellite constellation: including terminal mobility model, network model and communication model; S2. Define the performance indicators of service quality and experience quality, and construct the objective function; S3. Modeling the connection switching decision process: Modeling the connection switching decision process as a Markov decision process; S4. A spatiotemporal decision method for connection switching based on a dual-mask dual-depth Q network is proposed: the optimal switching strategy is obtained by using the continuous interaction between the agent and the environment; Action masks are used in the exploration phase and the exploitation phase to improve the learning efficiency of the agent in high-dimensional action space.

2. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 1 is characterized in that: S1 specifically includes the following steps: S11. Establish terminal mobility model A terminal mobility model is established based on the real environment, and the terminal's moving speed is v; S12. Establishing a network model Based on Starlink's two-line orbit data, the terminal's mobility model, and the deployment location of mobile communication base stations, a discrete-time satellite-ground fusion network model is established with a sampling period of δ in seconds. S13. Establishing a communication model Establish satellite-to-ground and ground communication models based on 3GPP and ITU-R technical reports: The satellite-to-ground link communication model is: Among them, SNR sa Indicates the signal-to-noise ratio in dB; Indicates the transmit power in dBm; Indicates the transmission gain in dB; Indicates the G / T value of the satellite antenna in dB·K -1 ; L sa represents the total path loss of the satellite-to-ground link in dB; k B Represents the Boltzmann constant, in dBm·K -1 ·Hz -1 ; B represents bandwidth, unit is Hz; Total path loss of satellite-to-ground link: L sa =L F +L P +L S +L A #(2) Among them, L F Indicates free space loss; L P Represents polarization loss; L S Indicates the shadow margin; L A It represents the total atmospheric loss, and the unit is dB; Ground link communication model: Among them, SNR te Indicates the signal-to-noise ratio in dB; Indicates the transmit power in dBm; Indicates the transmission gain in dB; G r Indicates the receiving gain of the base station, in dBi; L te It represents the total path loss of the ground link in dB; N0 represents the noise factor in dBm·Hz -1 ; B represents bandwidth, unit is Hz; Total path loss of the ground link: L te =L RMa +L P +L S +L B +L V #(4) Among them, L RMa represents the propagation loss under the line-of-sight model of rural macro base stations; L P Represents polarization loss; L S Indicates the shadow margin; L B Indicates human body loss; L V Indicates vegetation loss, the unit is dB; According to Shannon's formula, the satellite-to-ground and ground channel capacities are: Where C is the channel capacity in bps; is the dimensionless form of the signal-to-noise ratio.

3. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 1 is characterized in that: S2 specifically includes the following steps: S21. Define performance indicators for QoS and QoE Number of switches: the total number of switches in the entire transmission process, denoted as H; Average connection time: the average time of each connection (without switching), denoted as μ d ; Among them, d i Indicates the duration of the i-th connection, in units of discrete time sampling periods; for example, d8=3 means that the duration of the 8th connection is 3 sampling periods, i.e. 3δ, in units of seconds; Throughput: The average data rate during the entire transmission process, denoted as θ, in bps; LEO satellite delay: The propagation time of a data packet from the sender through the LEO satellite to the gateway, denoted as L; Where, l represents the delay of a single LEO satellite; R uplink R represents the propagation time of the data packet from the sender to the LEO satellite; backhaul It represents the propagation time from the LEO satellite to the gateway, in ms; Badness index: an index that quantifies the density of switching events during the entire transmission process, denoted as β; Where W represents the time window size; M represents the total number of all "bad" windows; D represents the threshold for determining whether a window is "bad"; S22. Constructing the objective function Objective function and constraints: Where T represents the total number of discrete time nodes in the entire transmission process; Represents the set of all available LEO satellites and base stations at time t.

4. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 1 is characterized in that: S3 specifically includes the following steps: S31. Define the state space State Space: Where N represents the total number of nodes; S32. Define action space Action Space: S33. Define reward function Reward function: R = w · sC HO #(12) Among them, R represents the immediate reward; w represents the weight vector of the performance indicator; s is the performance indicator vector; C HO is the switching cost; Switching cost: Where k represents the time weight coefficient; b represents the basic switching cost; C het I represents the heterogeneous switching cost; het is an indicator function; α represents the penalty coefficient when the maximum number of switching times is exceeded; 5. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 1 is characterized in that: S4 specifically includes the following steps: S41. Constructing a time-space switching decision framework Incorporating the time dimension into the switching decision framework, DDQN is a variant of the Q network in reinforcement learning: Where y represents the target Q value; represents the immediate reward for taking action a in state s; γ∈[0,1] represents the discount factor, which is used to balance the immediate reward and long-term return; Q target and Q represent the target network and the online network respectively; s' and a' represent the state and action at the next moment; θ target and θ represent the parameters of the target network and the online network, respectively; S42, initialization action mask The action mask is introduced in combination with the network model to shield the satellites that are unavailable at the current moment; Action Mask: Where M(t) represents the action mask vector at time t; θ i (t) represents the elevation angle of the i-th satellite at time t, in degrees; S43. Training BiMasQ model The training process of BiMasQ is divided into an exploration phase and an exploitation phase. In the exploration phase, the agent randomly selects actions in the action space after applying the action mask to fully explore the unknown environment. In the exploitation phase, the agent makes decisions through the online network and uses the learned knowledge to optimize the strategy performance. When the reward value of a single round converges, the optimal connection switching strategy is obtained. Loss function (mean square error): Parameter update: Among them, α represents the learning rate; S44. Apply the optimal connection switching strategy The online network with converged reward value provides the terminal with the optimal connection switching strategy. The input of the online network is the current state s, and the output is the optimal action a. * .

6. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 5 is characterized in that: S43 specifically includes the following steps: S431, Agent Exploration In the exploration phase, the agent continuously randomly selects actions in the action space after applying the action mask, forming a series of five-tuples consisting of the current state, current action, immediate reward, next moment state and completion flag, namely, state transition, expressed as (s t ,a t ,r t ,s t+1 , done), these state transfers are stored in the experience replay buffer; during training, experiences with a batch size of B are randomly selected from the experience replay buffer as training data; Action masks prevent the agent from exploring nodes that are unavailable at the current moment; S432, Intelligent Agent Utilization In the utilization phase, the agent takes the action corresponding to the maximum Q value. When selecting an action, the action mask prevents the state transition of unavailable nodes from being stored in the experience replay buffer. The action mask allows the agent to focus on learning the switching strategy without considering the availability of the node.

7. The satellite-ground fusion network multi-objective spatiotemporal switching decision method according to claim 5 is characterized in that: S44 specifically includes the following steps: S441, status acquisition The terminal obtains the current location of the terminal through the current time t or through the positioning module, and the node currently connected to the terminal to obtain the current state s; S442, Action Selection Input the current state s into the BiMasQ network to get the optimal action a * , that is, the target node to switch to at the next moment, or keep the current node connected; S443, Online Learning BiMasQ continuously collects and records the state transitions of terminals in real environments and conducts online learning to adapt to real application scenarios and satellite-ground fusion networks.

8. A satellite-ground fusion network multi-objective spatiotemporal switching decision system, characterized in that: The system includes a computer module, and utilizes the above-mentioned satellite-ground fusion network multi-objective spatiotemporal switching decision method.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-target spatiotemporal switching decision method for a satellite-ground fusion network are implemented.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a spatiotemporal decision-making method for switching a satellite-ground fusion network connection are implemented.

Citation Information

Cited By

  • Communication equipment multi-link automatic switching control method based on reinforcement learning

    CN121396704A

  • Non-cooperative satellite network switching behavior inference method and device based on time sequence model

    CN121887258A

  • Non-cooperative satellite network handover behavior inference method and device based on timing model

    CN121887258B