Satellite-ground integrated routing method based on QOS constraint
By using multi-objective deep reinforcement learning agents to select routing paths in LEO satellite networks, the problem that traditional routing algorithms cannot meet the needs of diversified QOS is solved, routing strategies are optimized, network performance and data transmission efficiency are improved, and high dynamic characteristics of LEO satellite networks are adapted to.
Patent Information
- Application Number
- CN202510536346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional routing algorithms cannot meet the diverse QOS needs in LEO satellite networks in the future, especially in high dynamic and high density environments, which leads to conflicts in independent routing decision-making processes of different services. The existing reinforcement learning algorithms do not fully consider the characteristics of LEO satellite networks, resulting in insufficient network performance optimization.
The agent is adopted to select the routing path based on multi-objective deep reinforcement learning, and re-plan the routing path during time slot switching through the agent. Combining the delay, packet loss rate, throughput and energy consumption models, the routing strategy is optimized, and the dynamic changes of the satellite network are adapted to the Actor-Critic training framework for routing decisions.
It effectively solves the differentiated sensitivity of different services to multiple QOS indicators, optimizes routing strategies, improves network performance and data transmission efficiency, adapts to the high dynamic characteristics of LEO satellite networks, and simplifies algorithm complexity.
Smart Images

Figure CN120264377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly to a space-ground integrated routing method based on QoS constraints. Background Art
[0002] With the emergence of autonomous vehicles and the development of the Internet of Things (IoT), a large number of widely deployed interconnected terminals have put forward diverse requirements for network communication, such as fast response and ubiquitous access. Low Earth Orbit (LEO) satellites are expected to make up for the shortcomings of terrestrial communication in terms of seamless coverage, especially in sparsely populated areas. In addition, compared with Geostationary Orbit (GEO) satellites, LEO satellites have significant advantages of low latency and high throughput. Currently, LEO satellite communication has become an important part of 6G.
[0003] However, LEO satellite constellations such as Starlink and OneWeb are composed of a large number of high-speed moving LEO satellites at an altitude of 200 - 3000 kilometers from the Earth's surface, which makes the LEO satellite network have high-density and high-dynamic characteristics. Therefore, traditional shortest-path-based routing algorithms designed for terrestrial networks and small satellite networks are no longer suitable for future LEO satellite networks. The frequent dynamics and ultra-density of LEO satellite networks pose significant challenges to the timely update and fast convergence of end-to-end routing strategy design. In addition, most traditional routing solutions aim to optimize a single Quality of Service (QoS) metric, such as latency, throughput, packet loss rate, or energy consumption metric, to reduce computational complexity, but cannot meet the diverse QoS requirements of various terrestrial applications. Future satellite network routing strategies should simultaneously meet the requirements of services for multiple metrics. Summary of the Invention
[0004] To improve the user experience, the present invention proposes a space-ground integrated routing method based on QoS constraints, including: using an agent to select a routing path. When two time slots are switched, the current satellite node determines the data packets that were not transmitted in the previous time slot, and judges whether the routing path in the previous time slot is blocked in the current time slot. If it is blocked, the current node is used as the starting routing point and the target node is used as the routing end point, and a routing path is reselected for the data packets not transmitted by the current node according to the agent.
[0005] Further, at the current satellite node, the minimum number of hops from adjacent satellites to the egress satellite, the latency required for one hop from the current satellite to an adjacent satellite, the node load level of the adjacent satellite, the link reliability from the current satellite to the adjacent satellite, and the Boolean value status indicating whether the adjacent satellite has already been a routing node in another routing path are used as the state vector to input into the agent, and the agent outputs the routing path between the current satellite and the destination satellite.
[0006] Compared with the prior art, the present invention has the following beneficial effects:
[0007] 1. In terms of intelligent routing, existing intelligent routing algorithms mostly focus on single-objective routing and only optimize single QoS metrics such as latency, throughput, packet loss rate, etc., and it is difficult to meet the diverse requirements of various services for multiple metrics in future satellite networks. Therefore, when dealing with the different sensitivities of different services to multiple QoS metrics, existing methods may lead to conflicts in the independent routing decision-making processes of different services. In contrast, the present invention designs a general utility function to reflect the sensitivities of various services to different QoS metrics according to the QoS requirements of different services, thereby solving the problem of conflicts in the independent routing decision-making processes of different time-sensitive services.
[0008] 2. In terms of algorithm implementation, although there are already satellite routing algorithms based on reinforcement learning, most of them do not fully consider the ultra-dense and highly dynamic characteristics of LEO satellite networks. The present invention uses multi-objective deep reinforcement learning for routing decisions, models the path selection problem as a decentralized partially observable Markov decision process, simplifies the satellite network state space, reduces the algorithm complexity. In this way, it can better adapt to the dynamic changes of satellite networks, optimize routing strategies, and improve network performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is the space-ground integrated network architecture in a space-ground integrated network routing method based on a constrained agent according to the present invention;
[0010] Figure 2 It is a schematic diagram of specific time slot division in a space-ground integrated network routing method based on a constrained agent according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0012] The present invention proposes a space-ground integrated routing method based on QoS constraints, including: using an agent to select a routing path. When two time slots are switched, the current satellite node determines the data packets that were not transmitted in the previous time slot, and determines whether the routing path in the previous time slot is blocked in the current time slot. If it is blocked, the current node is used as the starting routing point and the target node is used as the routing end point, and a routing path is reselected for the data packets that were not transmitted by the current node according to the agent.
[0013] In recent years, some companies and enterprises have been building their own LEO satellite constellations to provide low-latency and high-throughput global communication services, especially in the fields of ocean, desert, air and even space. All satellites operate in circular orbits at the same altitude and inclination, and the orbits are evenly distributed along the equator, while the satellites are evenly distributed within their respective orbits. The present invention adopts a similar constellation, which can be expressed as: W T / W P / W F / W h / W i , where W T 、W P 、W h and W i represent the total number of satellites, the number of orbital planes, the altitude of the orbit, and the orbital inclination respectively, and W F is the phase factor. According to the characteristics of the Walker constellation, the number of satellites on each orbital plane can be expressed as S = W T / W P .
[0014] Let G(V, E) represent the satellite network studied in the present invention, where V and E represent nodes and edges respectively. In addition, in the present invention, the node V refers to the set of low Earth orbit satellites, and the position of each satellite can be determined by {v i,j |0 ≤ i < W P , 0 ≤ j < S}. For convenience, v k=i×S+j is used to replace v i,j . Therefore, the nodes in the satellite network can be expressed as {v k ∈ V|k < W T}.
[0015] Due to the time-varying characteristics of the LEO satellite constellation, the present invention divides the system period T into a series of time intervals, that is, T rans (abbreviated as T r ) and T stable (abbreviated as T S ), which are also called the switching state and the stable state respectively. As Figure 2, a system cycle includes multiple stable states and switching states. In a stable state, its topological structure is considered fixed (i.e., topological snapshot), and the time interval of this state is [t 2n , t 2n+1 . That is, during the time from t 2n to t 2n+1 , the link is considered fixed; the switching state is responsible for completing the topological switch between adjacent topological snapshots within the time interval [t 2n-1 , t 2n . This state involves multiple Inter-Satellite Link (ISL) switches, that is, during the time from t 2n-1 to t 2n , the link switch is performed. The link switch includes link removal and link establishment, that is, removing the link that hinders communication in the current time slot and cannot communicate, and taking the satellite node where the data packet is currently located as the source node to re-establish the path between the source node and the destination node, and adding the links involved in this path to the path topological structure. After that, a two-dimensional array CS[E][N] is used to represent the connection relationship of the link between two nodes in a specific time slot. If CS[E][N]=1, it means that the E-th link is available in the N-th time slot; if it is 0, it means that the link is disconnected. Therefore, a routing selection matrix S is defined, whose elements are CS[E][N], and it has E rows and N columns:
[0016]
[0017] For example, considering a total of three available routes in four time slots, that is, N = 4, E = 3, the routing selection matrix S is as follows:
[0018]
[0019] The above matrix indicates that in the 1st and 2nd time slots, link 1 is in a connected state; in the 3rd time slot, link 2 is in a connected state and link 1 is disconnected; in the 4th time slot, link 3 is in a connected state and link 2 is disconnected.
[0020] To describe diverse QoS requirements, the present invention classifies applications into four types: delay-sensitive, high-reliability, throughput-sensitive, and energy-consumption-sensitive, which prioritize delay, packet loss rate, throughput, or energy consumption, respectively. Specifically, the user side includes a variety of application scenarios, such as Internet of Vehicles (IoV), intelligent manufacturing, smart home, smart healthcare, virtual reality (VR), etc. These application scenarios have different requirements for quality of service (QoS). For example, intelligent manufacturing is very sensitive to the packet loss rate. Even a slight packet loss can cause the production line to paralyze, resulting in unaffordable economic losses. VR has very strict delay requirements. Any delay exceeding 20 milliseconds will cause discomfort symptoms such as dizziness and nausea. In smart home applications, high throughput is the key to ensuring smooth transmission of multimedia content; while in satellite networks, due to limited energy of satellites, energy-consumption-sensitive tasks are crucial for extending satellite lifespan and reducing operating costs. To characterize the transmission cost, the present invention will elaborate on the satellite communication model in detail from four aspects: delay model, packet loss model, path load model, and energy consumption model.
[0021] (1) Transmission delay model
[0022] The end-to-end transmission delay includes queuing delay, propagation delay, forwarding delay, and handover delay. In the satellite network of the present invention, LEO satellites with limited buffers forward data packets in a first-in-first-out (FIFO) manner and follow the M / M / 1 / m queuing model, where M represents that the distribution of the arrival interval time of consecutive data packets and the distribution of service time follow an exponential distribution, 1 represents that there is only one queue in the router to process data packets, and m is the buffer capacity. In this queuing model, the occupancy ρ can be defined as:
[0023]
[0024] where λ represents the rate at which data packets arrive at the LEO satellite, and μ represents the data packet forwarding rate. The occupancy reflects the busyness of the satellite in processing data packets. When ρ is closer to 1, it indicates that the satellite is busier and the queuing delay may be larger. Therefore, the data packet forwarding delay of the LEO satellite can be represented by represented.
[0025] The average number of data packets queued in the LEO satellite buffer can be expressed as:
[0026]
[0027] The first term of the formula is the average number of queued data packets in the classical M / M / 1 queuing model (infinite buffer case); the second term It is a correction term considering the limited buffer capacity (i.e., the buffer capacity is m), which is used to more accurately calculate the average number of queued data packets in the case of a finite buffer; when ρ is close to 1 (i.e., the satellite load is high), the average number of queued data packets will increase significantly; when ρ is small, the average number of queued data packets is relatively small 1.
[0028] The effective data packet processing rate of LEO satellites can be expressed as:
[0029] λ s = μ(1 - P0)
[0030] where P0 is the probability that the LEO satellite is idle:
[0031]
[0032] Based on the above conditions, we can represent the queuing delay of the LEO satellite node v i as:
[0033]
[0034] The propagation delay of a data packet from the low Earth orbit satellite node v i to the satellite node v j can be calculated according to the space distance as follows:
[0035]
[0036] where the space distance between satellites is represented as In the formula, (x i , y i , z i ) and (x j , y j , z j ) represent the space coordinates of the satellite node v i and the satellite node v j respectively, and c represents the propagation speed of laser in space.
[0037] Since LEO satellites are moving at high speeds, time slot switching will inevitably cause link switching, resulting in handover delay, which will be regarded as a penalty component. When switching the free space optical link between satellites, signal reacquisition is an important factor causing delay, and this time mainly depends on the performance of the satellite's Acquisition, Pointing, and Tracking (APT) system, denoted as T e . In addition, during the link switching process, the satellite equipment needs to perform a series of operations, such as receiving and parsing handover instructions, reconfiguring communication channels, adjusting data transmission buffers, etc. Denote the fixed delay of the satellite equipment processing the handover operation as Tp For convenience, the present invention sets the handover time delay T e and the fixed delay T of the satellite device for processing the handover operation p as fixed values. Therefore, the penalty component generated by one handover is expressed as:
[0038]
[0039] In the solution of the present invention, when the e-th link occupies the n-th time slot at the current node, the value of CS[e][n] is 1, otherwise the value of CS[e][n] is 0. There are only two cases in the handover process from the current time slot to the next time slot, that is, the device needs to be switched in the next time slot and the device does not need to be switched in the next time slot. If a handover is required, it means that the next time slot is not processed at the current node, that is, the value of CS[e][n + 1] is 0. At this time If there is no need to switch time slots, it means that the next time slot is processed at the current node, that is, the value of CS[e][n + 1] is 1. At this time There is no penalty term.
[0040] Therefore, the one-hop delay from the low Earth orbit satellite node v i and the satellite node v j can be calculated as follows:
[0041]
[0042] The total end-to-end path delay can be expressed as:
[0043]
[0044] (2) Packet loss model
[0045] Considering that the buffer capacity of the low Earth orbit satellite is limited, the packet loss rate of the low Earth orbit satellite node v i can be expressed as:
[0046]
[0047] where ρ is the occupancy rate of the satellite buffer, reflecting the busyness of the satellite in processing data packets; m is the buffer capacity; when the satellite load is high (ρ is close to 1) and the buffer capacity is limited, the packet loss rate will increase; the larger the buffer capacity m, the relatively smaller the packet loss rate under the same load.
[0048] The end-to-end path packet loss rate P path can be expressed as:
[0049]
[0050] The above formula is calculated based on the packet loss rate P of each satellite on the path i by multiplying (1 - Pi ) Obtain the probability that data packets are not lost on the path, and then subtract this probability from 1 to obtain the end-to-end path packet loss rate.
[0051] (3) Throughput model
[0052] Since the throughput of each LEO satellite depends on the generated traffic load and adjacent nodes, it cannot represent the actual forwarding ability. Therefore, in the present invention, the delivery rate of each LEO satellite over a period of time is used to represent the throughput, so as to represent the forwarding ability. The end-to-end delivery rate can be expressed as:
[0053]
[0054] where, represents the number of data packets successfully transmitted from the low Earth orbit satellite node v i to the destination within a given time period, represents the total number of data packets sent by the low Earth orbit satellite node v i within a given time period.
[0055] (4) Energy consumption model
[0056] In the LEO satellite network, the satellite energy consumption mainly comes from the following three parts:
[0057] 1) Inter-satellite link transmission;
[0058] 2) Onboard CPU and processors;
[0059] 3) Self-control and self-maintenance of the onboard platform.
[0060] The energy consumed by the self-control and self-maintenance of the onboard platform has nothing to do with the size of the service data volume and is the energy that must be consumed. Therefore, this energy consumption is not considered in the energy consumption model of the present invention.
[0061] In the LEO satellite network, data propagation is mainly affected by free space loss. Based on the free space propagation theorem, p r represents the received power of the satellite, and the calculation formula is as follows:
[0062]
[0063] In the formula, p t represents the transmit power, G r represents the receive gain, G t represents the transmit gain, and L p represents the path loss of free space, and the calculation formula is as follows:
[0064]
[0065] Where λ = c / f is the operating wavelength, c is the speed of light, f is the carrier frequency, and d represents the propagation distance. According to the above formula, the transmission power p t can be calculated as follows:
[0066]
[0067] Where is a fixed parameter. Therefore, the link propagation energy consumption is:
[0068]
[0069] Where the propagation energy consumption k is only related to the propagation time t and the propagation distance d; since the propagation distances d, path losses L p , transmission power p t and received power p r of the space-ground link and the inter-satellite link are not the same, the constant is also not the same.
[0070] For any two adjacent satellite nodes v i transmitting data to satellite node v j the energy consumed for transmitting unit data on the formed link e i,j is expressed as
[0071]
[0072] The processing energy consumption of the satellite is divided into two parts: 1) the startup energy consumption generated by starting related functions such as the on-board CPU and data processor; 2) the data processing energy consumption generated by the data processor processing services. The startup energy consumption is independent of the size of the service and is represented by ε 0 , and the data processing energy consumption is related to the size of the service processed by the node and is represented by . Therefore, for satellite node v i the calculation formula for the processing energy consumption can be expressed as:
[0073]
[0074] Where ε 0 is a constant, ψ and χ are constants determined by the processor itself, and C represents the size of the service data volume processed by the node. The larger the amount of data processed, the greater the processing energy consumption generated. If the amount of data processed is small, the startup energy consumption generated by the processor will be much greater than the data processing energy consumption.
[0075] Therefore, the end-to-end energy consumption can be expressed as:
[0076]
[0077] After obtaining metrics such as latency, packet loss rate, throughput, and energy consumption, it is necessary to define QoS models for different services to reflect different requirements. The present invention defines a utility function to represent the sensitivity of different services to the three QoS metrics considered, and then formulates an optimization objective function.
[0078] Since the value ranges of the three QoS metrics considered are different, normalization is required to ensure fairness, as follows:
[0079]
[0080] Among them, q path represents any one of the metric values in T path , P path , D path , K path ; and q max represent the normalized value and the maximum value of each QoS metric, respectively. Then, the weight ω of each QoS metric is heuristically set according to the QoS requirements. The utility function of each path can be expressed as:
[0081]
[0082] Among them, and represent the normalized values of delay, packet loss rate, and throughput, respectively.
[0083] The multi-objective routing algorithm for multi-services proposed by the present invention aims to minimize the overall utility function under the consideration of transmission power and LEO satellite buffer size constraints, which can be expressed as:
[0084] min U pathi,j , i≠j
[0085] Constraint conditions: C1: w1, w2, w3, w4≥0
[0086] C2: 0≤ω t ≤ω max
[0087] C3: 0≤R t ≤R max
[0088] C4: N q ≤N b
[0089] In the formula represents the routing path from satellite node v i to satellite node v j ; ω t and R tThey are the satellite transmission power and the data transmission rate; ω max and R max represent the maximum transmission power and the data transmission rate respectively; N q and N b represent the data queue length and the buffer size of each LEO satellite respectively.
[0090] The most critical task of routing is to find a path that can satisfy the QOS constraints as much as possible. In the system of the present invention, each satellite acts as an intelligent agent and makes routing decisions independently. By observing the network status of adjacent satellites, the agent selects a suitable next-hop satellite as the intermediate node constituting the path. In this way, the routing discovery process proceeds hop by hop. The actions performed by each satellite are completely based on the current state of the network and do not depend on historical statistical data. This behavior can be effectively described as a Markov decision process: its characteristics can be modeled as a seven-tuple where represents the state of the low Earth orbit (LEO) satellite network, represents the set of actions of the intelligent agent ; represents the conditional probability distribution that the network transfers to the next state s′ after the intelligent agent α takes the action u in the current network state s; represents the reward value obtained by selecting the action set u in the state s; represents the local observable space of each intelligent agent, and the corresponding observation function is expressed as Therefore, the action observation history of each agent can be expressed as The corresponding policy can be expressed as The action-reward function corresponding to the path selection policy can be expressed as where is defined as the discounted return, and γ i is the discount factor. The specific settings are as follows:
[0091] (1) Environmental state: Especially during the state acquisition process before routing decision-making, the present invention assumes that the satellite only acquires the state information of adjacent satellites. Compared with the global acquisition of state information, it can greatly save communication signaling costs.
[0092] On the one hand, propagation delay is the main factor causing end-to-end delay in satellite networks. Therefore, the focus is on forwarding packets along the path with the fewest hops. By comparing the number of orbits and satellites of two satellites, the Bidirectional Manhattan Street Network (MSN) can determine the minimum number of hops. On the other hand, as the traffic in satellite networks continues to increase, queuing delay, as an important component of end-to-end delay, becomes increasingly important. By collecting status information, the one-hop delay from the current node to its adjacent node can be estimated. In addition, load imbalance may also lead to node and link bottlenecks, low network resource utilization efficiency, and affect transmission reliability. Therefore, the node load levels of adjacent satellites also need to be considered.
[0093] Therefore, satellite node v i The composite environmental state obtained at time slot n can be described as:
[0094]
[0095] Among them, H k , O k , A k , R k represent the minimum number of hops from adjacent satellites to the egress satellite, the time delay required for one hop from the current satellite to the adjacent satellite, the node load level of the adjacent satellite, and the link reliability to the adjacent satellite, respectively. In addition, in order to make the multiple paths finally found as non-intersecting as possible, the agent also needs to observe whether the adjacent satellite has become an intermediate node of another path, which can also be represented by the boolean state M k to represent.
[0096] (2) Action: Based on the observed environmental state The agent selects the next-hop forwarding direction u according to the policy α .
[0097] (3) Reward: When the agent takes action u based on the observed state, the environment returns an immediate reward to the agent. According to the received reward, the agent updates the policy until the algorithm converges in the learning phase. The reward value of each agent can be defined according to the utility function of a specific service on the end-to-end routing path, representing the satisfaction with different QOS metrics.
[0098] The algorithm involved in the present invention adopts the Actor-Critic training framework, where the Actor network estimates the policy function π(u|τ) according to the current environmental state s to execute actions, and the Critic network estimates the state value function according to the current environmental state and the next state after executing the action To evaluate the state. The Actor network and the Critic network optimize their respective networks by updating the network parameters θ and φ. θ and φ are updated by the loss functions L(θ) and L(φ) respectively.
[0099] Based on the characteristics of the above algorithm, the model training includes the following steps:
[0100] (1) Network definition: Construct the Actor and Critic neural networks and initialize their network parameters θ and φ;
[0101] (2) Network training: For each episode, there are some end-to-end routing requirements. For each requirement, multiple steps are needed to find the complete path. In each step, the Actor network generates an action selection probability function π(u|τ) according to the collected state information τ, selects the next-hop action u according to the probability distribution of the policy function (state transition probability distribution), then the environment returns the reward r to the agent and moves to the next state s'. Then the experience data <s, u, r, s', π(u|τ), done> is stored in the trajectory memory D.
[0102] (3) When the trajectory memory in the trajectory memory D reaches the long trajectory length ξ, network update is performed. First, randomly sample a batch of trajectory data from the trajectory memory D, and set the batch size to B.
[0103] For the update of the critic network, the present invention adopts the Time Difference (TD) learning method. For each sampled trajectory (s z , u z , r z , s z+1 ), calculate the target value y z = r z + γQ φ (s z+1 , u z+1 ), where u z+1 is the action generated by the actor network according to s z+1 . Then use the mean square error loss function to measure the difference between the target value and the current estimated value. Update the parameters θ of the critic network by the gradient descent method.
[0104] The update of the actor network depends on the feedback of the critic network. To calculate the loss of the actor network, the advantage function A(s z , u z ) = y z - Q φ (s z , u z ) is introduced, which represents the execution of the action u in the state s z z Degree of advantage relative to the average value. The loss function of the actor network The negative sign is used to convert the problem of maximizing the reward into the problem of minimizing the loss. After calculating the loss, the gradient descent method is used to update the parameters θ of the actor network. After the network update is completed, the trajectory memory D is cleared, and new trajectory data is collected again to enter the next round of training and update process. By continuously repeating such a process, the network parameters will gradually converge, and the agent will learn a better routing strategy, thereby improving the performance of the satellite network.
[0105] The trained PPO model is inserted into the satellite for online routing decisions.
[0106] For the stable state in time slot n, the ingress satellite periodically sends discovery packets (DPs) containing source and destination information. When a DP is received, the satellite makes an online routing decision. To find an available path as efficiently as possible, the routing decision adopts a deterministic strategy and selects the action with the highest probability in the PPO model. Then, the information in the packet is updated, such as the number of transmitted hops and the transmitted ratio. When the egress satellite is reached, the DP is copied into a feedback packet (BP) and fed back to the ingress satellite along with the recorded path. Once the BP is received, the ingress satellite forwards it to the source ground terminal.
[0107] When the time slot switches, there is a short switching state between two time slots. In this state, the ongoing routing task may face the problem that the existing path is disconnected and the data packet cannot be correctly and timely sent to the destination. If the method of discarding and resending half of the data packets is adopted, the transmission efficiency will be greatly reduced. The present invention will re-route the affected data packets to ensure that they can continue to move towards the destination. First, the satellite node will determine which data packets are affected by the link switch according to the current node information of the data packet. For these affected data packets, the satellite node will use the trained PPO model to make an online routing decision again with the current position of the data packet as the new starting point and the destination node of the data packet as the end point. As Figure 1, user U1 sends data packets to user U2 through a ground satellite base station and low-earth orbit satellites. The planned path in time slot n1 is that user U1 sends data to ground satellite base station G1, and then successively sends it to ground satellite base station G2 through low-earth orbit satellites S1, S2, S4, and S6. When the time slot is switched, the inter-satellite links between G1 and S1, and S3 and S4 in the original time slot n1 are disconnected, resulting in the inability to transmit data using the existing routing path. At this time, if the low-earth orbit satellite S1 has routing tasks that were not completed in the previous time slot, it needs to make an online routing decision again for re-routing to ensure the integrity of the receiving user's information. Therefore, when such a situation exists, that is, the current satellite node has routing tasks that were not completed in the previous time slot, the current satellite node is used as the source node of the routing, and the routing path between the source node and the destination node is re-planned, and the data completed and cached on the current satellite node, for example Figure 1 If there are still uncompleted routing tasks on the low-earth orbit satellite S1, the routing path is planned with the low-earth orbit satellite S1 as the source node and user U2 as the destination node for data transmission, and the routing path with user U1 as the source node and user U2 as the destination node is re-planned to complete other uncompleted data transmissions. After re-planning, there are two planned paths in time slot n1. The first path is the low-earth orbit satellite S1 - low-earth orbit satellite S3 - low-earth orbit satellite S5 - satellite base station G2 - user U2. This link is due to link switching, resulting in some data being transmitted to intermediate nodes and not reaching the destination node. The path needs to be re-planned to forward this part of the data to the destination node. The second path is user U1 - ground satellite base station G1 - low-earth orbit satellite S2 - low-earth orbit satellite S4 - low-earth orbit satellite S6 - satellite base station G2 - user U2. This path is used for the data sent by user U1 in the current time slot. Due to link switching, the routing path has changed. In the present invention, when a link switch occurs during time slot switching and the switched link affects the data transmission on a certain link, the path is re-planned by the agent. If no link switch occurs during time slot switching or the link switch does not affect the data transmission on the current link, there is no need to re-plan the routing path.
[0108] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A space-ground integrated routing method based on QoS constraints, which uses an agent to select a routing path, is characterized in that When two time slots are switched, the current satellite node determines the data packets that were not completely transmitted in the previous time slot, and determines whether the reason path in the previous time slot is blocked in the current time slot. If it is blocked, the current node is used as the starting point of the routing, and the destination node is used as the end point of the routing. Then, a routing path is re-selected for the data packets that were not completely transmitted at the current node according to the agent.
2. The routing method according to claim 1, characterized in that At the current satellite node, the minimum number of hops from an adjacent satellite to the egress satellite, the time delay required for one hop from the current satellite to the adjacent satellite, the node load level of the adjacent satellite, the link reliability from the current satellite to the adjacent satellite, and the boolean value status indicating whether the adjacent satellite has already been a routing node in another routing path are used as the state vector and input to the agent. The agent outputs the routing path between the current satellite and the destination satellite.
3. The routing method according to claim 1 or 2, characterized in that, When the agent selects an action to execute, the reward value for the action selected by the agent is the utility function of the service on the end-to-end routing path.
4. The routing method according to claim 3, wherein The utility function of the service on the end-to-end routing path is expressed as: Among them, U path is the utility function of a routing path; w1, w2, w3, and w4 are the normalized routing path delays the normalized routing path packet loss rate the normalized routing path throughput the normalized end-to-end energy consumption of the routing path weights.
5. The routing method according to claim 4, characterized in that, The routing path delay includes the delays of all links on the path. The delay of one link includes at least the queuing delay of the data packet at the starting node of the link, the propagation delay between the starting node and the destination node of the link, the forwarding delay of the starting node of the link, and the delay caused by time slot switching.
6. The routing method according to claim 5, wherein At the satellite node, a routing selection matrix is established for the routing path that needs to be forwarded at the current satellite node. Each row of the routing selection matrix represents the occupancy of the current satellite resources by a path in different time slots. The delay caused by time slot switching can be expressed as: Among them, T e represents the time delay of the handover time slot, and T p is the fixed delay setting for the satellite device to process the handover operation; CS[e][n] represents the occupancy of the nth time slot on the e-th link of the current satellite node. If it is occupied, the value of this element is 1, otherwise the value is 0.
7. The routing method according to claim 2, wherein The calculation of the routing path packet loss rate includes: Among them, P path represents the packet loss rate of the routing path path; v i is a satellite node on the routing path path; P i is the packet loss rate of the satellite node v i .
8. The routing method according to claim 2, wherein The calculation of the routing path throughput includes: Among them, D path represents the throughput of a routing path; v i is a satellite node on the routing path; represents the number of data packets successfully transmitted from the satellite node v i to the destination within a given time period; represents the total number of data packets sent by the satellite node v i within a given time period.
9. The routing method according to claim 2, wherein The calculation of the end-to-end energy consumption of the routing path includes: Among them, K path represents the end-to-end energy consumption of a routing path; e i,j represents the link from satellite node i to satellite node j on the routing path path; represents satellite node v i for satellite node v to transmit data to its adjacent satellite node v j the energy consumed per unit of data transmitted on the formed link e i,j ; v i is a satellite node on the routing path path; Ψ i is the energy consumption of the processing output of satellite node v i .