Low earth orbit satellite network fault-tolerant routing method and system based on deep reinforcement learning

By adopting a multi-agent collaboration framework with deep reinforcement learning in low-orbit satellite networks, the problems of high computing complexity and low routing decision efficiency in the existing technology are solved, and efficient and reliable routing decisions and fault recovery are achieved.

CN120075117AActive Publication Date: 2025-05-30XIDIAN UNIV +1

Patent Information

Application Number
CN202510219774.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art faces the problems of high computing complexity, untimely path updates, low routing decision-making efficiency and inability to effectively deal with large-scale network failures in low-orbit satellite networks.

Method used

Using a multi-agent collaborative framework based on deep reinforcement learning, we use detailed modeling of the satellite network topology to construct inter-domain and intra-domain routing frameworks, and use deep Q networks to perform reinforcement learning training in routing decision strategies, and optimize routing decisions to adapt to dynamic network environments.

Benefits of technology

It improves the reliability and efficiency of data transmission, reduces computing and communication overhead, and can adjust routing in real time and intelligently, and quickly adapt to network topology changes and failure recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075117A_ABST
    Figure CN120075117A_ABST
Patent Text Reader

Abstract

The invention provides a low earth orbit satellite network fault-tolerant routing method and system based on deep reinforcement learning, and belongs to the technical field of wireless communication. Specifically, the method comprises the following steps: judging whether a source satellite node and a target satellite node are in the same network domain or not based on domain information of a time slice, a source satellite node address and a target satellite node address; if the source satellite node and the target satellite node are in the same network domain, performing data transmission by adopting an intra-domain routing framework; if not, adopting an inter-domain routing framework to carry out data transmission; generating a routing decision strategy; and carrying out reinforcement learning training on the selected data transmission routing framework by adopting a deep Q network, and optimizing a routing decision strategy by updating an iterative Q value. According to the method, link reliability is modeled in detail, multi-agent collaborative deep reinforcement learning network collaborative intra-domain and inter-domain routing is performed, dynamic changes of a satellite network are responded in real time, data transmission reliability and efficiency are improved, and calculation and communication overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a fault-tolerant routing method and system for a low-earth orbit satellite network based on deep reinforcement learning. Background Art

[0002] Compared with the geosynchronous orbit satellite network, the low-earth orbit satellite network has the advantages of low communication latency, wide coverage, and low transmission power of communication devices. With the growth of global communication services, the low-earth orbit satellite network is increasingly widely used in military, commercial, and civilian applications. Therefore, how to achieve seamless global network coverage and provide users with anytime and anywhere connectivity has become a major challenge faced by the current development of the Internet. Compared with the terrestrial network, the satellite network, with its wide coverage, has gradually become the key solution to the above problems. Among them, LEO satellites, due to their low orbital altitude, exhibit excellent communication capabilities.

[0003] In recent years, the construction of low-earth orbit satellite megaconstellations has attracted extensive attention in the industrial and academic circles [2] 。Multiple enterprises and countries have successively launched the design and deployment of megaconstellations, including typical representatives such as Starlink, OneWeb, Kuiper, and China's "Hongyan". Starlink was proposed by SpaceX and plans to launch 42,000 low-earth orbit satellites. After the first-phase constellation of 1,584 Ku / Ka-band satellites was launched into orbit, the Federal Communications Commission of the United States approved adjusting the orbital altitude to below 570 kilometers. The OneWeb constellation adopts a hybrid polar and inclined orbit configuration and consists of Ku / Ka-band satellites with an orbital altitude of 1,200 kilometers, aiming to achieve broadband network coverage on a global scale. The Kuiper constellation is led by Amazon and plans to deploy 3,236 Ka-band low-earth orbit satellites to provide low-cost broadband services, Internet protocol transmission, and wireless backhaul support using advanced communication technologies. These megaconstellation projects, with their significant advantages of low latency, global coverage, and high bandwidth, are driving a revolutionary change in the global Internet access method, making the LEO satellite megaconstellation an essential part of building the next-generation Internet.

[0004] Due to volume and weight limitations, the computing and communication resources of satellites are relatively limited. In a giant LEO satellite constellation, there are many challenges in achieving large-scale cooperation among heterogeneous satellite nodes, especially in efficiently and reliably transmitting messages and scheduling tasks in a dynamic network environment. Most existing research focuses on small-scale task allocation and resource scheduling within a single domain, and lacks in-depth exploration of cross-domain and cross-node cooperation mechanisms and communication optimization.

[0005] The patent application with the publication number CN117614882A provides a routing decision method and device for a low-earth orbit satellite network based on multi-agent. According to the target satellite address and the pre-established routing strategy, it determines the best transmission path from the current satellite address to the target satellite address by considering the state information of the source satellite node, the destination satellite node, the network state, and adjacent satellites. This patent uses the Markov Decision Process (MDP) to model the routing problem and trains a neural network using a deep reinforcement learning model to optimize the routing decision, effectively handling the challenges brought by the frequent change of satellite positions, but still relying on the pre-established routing strategy. Due to the strong dynamics of satellite positions and network topologies, existing methods face problems such as high computational complexity and slow path updates, especially in large-scale satellite networks, which limits the real-time performance and flexibility.

[0006] The patent application with the publication number CN119300113A addresses the problem of poor reliability in satellite node control in a multi-layer SDN architecture. It provides an inter-satellite reliable routing method based on a fault domain model. According to the real-time state and failure reasons of inter-satellite links, satellite nodes are classified into faulty nodes, boundary nodes, unreachable nodes, unreliable nodes, and available nodes. The scope of the fault area is defined, and appropriate routing paths are selected based on different node states. At the same time, multiple attributes such as the signal-to-noise ratio utility function, link length utility function, and buffer queue utility function are used to calculate the link quality to quantify the reliability, delay, and load capacity of the link, thereby optimizing the path selection and reducing the system packet loss rate and delay. However, in the face of large-scale networks and complex fault situations, the routing selection process may still incur high computational and communication overheads due to excessive link state calculations and information transmissions. This results in low path calculation efficiency and high latency, especially in an environment with dynamic changes and frequent node failures.

[0007] The patent application with the publication number CN117811642A proposes a random link failure routing optimization method in a low-earth orbit satellite network. It uses the pheromone concentration update mechanism of the ant colony algorithm to simulate the path selection process. By dynamically calculating the selection probability of adjacent satellites and based on factors such as pheromone concentration and path desirability values, it optimizes the routing selection to avoid falling into local optimal solutions. At the same time, to ensure communication reliability through alternative paths in case of network node or link failures, this patent calculates the betweenness centrality of nodes and edges in the network, identifies critical nodes or critical paths, and calculates backup paths for critical paths. It fails to fully address the high complexity and high computational overhead in large-scale low-earth orbit satellite networks. As the number of satellites increases, the overheads of information collection, routing decision-making, and fault recovery will increase significantly, and existing methods may not be able to adapt to large-scale and highly dynamic low-earth orbit satellite networks.

[0008] Therefore, in the environment of low Earth orbit satellite networks, especially giant constellations, the present invention provides an efficient, reliable, and fault-tolerant routing decision method, which is an effective solution to some key technical challenges faced by current low Earth orbit satellite networks. Summary of the Invention

[0009] The purpose of the present invention is to overcome the disadvantages of high computational complexity, untimely path update, low routing decision efficiency, and inability to effectively handle large-scale network failures in the prior art, and to provide a fault-tolerant routing method and system for low Earth orbit satellite networks based on deep reinforcement learning.

[0010] To achieve the above purpose, the present invention adopts the following technical solutions: A fault-tolerant routing method for low Earth orbit satellite networks based on deep reinforcement learning, comprising the following steps: Model the links of the satellite network topology structure; Based on the established link model, construct a multi-agent collaborative deep reinforcement learning framework, including an inter-domain routing framework and an intra-domain routing framework; Obtain the data packet to be sent, the domain division information of the time slice, the source satellite node address, and the destination satellite node address; Based on the domain division information of the time slice, the source satellite node address, and the destination satellite node address, determine whether the source satellite node and the destination satellite node are in the same network domain; if the source satellite node and the destination satellite node are in the same network domain, use the intra-domain routing framework for data transmission; if not, use the inter-domain routing framework for data transmission; generate a routing decision strategy; Use a deep Q-network to perform reinforcement learning training on the selected data transmission routing framework, and optimize the routing decision strategy by updating and iterating the Q value; Evaluate the performance of the optimized routing decision strategy in different link availability environments to verify the learning effect.

[0011] In the step of modeling the links of the satellite network topology structure, it includes an inherent link model and an irregular link model. The satellite node and the next-hop satellite node The definition of the inherent link model between them is as follows:

[0012] Among them, represents the inter-satellite link between satellite node and ; represents the reliability of the inter-satellite link; Assume that the current time is , and the size of the historical time window is , the irregular link model is shown as follows:

[0013] Wherein, indicates whether the link operates in each time slot. If so, otherwise ; The overall model of each link is expressed as follows:

[0014] Where , is a parameter for adjusting the weight.

[0015] In the established link model, the estimated remaining propagation delay ERPD is introduced. By sharing the spatial position coordinates between adjacent satellite nodes, the laser propagation delay from the candidate next-hop satellite node to the target satellite node is calculated. The specific formula is as follows:

[0016] Wherein, and respectively represent the spatial position coordinates of the LEO satellite and , and c represents the propagation speed of laser in space.

[0017] In the step of constructing the multi-agent collaborative deep reinforcement learning framework based on the established link model, including the inter-domain routing framework and the intra-domain routing framework, the established intra-domain routing framework is: State space: The state vector of the intra-domain relay node is defined as Where is the set of candidate next-hop nodes for the intra-domain relay node ; Where represents the reliability of the link between the satellite relay node and the next-hop satellite node , represents the from the next-hop satellite node to the target satellite node , represents the occupancy rate of the transmission queue of the link between the relay node and the next-hop satellite node ; Action space: The IDs of the four LEO satellite nodes connected to the relay node through four inter-satellite links; Reward function: Define the time slot The relay node transmits the data packet Forward to the next-hop satellite node The reward function is as follows:

[0018] where, represents the penalty value for the agent in case of packet loss, represents the normalized link reliability, represents the normalized , represents the normalized link transmission queue occupancy rate between the relay node and the next-hop satellite node , is the normalized link reliability, is the normalized reliability, is the weighting coefficient of the normalized link transmission queue occupancy rate; Considering the long-term impact of a single action on subsequent behaviors, the cumulative discounted reward calculation method is as follows:

[0019] where, represents the overall reliability of the normalized link, and the discount factor is used to balance the immediate reward and the future reward.

[0020] The established inter-domain routing framework is as follows: State space: Define the state vector of the domain head satellite node as where is the set of boundary nodes of the domain where the satellite node is located; , where represents the average link reliability of the domain boundary satellite node , represents the from the domain boundary satellite node to the target satellite node , represents the average link transmission queue occupancy rate of the domain boundary satellite node ; Action space: Any node in the boundary node set ; Reward function: Define the reward function for the domain head satellite node to specify the transfer of cross-domain data packets through the domain boundary node in the time slot as follows: ​

[0021] Among them, represents the penalty value for the proxy in case of packet loss, represents the domain boundary node The average link reliability of four inter-satellite links, represents the domain boundary node The distance to the target satellite node of , represents the domain boundary node The average link transmission queue occupancy rate of four inter-satellite links, is the average link reliability, is reliability, is the weighting coefficient of the average link transmission queue occupancy rate.

[0022] Based on the time-slot-based domain information, the source satellite node address, and the destination satellite node address, it is determined whether the source satellite node and the destination satellite node are in the same network domain. If the source satellite node and the destination satellite node are in the same network domain, the intra-domain routing framework is used for data transmission; if not, the inter-domain routing framework is used for data transmission; in the step of generating the routing decision strategy, When using the intra-domain routing framework for data transmission, each satellite node acts as an independent agent and executes routing determination actions based on the state space including link reliability; When using the inter-domain routing framework for data transmission, the inter-domain proxy node selects the optimal boundary node through the following parameters: the average link reliability of the domain where the target boundary node is located, the ERPD of the boundary node to the destination node, and the average transmission queue occupancy rate of the boundary node.

[0023] After the inter-domain proxy selects the optimal boundary node, the current routing is converted to intra-domain routing, and the inter-domain routing framework will call the intra-domain routing framework to complete the intra-domain routing to the boundary node; when the data packet is forwarded to the boundary node, the inter-domain routing framework is called again, and this process is repeated until the data packet reaches the destination node.

[0024] In the step of using the deep Q-network to perform reinforcement learning training on the selected data transmission routing framework and optimizing the routing decision strategy by updating and iterating the Q value, the specific method is as follows: Initialize the evaluation network and the target network, and set the optimizer and the loss function; At the beginning of the training, create an empty experience replay pool. After each execution of the data transmission routing selection action, store the current state, the selected action, the obtained reward, and the next state after executing the action as an experience in the experience replay pool. As the training progresses, the experience data in the experience replay pool gradually accumulates; Based on the current data transmission routing status, the agent selects the action with the maximum Q value as the optimal routing decision using the greedy strategy; Execute the optimal routing decision. The environment updates the node status according to the executed action and gives the corresponding reward value; Combine the state, action, reward, and next state after each action execution as an experience and store it in the experience replay pool; Randomly extract an experience sample from the experience replay pool, calculate the current Q value and the target Q value, use the mean squared error as the loss function, and update the parameters of the evaluation network through the optimizer to optimize the routing decision strategy.

[0025] In the step of evaluating the performance of the optimized routing decision strategy in different link availability environments and verifying the learning effect, the specific method is as follows: Load the optimized routing decision strategy, enter the evaluation mode, and select the action with the maximum Q value as the routing decision; Judge whether the source satellite node address and the destination satellite node address are in the same network domain, and select the corresponding routing framework for routing decision according to the judgment result; Calculate the corresponding reward according to the selected routing decision and output the test result.

[0026] A low-earth orbit satellite network fault-tolerant routing system based on deep reinforcement learning, including: A link modeling module for modeling the links of the satellite network topology structure; A routing framework construction module for constructing a multi-agent collaborative deep reinforcement learning framework based on the established link model, including an inter-domain routing framework and an intra-domain routing framework; An acquisition module for acquiring the data packet to be sent, the domain information of the time slice, the source satellite node address, and the destination satellite node address; A routing decision strategy generation module for judging whether the source satellite node and the destination satellite node are in the same network domain based on the domain information of the time slice, the source satellite node address, and the destination satellite node address; if the source satellite node and the destination satellite node are in the same network domain, use the intra-domain routing framework for data transmission; if not, use the inter-domain routing framework for data transmission; generate a routing decision strategy; A routing decision strategy optimization module for performing reinforcement learning training on the selected data transmission routing framework using a deep Q network and optimizing the routing decision strategy by updating and iterating the Q value; An evaluation and verification module for evaluating the performance of the optimized routing decision strategy in different link availability environments and verifying the learning effect.

[0027] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a fault-tolerant routing method and system for a low-earth orbit satellite network based on deep reinforcement learning, including the following steps: modeling the links of the satellite network topology structure; based on the established link model, constructing a multi-agent collaborative deep reinforcement learning framework, including an inter-domain routing framework and an intra-domain routing framework; obtaining the data packet to be sent, the domain division information of the time slice, the source satellite node address, and the destination satellite node address; based on the domain division information of the time slice, the source satellite node address, and the destination satellite node address, determining whether the source satellite node and the destination satellite node are in the same network domain; if the source satellite node and the destination satellite node are in the same network domain, using the intra-domain routing framework for data transmission; if not, using the inter-domain routing framework for data transmission; generating a routing decision strategy; using a deep Q-network to perform reinforcement learning training on the selected data transmission routing framework, and optimizing the routing decision strategy by updating and iterating the Q value; evaluating the performance of the optimized routing decision strategy in different link availability environments to verify the learning effect. By modeling the links in detail, the multi-agent collaborative deep reinforcement learning network collaborates on intra-domain and inter-domain routing, responds to the dynamic changes of the satellite network in real time, improves the reliability and efficiency of data transmission, and reduces the computing and communication overhead; combining deep reinforcement learning and dynamic routing strategies can adjust the routing in real time and intelligently, ensuring that when the network topology changes or a satellite node fails, the routing strategy can quickly adapt and restore the network connection.

[0028] Furthermore, through intra-domain and inter-domain collaborative routing decisions, significant advantages are shown in the routing decisions of cross-domain traffic. The intra-domain routing framework is responsible for path optimization within the domain, while the inter-domain routing framework is responsible for cross-domain path selection. This design not only improves the efficiency of cross-domain routing decisions but also reduces the overall computational complexity and communication latency. Compared with the single global decision-making mode in the prior art, the multi-agent collaborative framework of the present invention can more effectively handle the collaborative problems of large-scale and heterogeneous satellite nodes.

[0029] Furthermore, the application of the deep Q-network enables the system to gradually optimize the routing decision by learning experience, greatly enhancing the fault tolerance and reliability of the network.

[0030] Furthermore, with the gradual expansion of the scale of the low-earth orbit satellite network, the construction of giant constellations has caused computational complexity problems in traditional routing methods. When faced with a large number of satellites in the prior art, problems such as information collection delay and excessive computational overhead are often encountered. However, the multi-agent collaborative deep reinforcement learning framework adopted by the present invention, through a hierarchical and domain-based structure and local routing decisions, can efficiently handle large constellation networks composed of thousands of satellites, avoiding the inefficiency caused by computational complexity and information collection delay. Description of the Drawings

[0031] Figure 1 This is the flowchart of the present invention; Figure 2 This is the ERPD diagram of the present invention; Figure 3 This is the schematic diagram of the clustering control architecture of the satellite network of the present invention; Figure 4 This is the schematic diagram of the system architecture of the present invention. Detailed implementation manners

[0032] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0033] Embodiment 1 As Figure 1 shown, a low-earth orbit satellite network fault-tolerant routing method based on deep reinforcement learning includes the following steps: S1: Model the links of the satellite network topology structure, including the inherent link model and the irregular link model; S2: Based on the established link model, construct a multi-agent collaborative deep reinforcement learning framework, including an inter-domain routing framework and an intra-domain routing framework; S3: Obtain the data packet to be sent, the domain division information of the current time slice, the source satellite node address, and the destination satellite node address; S4: Based on the domain division information of the current time slice, the source satellite node address, and the destination satellite node address, determine whether the source satellite node and the destination satellite node are in the same network domain; if the source satellite node and the destination satellite node are in the same network domain, use the intra-domain routing framework for data transmission; if not, use the inter-domain routing framework for data transmission; generate a routing decision strategy; S5: Use the deep Q-network to perform reinforcement learning training on the selected data transmission routing framework, and optimize the routing decision strategy by updating and iterating the Q value; S6: Evaluate the performance of the optimized routing decision strategy in different link availability environments to verify the learning effect.

[0034] Specifically, in S1, due to the flexibility of the satellite network topology structure, the link characteristics change frequently. As the basic unit for realizing data transmission, there are obvious differences in the characteristics such as delay, packet loss, and bandwidth of each path. Therefore, when designing and optimizing the reliable routing scheme of the satellite network, it is necessary to accurately model the link attributes.

[0035] The present invention models link reliability as two parts: inherent link reliability and irregular link reliability. Due to the inter-satellite distance and phase change, the stability of the inter-orbit ISL is poor, while the state of the intra-orbit ISL is relatively stable. The satellite nodes and the next-hop satellite node The definition of the inherent link reliability model between them is as follows:

[0036] Wherein, represents the inter-satellite link between satellite node and ; represents the inter-satellite link reliability.

[0037] In addition, potential signal interference from antenna pointing can cause irregular failures of the ISL. The historical moment failures are used to determine the irregular link reliability. Assume the current moment is , and the historical moment window size is , then the irregular link reliability is defined as follows:

[0038] Where represents whether the link is operating in each time slot. If so, , otherwise .

[0039] The overall reliability model of each link can be expressed as the following formula:

[0040] Where is a parameter for adjusting the weight.

[0041] Furthermore, the present invention introduces the Estimated Residual Propagation Delay (ERPD) for forwarding satellite-driven routing selection inspired by the position-assisted routing protocol. The ERPD is defined as the laser propagation delay for transmitting a data packet from the next-hop satellite node to the target satellite node . To calculate the ERPD of the intermediate forwarding satellite , its adjacent satellite set needs to share the spatial position coordinate information with . Then for any , the ERPD from to is as follows:

[0042] Wherein, and respectively represent the spatial position coordinates of LEO satellites and , and c represents the propagation speed of laser in space. The above formula illustrates how to calculate the ERPD of the intermediate relay satellite . is the source node, is the destination node, is the set of candidate next-hop nodes. If is selected as 's next hop, then 's ERPD is . This parameter is passed in as a state for subsequent deep reinforcement learning and is used to measure the distance between the next-hop node and the destination node. The larger the proportion of this parameter, the more inclined the agent is to select a smaller ERPD value. Taking Figure 2 as an example, then is selected as the next-hop node.

[0043] Due to the different user densities in the satellite coverage area and the increasing data transmission requirements, as the satellite moves, the traffic it receives from the ground gateway station or user nodes also changes continuously, resulting in a high growth rate and uneven distribution of the LEO satellite network load. Therefore, the present invention also considers the transmission queue occupancy rate of the link between satellite and satellite , as shown in the following formula:

[0044] where, represents the number of data packets in the current link transmission queue, represents the maximum data packet limit of the current link transmission queue.

[0045] Specifically, in S2, as shown in Figure 3 , the hierarchical and domain-based software-defined satellite network management and control architecture adopted by the present invention is such that the slave controller is responsible for processing the collection of frequent network and traffic status information within the domain and the design of in-domain routing policies, shielding the messages of this domain from the master controller, thereby alleviating the pressure and overhead of the master controller, effectively coping with the robustness of routing in the scenario of controller disconnection, and improving network performance.

[0046] A multi-agent collaborative deep reinforcement learning framework is designed. The framework includes two layers of deep reinforcement learning modules - the inter-domain routing framework and the intra-domain routing framework, which are used for routing decisions of cross-domain traffic and intra-domain traffic respectively. In the intra-domain routing framework, each LEO satellite is regarded as an intra-domain deep reinforcement learning (DRL) agent, that is, an intra agent. Each intra agent can only perceive the position status, queue status, link status, etc. of its single-hop adjacent nodes. In the inter-domain routing framework, each domain head satellite, that is, the domain controller, is regarded as an inter-domain deep reinforcement learning agent, that is, an inter agent. Each inter agent can only perceive the position status, queue status, link status, etc. of all nodes within its domain. Therefore, whether it is an intra agent or an inter agent, the dynamic routing decision problem of each agent can be modeled as a partially observable Markov decision process (POMDP).

[0047] 1) Intra-domain routing framework State space: The state vector of the intra-domain agent, that is, the relay node is defined as , where is the set of candidate next-hop nodes for the intra-domain relay node ; , where represents the reliability of the link between the satellite relay node and the next-hop satellite node , represents the from the next-hop satellite node to the target satellite node , represents the occupancy rate of the transmission queue of the link between the relay node and the next-hop satellite node ; Action space: The IDs of the four LEO satellite nodes connected to the current relay node through four inter-satellite links.

[0048] Reward function: To ensure that each intra-domain agent, that is, each satellite, learns the optimal routing decision, define the time slot the relay node for forwarding the data packet to the next-hop satellite node The reward function is as follows:[[]]

[0049] Among them, represents the penalty value for the proxy in case of packet loss, represents the normalized link reliability, represents normalization , represents the transit node and the next-hop satellite node the normalized link transmission queue occupancy rate between them, is the normalized link reliability, is normalization reliability, is the weighting coefficient of the normalized link transmission queue occupancy rate; Considering the long-term impact of a single action on subsequent behaviors, the cumulative discounted reward calculation method is as follows:

[0050] Among them, represents the overall reliability of the normalized link, and the discount factor is used to balance the immediate reward and the future reward.

[0051] 2) Inter-domain routing framework State space: The inter-domain proxy, that is, the state vector of the domain head satellite node is defined as , where is the set of boundary nodes of the domain where the satellite node is located. If the four LEO satellite nodes adjacent to a node are not in the same domain as this node, then this node is a boundary node. , where represents the average link reliability of the domain boundary satellite node , represents the from the domain boundary satellite node to the target satellite node , represents the average link transmission queue occupancy rate of the domain boundary satellite node .

[0052] Action space: Any node in the boundary node set ; Reward function: To ensure that each inter-domain proxy, that is, each domain head satellite learns the optimal routing decision, define the time slot the domain head satellite node , specifies the reward function for the transit cross-domain data packet through the domain boundary node as follows:

[0053] Among them, represents the penalty value for the proxy in the case of packet loss, represents the domain boundary node the average link reliability of four inter-satellite links, represents the domain boundary node the distance to the target satellite node of , represents the domain boundary node the average link transmission queue occupancy rate of four inter-satellite links, is the average link reliability, is reliability, is the weighted coefficient of the average link transmission queue occupancy rate.

[0054] Considering the long-term impact of a single action on subsequent behaviors, the cumulative discounted reward calculation method is shown as follows:

[0055] Among them, represents the overall reliability of the normalized link, and the discount factor is used to balance the immediate reward and the future reward.

[0056] Specifically, in S4, if the source satellite node and the destination satellite node are in the same network domain, the intra-domain routing framework is adopted; in the intra-domain routing framework, each satellite node acts as an independent agent to make decisions; the proxy makes action decisions based on the state space information, including link reliability, ERPD, and transmission queue occupancy rate. If the source satellite node and the destination satellite node are not in the same network domain, the inter-domain routing framework is adopted; in the inter-domain routing framework, the inter-domain proxy node will select a boundary node as the target for data transmission. To select the best boundary node, the inter-domain proxy needs to make decisions according to the current network state, including the average link reliability of the boundary node, the ERPD of the boundary node's distance to the target node, and the average transmission queue occupancy rate of the boundary node.

[0057] After the inter-domain proxy selects the optimal boundary node, the current routing is converted to intra-domain routing, and the inter-domain routing framework will call the intra-domain routing framework to complete the intra-domain routing to the boundary node; when the data packet is forwarded to the boundary node, the inter-domain routing framework is called again, and this process is repeated until the data packet reaches the destination node.

[0058] Specifically, in S5, a deep Q-network is used to perform reinforcement learning training on the selected data transmission routing framework, and the routing decision-making strategy is optimized by updating and iterating the Q value. The specific training process is as follows: S51: Parameter initialization: Initialize the evaluation network and the target network, and set the optimizer and the loss function; S52: Experience replay pool: At the beginning of training, create an empty experience replay pool. After each execution of the data transmission routing selection action, store the current state, the selected action, the obtained reward, and the next state after the execution of the action as an experience in the experience replay pool. As training progresses, the experience data in the experience replay pool gradually accumulates; S53: Select action: The agent selects an action based on the current state using the greedy strategy, that is, selects an action most relevant to the current state (selects the action with the largest Q value). At the same time, to maintain exploration, the agent also randomly selects an action with a certain probability for trial; S54: Execute action: Execute the optimal routing decision. The environment updates the node state according to the executed action and gives the corresponding reward value; S55: Store experience: Store the state, action, reward, and next state combination after each execution of the action as an experience in the experience replay pool; S56: Update network: When there is a sufficient amount of experience data accumulated in the experience replay pool, randomly extract a mini-batch of experience samples from the experience replay pool, calculate the current Q value and the target Q value, use the mean squared error as the loss function, and update the parameters of the evaluation network through the optimizer to optimize the routing decision-making strategy; Every certain number of training steps, copy the parameters of the evaluation network to the target network to maintain the stability of the target network and improve the learning efficiency.

[0059] Specifically, in S6, evaluate the performance of the optimized routing decision-making strategy in different link availability environments and verify the learning effect. The specific method is as follows: S61: Evaluate the model: Load the optimized routing decision-making strategy, enter the evaluation mode, and select the action with the largest Q value as the routing decision; S62: Determine whether the source satellite node address and the destination satellite node address are in the same network domain, and select the corresponding routing framework for routing decision according to the judgment result; S63: Calculate the reward: Test under the combined scenarios of different link failure rates and node load rates, and verify the adaptability and optimization effect of the routing strategy in a complex network environment according to test metrics such as the actual packet loss rate, end-to-end delay, and throughput.

[0060] Since this routing scheme is a fault-tolerant routing, the test results include packet loss rates, end-to-end delays, delay jitters, and throughput under scenarios with different load rates (10% - 100%) and different link failure probabilities (0% - 40%). The advantages of this scheme are measured by comparing these metrics such as packet loss rates with typical routing schemes, such as OSPF, ECMP, or other satellite network fault-tolerant routing schemes. For a single intra-domain routing agent / inter-domain routing agent, the larger the reward value, the better the routing strategy is considered.

[0061] Embodiment 2 A low-earth orbit satellite network fault-tolerant routing system based on deep reinforcement learning, comprising: A link modeling module for modeling the links of the satellite network topology structure; A routing framework construction module for constructing a multi-agent collaborative deep reinforcement learning framework based on the established link model, including an inter-domain routing framework and an intra-domain routing framework; An acquisition module for acquiring the data packet to be sent, the domain information of the time slice, the source satellite node address, and the destination satellite node address; A routing decision strategy generation module for determining whether the source satellite node and the destination satellite node are in the same network domain based on the domain information of the time slice, the source satellite node address, and the destination satellite node address; if the source satellite node and the destination satellite node are in the same network domain, use the intra-domain routing framework for data transmission; if not, use the inter-domain routing framework for data transmission; generate a routing decision strategy; A routing decision strategy optimization module for performing reinforcement learning training on the selected data transmission routing framework using a deep Q network, and optimizing the routing decision strategy by updating and iterating the Q value; An evaluation and verification module for evaluating the performance of the optimized routing decision strategy in different link availability environments and verifying the learning effect.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A fault-tolerant routing method for low-orbit satellite networks based on deep reinforcement learning, characterized in that: The steps include: Modeling satellite network topology links; Based on the established link model, a multi-agent collaborative deep reinforcement learning framework is constructed, including an inter-domain routing framework and an intra-domain routing framework; Obtain the data packet to be sent, the domain information of the time slice, the source satellite node address and the destination satellite node address; Based on the domain information of the time slice, the source satellite node address and the destination satellite node address, determine whether the source satellite node and the destination satellite node are in the same network domain; If the source satellite node and the destination satellite node are in the same network domain, the intra-domain routing framework is used for data transmission; If they are not in the same network domain, an inter-domain routing framework is used for data transmission; Generate routing decision strategies; A deep Q network is used to perform reinforcement learning training on the selected data transmission routing framework, and the routing decision strategy is optimized by updating the iterative Q value; The performance of the optimized routing decision strategy in different link availability environments is evaluated to verify the learning effect.

2. According to claim 1, a fault-tolerant routing method for low-orbit satellite networks based on deep reinforcement learning is characterized in that: The step of modeling the satellite network topology structure link includes an inherent link model and an irregular link model. and the next-hop satellite node The inherent link model between is defined as follows: in, Satellite node and Intersatellite links, Indicates intersatellite link reliability; Assume that the current time is , the historical moment window size is , then the irregular link model is as follows: in, Indicates whether the link is running in each time slot. If so, ,otherwise ; The overall model of each link is expressed as follows: in , is the parameter for adjusting the weight.

3. According to claim 2, a low-orbit satellite network fault-tolerant routing method based on deep reinforcement learning is characterized in that: The estimated residual propagation delay ERPD is introduced into the established link model. The laser propagation delay from the candidate next-hop satellite node to the target satellite node is calculated by sharing the spatial position coordinates between adjacent satellite nodes. The specific formula is as follows: in, and Represents LEO satellites and The spatial position coordinates of , c represents the propagation speed of the laser in space.

4. According to claim 3, a fault-tolerant routing method for low-orbit satellite networks based on deep reinforcement learning is characterized in that: In the step of constructing a multi-agent collaborative deep reinforcement learning framework based on the established link model, including an inter-domain routing framework and an intra-domain routing framework, the established intra-domain routing framework is: State space: transit nodes within the domain The state vector is defined as ,in It is a transit node within the domain The set of candidate next-hop nodes; ,in Indicates a satellite transfer node Next-hop satellite node The reliability of the link between Indicates the next hop satellite node To the target satellite node of , Represents a transit node Next-hop satellite node The sending queue occupancy rate of the link between them; Action Space: With Transfer Nodes Four LEO satellite node IDs connected by four intersatellite links; Reward Function: Defining Time Slots Transfer Node Packet Forward to the next hop satellite node The reward function is as follows: in, Indicates the penalty value for the proxy in case of packet loss, represents the normalized link reliability, Represents normalization , Represents a transit node Next-hop satellite node Normalized link send queue occupancy between, To normalize the link reliability, Normalized reliability, is the weighted coefficient of the normalized link send queue occupancy; Considering the long-term impact of a single action on subsequent behaviors, the cumulative discounted reward calculation method is as follows: in, represents the overall reliability of the normalized link, the discount factor Used to strike a balance between immediate rewards and future rewards.

5. According to claim 4, a fault-tolerant routing method for low-orbit satellite networks based on deep reinforcement learning is characterized in that: The inter-domain routing framework established is: State space: The domain first satellite node The state vector is defined as ,in For satellite nodes The set of boundary nodes of the domain; ,in Satellite nodes representing domain boundaries The average link reliability is Satellite nodes representing domain boundaries To the target satellite node of , Satellite nodes representing domain boundaries Average link send queue occupancy rate; Action space: a set of boundary nodes Any node in ; Reward Function: Defining Time Slots Domain Head Satellite Node Specify the node that passes through the domain boundary Relay cross-domain data packets The reward function is as follows: in, Indicates the penalty value for the proxy in case of packet loss, Domain boundary node The average link reliability of the four intersatellite links, Domain boundary node Distance to target satellite node of , Domain boundary node The average link transmission queue occupancy rate of the four intersatellite links, is the average link reliability, for reliability, is the weighting coefficient of the average link send queue occupancy.

6. According to claim 1, a fault-tolerant routing method for low-orbit satellite networks based on deep reinforcement learning is characterized in that: The domain information based on the time slice, the source satellite node address and the destination satellite node address is used to determine whether the source satellite node and the destination satellite node are in the same network domain. If the source satellite node and the destination satellite node are in the same network domain, an intra-domain routing framework is used for data transmission; If they are not in the same network domain, an inter-domain routing framework is used for data transmission; In the step of generating a routing decision strategy, When using the intra-domain routing framework for data transmission, each satellite node acts as an independent agent and performs routing decision actions based on the state space including link reliability; When using the inter-domain routing framework for data transmission, the inter-domain proxy node selects the optimal border node based on the following parameters: average link reliability of the domain where the target border node is located, ERPD from the border node to the destination node, and average send queue occupancy of the border node.

7. The method for fault-tolerant routing of low-orbit satellite networks based on deep reinforcement learning according to claim 6, characterized in that: After the inter-domain proxy selects the optimal border node, the current route is converted to intra-domain route, and the inter-domain routing framework will call the intra-domain routing framework to complete the intra-domain routing to the border node; When the data packet is forwarded to the border node, the inter-domain routing framework continues to be called and the process is repeated until the data packet reaches the destination node.

8. The method for fault-tolerant routing of low-orbit satellite networks based on deep reinforcement learning according to claim 1, characterized in that: In the step of using a deep Q network to perform reinforcement learning training on the selected data transmission routing framework and optimizing the routing decision strategy by updating the iterative Q value, the specific method is as follows: Initialize the evaluation network and target network, and set the optimizer and loss function; At the beginning of training, an empty experience replay pool is created. After each data transmission route selection action is executed, the current state, the selected action, the reward obtained, and the next state after the action is executed are stored in the experience replay pool as an experience. As the training progresses, the experience data in the experience replay pool is gradually accumulated; Based on the current data transmission routing status, the agent uses a greedy strategy to select the action with the maximum Q value as the optimal routing decision; Execute the optimal routing decision, the environment updates the node status according to the executed action and gives the corresponding reward value; The combination of state, action, reward and next state after each action is executed is stored in the experience replay pool as an experience; An experience sample is randomly drawn from the experience replay pool, the current Q value and the target Q value are calculated, the mean square error is used as the loss function, the parameters of the evaluation network are updated through the optimizer, and the routing decision strategy is optimized.

9. The method for fault-tolerant routing of low-orbit satellite networks based on deep reinforcement learning according to claim 8, characterized in that: In the step of evaluating the performance of the optimized routing decision strategy in different link availability environments and verifying the learning effect, the specific method is as follows: Load the optimized routing decision strategy, enter the evaluation mode, and select the action with the largest Q value as the routing decision; Determine whether the source satellite node address and the destination satellite node address are in the same network domain, and select the corresponding routing framework to make routing decisions based on the judgment result; Calculate the corresponding reward based on the selected routing decision and output the test results.

10. A low-orbit satellite network fault-tolerant routing system based on deep reinforcement learning, characterized in that: include: A link modeling module, used to model the satellite network topology links; The routing framework building module is used to build a multi-agent collaborative deep reinforcement learning framework based on the established link model, including an inter-domain routing framework and an intra-domain routing framework; An acquisition module is used to acquire the data packet to be sent, the domain information of the time slice, the source satellite node address and the destination satellite node address; A routing decision strategy generation module is used to determine whether the source satellite node and the destination satellite node are in the same network domain based on the domain information of the time slice, the source satellite node address and the destination satellite node address; If the source satellite node and the destination satellite node are in the same network domain, the intra-domain routing framework is used for data transmission; If they are not in the same network domain, an inter-domain routing framework is used for data transmission; Generate routing decision strategies; The routing decision strategy optimization module is used to use a deep Q network to perform reinforcement learning training on the selected data transmission routing framework, and optimize the routing decision strategy by updating the iterative Q value; The evaluation and verification module is used to evaluate the performance of the optimized routing decision strategy in different link availability environments and verify the learning effect.

Citation Information

Patent Citations

  • Multi-agent-based low-orbit satellite network routing decision-making method and device

    CN117614882A

  • Random link fault routing optimization method in low earth orbit satellite network

    CN117811642A

  • Inter-satellite reliable routing method based on fault domain model

    CN119300113A

  • Software-defined satellite-ground convergence network QoE perception routing architecture based on deep reinforcement learning

    CN114173392A

  • Method for optimizing routing decision and wavelength allocation on sub-domain satellite

    CN116232425A

Cited By

  • Large-scale low-orbit satellite network domain division intelligent routing method

    CN120567276A

  • A Domain-Specific Intelligent Routing Method for Large-Scale Low-Earth Orbit Satellite Networks

    CN120567276B

  • Remote sensing task on-orbit processing and intelligent scheduling method and system based on satellite-ground cooperation

    CN121436597A

  • On-orbit processing and intelligent scheduling method and system for remote sensing missions based on satellite-ground collaboration

    CN121436597B

  • Service quality sensing routing method and system for large-scale low earth orbit satellites

    CN121567196A