A Trusted Load Balancing Routing Method, System, Device and Medium for Low Earth Orbit Satellite Networks
Through the multi-agent D3QN deep reinforcement learning method, combined with trust evaluation and dynamic routing decisions, the problems of trust and load balancing in low-orbit satellite networks are solved, and efficient routing security and load balancing are achieved, with strong adaptability and can meet the trust and delay requirements of services.
Patent Information
- Application Number
- CN202310379081.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-11
AI Technical Summary
The existing low-orbit satellite network routing algorithms fail to effectively combine trusted routing and load balancing, making it difficult to meet the trust requirements and delay-sensitive requirements of services in scenarios where user distribution is uneven and network topology changes rapidly, and the impact of existing solutions on malicious nodes is inevitable.
The multi-agent D3QN deep reinforcement learning method is adopted, and through the trust evaluation model and dynamic routing decision-making, combining the trust value, queue utilization and delay of inter-star links, a fully distributed multi-agent architecture is designed to optimize the routing path to meet the trust and delay requirements.
It realizes efficient trusted load balancing routing in low-orbit satellite networks, reduces the maximum queue utilization rate of the network, improves routing security and service accessibility, is highly adaptable, and can quickly respond to network state changes.
Smart Images

Figure CN116390164B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite communication, and particularly relates to a trusted load balancing routing method, system, device and medium for a low-earth orbit satellite network. Background Art
[0002] The emerging 6G Space-Air-Ground Integrated Network (SAGIN) can support high-speed wide-coverage connections and give rise to a large number of new applications. As an important part of the 6G SAGIN, the Low Earth Orbit Satellite Network (LEO-SN) can cooperate with the ground network to achieve high-quality communication for various services. In the LEO-SN, untrusted nodes may misbehave by discarding data packets or launching other malicious attacks, resulting in routing misbehavior. Trusted routing schemes use trust evaluation mechanisms to design routing paths based on the trust values of nodes, avoiding the degradation of Quality of Service (QoS) caused by malicious nodes. In addition, load balancing routing is also an important way to avoid network congestion and thus improve the QoS of satellite applications. Since satellite networks are completely different from ground networks, many new issues need to be considered when designing the routing scheme for LEO-SN. First, the topological structure of LEO-SN changes rapidly with the high-speed movement of satellites, and node behaviors are highly dynamic and uncertain. Second, due to the uneven distribution of ground users, the traffic distribution in the satellite network is unbalanced. Finally, the length of inter-satellite links often changes at different latitudes, resulting in uncertainties in communication delay and trust evaluation. Therefore, it is necessary to design a trusted load balancing routing scheme in the scenario of a low-earth orbit satellite network to solve the problems of trusted routing and load balancing routing in the low-earth orbit satellite network and improve the network QoS performance.
[0003] In the existing related research, there is no research on combining trusted routing with load balancing routing and optimizing the solution through deep reinforcement learning training. The reasons are as follows: 1. In the existing research on low-earth orbit satellite network routing algorithms, the research on the security of inter-satellite routing is relatively less. The few trusted routing schemes for low-earth orbit satellite networks only make certain optimizations for the security of inter-satellite routing and do not consider other related communication performances of satellite routing, such as delay, packet loss rate, queue utilization rate, etc., and it is difficult to meet the current demand situation of satellite communication; 2. Most of the existing low-earth orbit satellite routing algorithms are based on virtual topologies. The virtual topology divides the satellite topology into static topologies within several time slices according to the periodicity of satellite movement, thus shielding the dynamicity of the satellite topology. However, the low-earth orbit satellite topology changes frequently. A larger segmentation interval is difficult to capture the topology changes, and a smaller segmentation interval will generate a large amount of redundant topology data. How to enable the routing algorithm to be implemented in the frequently changing satellite topology is the premise for improving satellite QoS.
[0004] The patent application No.
CN202210143466.1
[0005] The patent application No.
CN202210127411.1
[0006] The above two solutions both have the following problems:
[0007] They are less efficient in solving the load balancing problem of low-earth orbit satellite networks. In scenarios where user distribution is uneven and network topologies change rapidly, the load balancing performance of the above solutions is poor. They only consider the routing optimization of the same type of service and rarely consider the delay constraints of multiple delay-sensitive services. Summary of the Invention
[0008] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a method, system, device and medium for reliable load balancing routing in a low-earth orbit satellite network. A fully distributed reliable load balancing routing scheme for low-earth orbit satellites based on multi-agent D3QN (Dueling Double DeepQ Network) learning is proposed. Multiple agents are organized to make decisions based on the trust value, delay, and queue utilization of nodes, and generate routes hop by hop. It has good dynamic adaptability and scalability, can be deployed on various satellite constellations, can meet the trust requirements and delay-sensitive requirements of services. This scheme solves the problems of reliable routing and load balancing routing in the low-earth orbit satellite scenario, and can reduce the maximum queue utilization of the network while meeting the trust requirements and delay-sensitive requirements of services, and balance the network load.
[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0010] A method for reliable load balancing routing in a low-earth orbit satellite network includes the following steps:
[0011] Step 1: Construct a communication scenario for the low-earth orbit satellite network; establish a topology model of the low-earth orbit satellite network according to the low-earth orbit satellite network, construct network parameters of the topology model according to the network environment, and evaluate the trust degree of each satellite using the decentralized DTMS trust evaluation model;
[0012] Step 2: Model the quality of service constraints for reliable load balancing routing of low-earth orbit satellites as constraints during model solving;
[0013] Step 3: Model the objective function for reliable load balancing routing of low-earth orbit satellites and construct the objective function during model solving;
[0014] Step 4: Establish a fully distributed multi-agent architecture in the topology model of the low-earth orbit satellite network established in Step 1;
[0015] Step 5: Establish a D3QN deep reinforcement learning training model in the fully distributed multi-agent architecture established in Step 4. Using the constraint conditions obtained in Step 2 and the objective function obtained in Step 3, set the state, action, and reward of the training model, and train the routing decision model for the multi-agent D3QN learning model;
[0016] Step 6: Design a reliable load balancing routing strategy for the low-earth orbit satellite network according to the routing decision model trained in Step 5.
[0017] The specific method for constructing the communication scenario of the low-earth orbit satellite network in Step 1 is as follows:
[0018] Construct the LEO-SN topology as an undirected graph G=(V,E), where V={v1,v2,…,v n} represents the set of satellite nodes, where v i is the i-th satellite node, and n is the total number of satellite nodes. represents the set of inter-satellite links. represents the inter-satellite link i between satellite v j and satellite v i,j Define τ i,j and ω to describe the transmission delay and bandwidth of the inter-satellite link i respectively. Define ξ i as the queue utilization rate on satellite node v
[0019]
[0020] where is the number of current data packets in v i , is the total queue capacity in v i . ξ i reflects the load condition of the satellite node, that is, the higher ξ i , the heavier the load and the more prone to congestion.
[0021] After selecting the routing path, construct a one-dimensional array P = [v s , …, v m-1 , v m , v m+1 , …, v d to represent the satellite elements of the routing path, where v s represents the source node, v m represents the intermediate node, and v d represents the destination node. Represent the links in path P as L P = [e1, e2, …, e k , …, e K , where e k represents the k-th link that makes up the path. The transmission delay on the routing path can be calculated as:
[0022]
[0023] where τ k represents the transmission delay of link e k . Use the bottleneck link bandwidth to represent the bandwidth of the path. The formula is:
[0024]
[0025] where ω k represents the bandwidth of link e kThe bandwidth, the maximum queue utilization rate and the average queue utilization rate of path P can be respectively:
[0026]
[0027]
[0028] Among them, l represents the number of nodes in path P;
[0029] Use the decentralized DTMS trust evaluation model to evaluate the trust values of each satellite based on the historical behavior of the satellites, and obtain the trust value Γ of each satellite in each state i .
[0030] The modeling of the quality of service constraints for the trusted load balancing routing of low-earth orbit satellites in step 2 is specifically as follows:
[0031] For each node v in the routing path i , its trust value should satisfy:
[0032]
[0033] Among them, is the trust threshold of the mth service;
[0034] Use the penalty strategy based on delay constraints to improve the quality of each service. If the delay of the mth flow is greater than its QoS threshold then this routing action will be severely punished, where the reward function leads to negative feedback, and the QoS threshold is not a fixed constant. It varies according to ensuring different delay requirements. The constraint of variable delay is:
[0035]
[0036] The modeling of the trusted load balancing routing objective function for low-earth orbit satellites in step 3, constructing the objective function when solving the model; specifically as follows:
[0037] The objective function can be expressed as:
[0038]
[0039] C1: λ1 + λ2 = 1, λ1, λ2 ≥ 0
[0040] C2:
[0041] C3: Among them, λ1 and λ2 are the weights of the objective function, C1 is the weight constraint, C2 is the delay constraint, and C3 is the trust constraint.
[0042] The specific method for establishing the fully distributed multi-agent architecture in step 4 is as follows:
[0043] Design a fully distributed multi-agent architecture, model the dynamic routing decision problem of each satellite as a partially observable Markov decision process (POMDP), regard each satellite as an agent, without parameter sharing among agents, train independently, and construct the routing decision of the entire network as a multi-agent POMDP.
[0044] The specific method for step 5 is as follows:
[0045] Combine the multi-agent POMDP model with the D3QN algorithm to design the trustworthy load-balanced routing (TLBR) for LEO-SN. The agent obtains the current state through interaction with surrounding nodes, and derives the reward value through the state transition function and the reward function after making a decision.
[0046] 5.1: State space setting;
[0047] The state of satellite node v i at time t is:
[0048]
[0049] where d is a binary vector of dimension 1×N. If v i is the destination node, the i-th element in d is set to 1, and other elements are 0. The shortest hop count from the current node v i to the destination node is K i , and the shortest hop count from neighbor node j to the destination node is K j , ξ j is the queue utilization rate of neighbor node j, indicating the load situation of node j, and τ ij and ω ij respectively describe the link delay and bandwidth from v i to the j-th neighbor node, and Γ j is the trust value of the j-th neighbor node;
[0050] 5.2: Action space setting;
[0051] The action space of the current satellite node v i is represented as a i ={j}. When a i =j, it means that the agent selects the j-th neighbor node as the next-hop node;
[0052] 5.3: Reward function setting;
[0053] The specific reward design is discussed as follows:
[0054] (1) Trust reward: The reward related to trust is designed as:
[0055]
[0056] Among them, is the proportional adjustment parameter, and θ1 is a negative value, which is used to punish the decision that does not meet the trust requirement;
[0057] (2) QoS reward: The reward related to the delay service quality is:
[0058]
[0059] Among them, is the proportional adjustment parameter, θ2 is a negative value, which is used to punish the decision that does not meet the Qos requirement, and l s is the hop count of the shortest path of the current service;
[0060] (3) Load reward: The reward related to the load is:
[0061]
[0062] Among them, is the proportional adjustment parameter;
[0063] (4) Path reward: The reward related to the routing is:
[0064]
[0065] Among them, is the scale adjustment parameter. Using K i and K j can accelerate the convergence of the routing scheme; when K i -K j > 0, it means that the current decision result is closer to the destination node, and a positive reward is given to this decision. Otherwise, when K i -K j < 0, it means that the current decision result is far from the destination node, and a negative reward is given;
[0066] Therefore, the overall reward can be described as:
[0067] r i = ρ1r1 + ρ2r2 + ρ3r3 + ρ4r4
[0068] Among them, ρ1, ρ2, ρ3, and ρ4 are positive weights, where ρ1 + ρ2 + ρ3 + ρ4 = 1.
[0069] The method for designing the trusted load balancing routing strategy of the low-earth orbit satellite network in step 6 is specifically as follows:
[0070] The trusted load balancing routing algorithm is divided into a training stage and a decision-making stage;
[0071] 6.1: In the training stage, at the beginning of the training, the experience replay pool and network parameters are initialized. Source node and destination node pairs are randomly generated in the LEO-SN topology. The routing decision is made by the source node. During the decision-making process, the agent makes decisions according to the following dynamic greedy strategy:
[0072]
[0073] where step is the number of training steps, is a constant that can adjust the decreasing speed of the decision probability in the dynamic greedy strategy. σ is a random number in the range of [0, 1]. After each decision, the reward value is calculated according to the reward function. After an agent makes a decision, the current sample is stored in the experience replay pool for future use. The estimation network updates the network weights according to the gradient descent method. The target network updates the network weights from the estimation network every C steps. When the destination node is found, one routing training ends. Repeat randomly generating source and destination node pairs for training until the model training converges;
[0074] 6.2: In the decision-making stage, according to the decision-making model that converges during training: a. Initialize the routing path; b. Input the source node and destination node of the current service; c. The agent of the current node makes the next-hop decision according to the surrounding environmental state; d. Add the decision node to the routing path; e. Take the next-hop node as the new current node; f. Judge whether the current node is the destination node. If it is the destination node, end the routing to obtain the routing path. If it is not the destination node, jump to step c to continue the decision-making.
[0075] The present invention also provides a system for implementing the trusted load balancing routing method for the low-earth orbit satellite network, including:
[0076] Low-earth orbit satellite network topology construction module: used to construct the low-earth orbit satellite network topology model, and construct the low-earth orbit satellite network into a topology graph with topological relationships and network parameters;
[0077] Multi-agent D3QN training module: used to implement the training of the multi-agent D3QN decision-making model. According to the set constraints and objective functions, use the data set to train the multi-agent D3QN decision-making model so that it can decide the best routing path according to the state environment of the node;
[0078] Trusted load balancing routing policy decision-making module: used to implement the trusted load balancing routing policy, and plan the routing path for services with different types of service requirements according to the trained multi-agent D3QN decision-making model.
[0079] The present invention also provides a trusted load balancing routing device for a low-earth orbit satellite network, including:
[0080] A memory for storing computer programs;
[0081] A processor for implementing the described trusted load balancing routing method for a low-earth orbit satellite network when executing the computer programs.
[0082] The present invention also provides a computer-readable storage medium storing a computer program, which can implement a trusted load balancing routing method for a low-earth orbit satellite network when executed by a processor.
[0083] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0084] 1. Currently, there is no relevant research on combining trusted routing and load balancing routing in a low-earth orbit satellite network and optimizing the solution through deep reinforcement learning training. The present invention combines trusted routing and load balancing routing in a low-earth orbit satellite network, considers the load situation of the satellite network while taking into account the security of satellite routing, and improves the QoS of satellite services by designing an optimal objective function and constraint conditions.
[0085] 2. From the perspective of on-board service processing of low-earth orbit satellites, the present invention sets different credibility thresholds and delay-sensitive constraints according to different on-board services, and the routing decision changes dynamically with different services, so that each routing path can meet the delay requirements and security requirements of the services, improving the reach rate of satellite services.
[0086] 3. From the perspective of low-earth orbit satellite network routing technology, the present invention adopts a fully distributed multi-agent architecture, takes each satellite as an agent, and makes routing decisions independently according to the surrounding environmental states, which has high adaptability to the highly dynamic satellite network topology; at the same time, the agent decision-making mechanism can also quickly respond to the state changes in the network environment, can timely detect the load changes in the network, and optimize the network load balancing.
[0087] 4. The multi-agent D3QN training model provided in step 5 of the present invention can train routing decisions according to states, actions, and rewards, and has the advantages of fast convergence speed and good decision-making performance.
[0088] 5. The trusted load balancing routing strategy provided in step 6 of the present invention can quickly make routing decisions according to satellite trust levels and network congestion states, and has the advantages of high credibility and good load balancing degree.
[0089] In summary, compared with the prior art, the present invention has the advantages of high routing security, low packet loss rate, high reach rate, good load balancing degree, high dynamic adaptability, and high QoS service quality. Description of the Drawings
[0090] Figure 1 This is the flowchart of the design scheme of the present invention.
[0091] Figure 2 This is the fully distributed multi-agent architecture of the low-earth orbit satellite network of the present invention.
[0092] Figure 3 This is the flowchart of the trusted load balancing routing strategy for the low-earth orbit satellite network of the present invention.
[0093] Figure 4 This is the graph of the routing delay performance and delay constraint performance of the present invention.
[0094] Figure 5 This is the graph of the routing packet loss rate and routing trust degree performance of the present invention.
[0095] Figure 6 This is the load balancing performance graph of the maximum queue utilization rate and average queue utilization rate of the present invention. Detailed Implementation Manner
[0096] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0097] The present invention provides a trusted load balancing routing method, system, device and medium for a low-earth orbit satellite network. This solution combines multi-agent with D3QN, with each satellite as an agent, and the agent makes routing decisions based on the status information of adjacent nodes. This solution first uses a trust evaluation model to evaluate the trust degree of satellites based on their historical behaviors, and then designs the states, actions and rewards of deep reinforcement learning, using the trust degree of satellite nodes, queue utilization rate, and QoS parameters such as the delay and bandwidth of inter-satellite links as the basis for agent decision-making. By setting relevant objective functions and constraint conditions for trust degree, load, delay, etc., a trusted load balancing routing decision model for the low-earth orbit satellite scenario is finally trained. The QoS parameters such as the credibility, load balancing performance, and delay constraint of the routing path obtained through the decision of the decision model are effectively improved.
[0098] As Figure 1 shown, a trusted load balancing routing method for a low-earth orbit satellite network includes the following steps:
[0099] Step 1: Construct the communication scenario of the low-earth orbit satellite network; establish a topology model of the low-earth orbit satellite network according to the low-earth orbit satellite network, construct network parameters of the topology model according to the network environment, and use the decentralized DTMS trust evaluation model to evaluate the trust degree of each satellite.
[0100] Further, the specific method for constructing the communication scenario of the low-earth orbit satellite network in Step 1 is as follows:
[0101] The present invention constructs the LEO-SN topology as an undirected graph G=(V, E) to characterize relevant parameters, where V={v1, v2, …, v n} represents the set of satellite nodes, v i is the i-th satellite node, and n is the total number of satellite nodes. represents the set of inter-satellite links. represents the inter-satellite link between satellite v i and satellite v j . The present invention defines τ i,j and ω i,j to describe the transmission delay and bandwidth of the inter-satellite link respectively. The present invention defines ξ i as the queue utilization rate on satellite node v i , which can be expressed as:
[0102]
[0103] where, is the number of current data packets in v i , is the total queue capacity in v i . ξ i reflects the load condition of the satellite node, that is, the higher ξ i is, the heavier the load is, and the more likely congestion occurs.
[0104] After selecting the routing path, the present invention constructs a one-dimensional array P=[v s , …, v m-1 , v m , v m+1 , …, v d to represent the satellite elements of the routing path, where v s represents the source node, v m represents the intermediate node, and v d represents the destination node. The present invention represents the links in path P as L P =[e1, e2, …, e k , …, e K , where e k represents the k-th link constituting the path. Then the transmission delay on the routing path can be calculated as:
[0105]
[0106] where, τ k represents the transmission delay of link e k . To consider the overall performance of the path and maintain a stable transmission rate, the present invention uses the bottleneck link bandwidth to represent the bandwidth of the path, and the formula is:
[0107]
[0108] Among them, ω k represents the bandwidth of link e k , then the maximum queue utilization rate and average queue utilization rate of path P can be respectively:
[0109]
[0110]
[0111] Among them, l represents the number of nodes in path P;
[0112] Use the decentralized DTMS trust evaluation model to evaluate the trust values of each satellite according to the historical behaviors of the satellites, and obtain the trust value Γ of each satellite in each state i .
[0113] Step 2: Model the quality of service constraints for the trusted load balancing routing of low-earth orbit satellites as the constraints during model solving.
[0114] Furthermore, the method for modeling the quality of service constraints for the trusted load balancing routing of low-earth orbit satellites in Step 2 is specifically as follows:
[0115] Since satellite services are diverse, more important and confidential data should be distinguished and transmitted through more trusted nodes and paths. To meet different trust requirements, the present invention dynamically evaluates the trust values of nodes and uses the DTMS trust evaluation model to evaluate historical behaviors such as attack behaviors and forwarding behaviors, which will affect the security and packet loss rate of data forwarding. Therefore, when making decisions on routing paths, the present invention determines the trust values of the next-hop nodes to meet the trust requirements of different services. For each node v in the routing path i , its trust value should satisfy:
[0116]
[0117] Among them, is the trust threshold of the mth service;
[0118] In addition, considering the high mobility of satellites and the dynamic topology of LEO-SNs, if the link conditions change, the waiting time of services may not be guaranteed. To meet the different QoS requirements of different services, the present invention sets variable QoS constraints for various services, and the present invention mainly focuses on delay-sensitive applications. When designing a routing scheme, the present invention uses a penalty strategy based on delay constraints to improve the quality of each service. Specifically, if the delay of the mth flow is greater than its QoS threshold Then the routing action will be severely punished, where the reward function results in negative feedback. In addition, the QoS threshold of the present invention is not a fixed constant, and it varies according to ensuring different latency requirements. The variable latency constraint is:
[0119]
[0120] Step 3: Model the target function of the trusted load balancing routing for low-earth orbit satellites, and construct the target function for solving the model.
[0121] Furthermore, the modeling of the target function of the trusted load balancing routing for low-earth orbit satellites in Step 3, and constructing the target function for solving the model; the specific method is:
[0122] In the present invention, the objective of the present invention is to minimize the maximum queue utilization and path latency, while ensuring the trust and latency requirements of each service. The optimization problem can be expressed as:
[0123]
[0124] C1: λ1 + λ2 = 1, λ1, λ2 ≥ 0
[0125] C2:
[0126] C3: Among them, λ1 and λ2 are the weights of the target function. C1 is the weight constraint, C2 is the latency constraint, and C3 is the trust constraint. The present invention adds the latency constraint (C2) to the optimization problem so that the obtained solution will have the ability to handle latency-sensitive services. The constraint (C3) represents the trust requirements of satellite services.
[0127] Step 4: Establish a fully distributed multi-agent architecture in the low-earth orbit satellite network topology model established in Step 1.
[0128] Furthermore, the establishment of the fully distributed multi-agent architecture in Step 4, the specific method is:
[0129] Such as Figure 2As shown in the figure, the present invention designs a fully distributed multi-agent architecture to enable the proposed solution of the present invention to have good adaptability in dynamic network topologies and environments. Compared with centralized routing algorithms, distributed routing algorithms do not require collecting global information of the network. Satellite nodes only need to collect relevant status information of adjacent nodes and adjacent links for decision-making. In a fully distributed multi-agent architecture, each agent is independently trained. Therefore, it can adapt to real-time changes in satellite links and network topologies. The present invention models the dynamic routing decision problem of each satellite as a partially observable Markov decision process (POMDP), and regards each satellite as an agent. There is no parameter sharing among agents, and they are independently trained. The routing decision of the entire network is constructed as a multi-agent POMDP.
[0130] Step 5: Establish a D3QN deep reinforcement learning training model in the fully distributed multi-agent architecture established in Step 4. Using the constraint conditions obtained in Step 2 and the objective function obtained in Step 3, set the state, action, and reward of the training model, and train the routing decision model of the multi-agent D3QN learning model.
[0131] Further, the specific method of Step 5 is as follows:
[0132] The present invention combines a multi-agent POMDP model with a D3QN algorithm to design a trustworthy load balancing routing (TLBR) for LEO-SN. Agents obtain the current state through interaction with surrounding nodes, and derive a reward value through a state transition function and a reward function after making a decision.
[0133] 5.1: State space setting;
[0134] Satellite node v i The state at time t can be described as:
[0135]
[0136] where d is a binary vector of dimension 1×N. If v i is the destination node, the i-th element in d is set to 1, and other elements are 0. The shortest hop count from the current node v i to the destination node is denoted as K i , and the shortest hop count from neighbor node j to the destination node is denoted as K j . ξ j is the queue utilization rate of neighbor node j, indicating the load condition of node j, and τ ij and ω ij respectively describe the link delay and bandwidth from v i to the j-th neighbor node. Γ j is the trust value of the j-th neighbor node.
[0137] 5.2: Action Space Setting;
[0138] The satellite agent can select its neighbor nodes according to the current state for the next-hop routing. Therefore, the current satellite node v i 's action space is represented as a i = {j}. When a i = j, it means that the agent selects the j-th neighbor node as the next-hop node.
[0139] 5.3: Reward Function Setting;
[0140] The key to the D3QN scheme is the design of the reward function. The goal of this invention is to minimize the queue utilization while ensuring the trust and service quality requirements. The specific reward design is discussed as follows.
[0141] (1) Trust Reward: The reward related to trust is designed as:
[0142]
[0143] where is the proportional adjustment parameter, and θ1 is a negative value, which is used to punish the decisions that do not meet the trust requirements.
[0144] (2) QoS Reward: The reward related to the delay service quality is designed as:
[0145]
[0146] where is the proportional adjustment parameter, and θ2 is a negative value, which is used to punish the decisions that do not meet the Qos requirements. l s is the number of hops of the shortest path of the current service.
[0147] (3) Load Reward: The reward related to the load is designed as:
[0148]
[0149] where is the proportional adjustment parameter. If the queue utilization is low, the agent will obtain a greater reward.
[0150] (4) Path Reward: In order to make the next-hop closer to the destination node, the routing-related reward of this invention is designed as:
[0151]
[0152] where is the scale adjustment parameter. K i and K jIt is used to guide the agent to find the target node faster. This reward can accelerate the convergence speed of the algorithm and avoid the occurrence of ping-pong routing. When K i -K j > 0, it means that the current decision result is closer to the destination node, and a positive reward is given to this decision. Otherwise, when K i -K j < 0, it means that the current decision result is far from the destination node, and a negative reward is given. Therefore, using K i and K j can accelerate the convergence of the routing scheme.
[0153] Therefore, the overall reward can be described as:
[0154] r i = ρ1r1 + ρ2r2 + ρ3r3 + ρ4r4
[0155] where ρ1, ρ2, ρ3, and ρ4 are positive weights, and ρ1 + ρ2 + ρ3 + ρ4 = 1.
[0156] Step 6: According to the routing decision model obtained by training in Step 5, design a trusted load balancing routing strategy for the LEO satellite network.
[0157] Furthermore, the method for designing the trusted load balancing routing strategy for the LEO satellite network in Step 6 is as follows:
[0158] When designing the trusted load balancing routing strategy for the LEO satellite network, the trusted load balancing routing algorithm is divided into a training stage and a decision-making stage.
[0159] 6.1: In the training stage, at the beginning of training, initialize the experience replay pool and network parameters. Randomly generate source node and destination node pairs in the LEO-SN topology, and the routing decision is made by the source node. During the decision-making process, the agent makes decisions according to the following dynamic greedy strategy:
[0160]
[0161] where step is the number of training steps, is a constant that can adjust the decreasing speed of the decision probability in the dynamic greedy strategy. σ is a random number in the range of [0, 1]. After each decision, calculate the reward value according to the reward function. After the agent makes a decision once, store the current sample in the experience replay pool for later use. The estimation network updates the network weights according to the gradient descent method, and the target network updates the network weights from the estimation network every C steps. When the destination node is found, one routing training ends. Repeat randomly generating source and destination node pairs for training until the model training converges.
[0162] Such as Figure 3As shown in the figure, 6.2: In the decision-making stage, according to the decision model that converges during training: a. Initialize the routing path; b. Input the source node and destination node of the current service; c. The agent of the current node makes the next-hop decision based on the surrounding environment state; d. Add the decision node to the routing path; e. Take the next-hop node as the new current node; f. Determine whether the current node is the destination node. If it is the destination node, end the routing to obtain the routing path. If it is not the destination node, jump to step c to continue the decision-making.
[0163] The present invention also provides a system for implementing the above-mentioned trusted load balancing routing method for low-earth orbit satellite networks, including:
[0164] Low-earth orbit satellite network topology construction module: Used to implement the construction of the low-earth orbit satellite network topology model in step 1, and construct the low-earth orbit satellite network into a topology graph with topological relationships and network parameters.
[0165] Multi-agent D3QN training module: Used to implement the training of the multi-agent D3QN decision model in steps 4 and 5. According to the constraint conditions and objective functions set in steps 2 and 3, use the data set to train the multi-agent D3QN decision model so that it can decide the best routing path according to the state environment of the node.
[0166] Trusted load balancing routing strategy decision module: Used to implement the trusted load balancing routing strategy in step 6. According to the multi-agent D3QN decision model trained in step 5, plan the routing paths for services with different types of service requirements.
[0167] Perform performance analysis and parameter tuning on the above-mentioned trusted load balancing routing strategy for low-earth orbit satellite networks, and compare the routing performance analysis of the proposed scheme with several benchmark schemes.
[0168] As Figure 4 shown, the delay performance of different routing schemes was tested. 100 test flows with different quality of service requirements were randomly generated. The red dashed line is the delay constraint of the current flow. Obviously, the proposed scheme can ensure the quality of service requirements of all flows, while the DQN scheme and the Dijkstra-Qu scheme fail for some flows. The path delay of the scheme under different numbers of malicious nodes was given, and it was found that the path delay increases with the increase of malicious nodes. In addition, the delay performance of this scheme is very close to that of the Dijkstra scheme, which only aims to minimize the delay. The experimental results show that this scheme can ensure the quality of service of different delay-sensitive services.
[0169] As Figure 5As shown, the packet loss performance under different scenarios was tested when there were multiple malicious nodes in the system. As the number of malicious nodes increased, the packet loss rate increased. This is because improper behaviors such as node discarding and attacks may lead to data forwarding failures. Compared with the other three benchmark scenarios, this scenario showed the lowest packet loss rate and the highest trust value. For example, when there were 20% and 30% malicious nodes in the system, the packet loss rate of this scenario was 29% and 24% lower than that of the benchmark scenarios respectively. The results indicate that this trusted routing scenario has better confidentiality against attacks from malicious nodes.
[0170] As Figure 6 shown, the maximum and average queue utilization under different routing scenarios were tested. As shown in the figure, the load balancing performance of this scenario was better than that of the DQN scenario, verifying the superiority of the designed multi-agent D3QN model. Among all the scenarios, the Dijkstra scenario showed the worst load balancing performance because it only focused on link latency. The Dijkstra-Qu scenario achieved good load balancing performance at the cost of sacrificing latency performance. The results indicate that this scenario improves the load balancing performance while maintaining a low transmission delay and meeting the trust requirements of the service. In addition, by testing several different satellite topologies, the satellite mobility was considered in the simulation. The experimental results show that this scenario has stable performance and good robustness and effectiveness in a highly dynamic LEO-SN environment.
[0171] The present invention also provides a trusted load balancing routing device for a low-earth orbit satellite network, including:
[0172] A memory for storing computer programs;
[0173] A processor for implementing the described trusted load balancing routing method for a low-earth orbit satellite network when executing the computer programs.
[0174] The present invention also provides a computer-readable storage medium storing a computer program, which can implement a trusted load balancing routing method for a low-earth orbit satellite network when executed by a processor.
Claims
1. A trusted load balancing routing method for a low-earth orbit satellite network, characterized in that: It includes the following steps: Step 1: Construct a low-earth orbit satellite network communication scenario; establish a low-earth orbit satellite network topology model according to the low-earth orbit satellite network, construct network parameters of the topology model according to the network environment, and use the decentralized DTMS trust evaluation model to evaluate the trust level of each satellite; Step 2: Model the quality of service constraints of the trusted load balancing routing for low-earth orbit satellites, which are used as constraint conditions during model solving; Step 3: Model the objective function of the trusted load balancing routing for low-earth orbit satellites to construct the objective function during model solving; Step 4: Establish a fully distributed multi-agent architecture in the low-earth orbit satellite network topology model established in Step 1; The specific method for establishing the fully distributed multi-agent architecture is as follows: Design a fully distributed multi-agent architecture, model the dynamic routing decision problem of each satellite as a partially observable Markov decision process POMDP, regard each satellite as an agent, there is no parameter sharing among agents, and they are trained independently, and construct the routing decision of the entire network as a multi-agent POMDP; Step 5: Establish a D3QN deep reinforcement learning training model in the fully distributed multi-agent architecture established in Step 4, use the constraint conditions obtained in Step 2 and the objective function obtained in Step 3, set the state, action and reward of the training model, and train the routing decision model for the multi-agent D3QN learning model; Step 6: Design a trusted load balancing routing strategy for the low-earth orbit satellite network according to the routing decision model trained in Step 5; The specific method for designing the trusted load balancing routing strategy for the low-earth orbit satellite network is as follows: The trusted load balancing routing algorithm is divided into a training stage and a decision-making stage; 6.1: In the training stage, initialize the experience replay pool and network parameters at the beginning of training, randomly generate source node and destination node pairs in the LEO-SN topology, and the routing decision is made by the source node. During the decision-making process, the agent makes decisions according to the following dynamic greedy strategy: where step is the number of training steps, is a constant that can adjust the decreasing speed of the decision probability in the dynamic greedy strategy. σ is a random number in the range of [0, 1]. After each decision, the reward value is calculated according to the reward function. After an agent makes a decision, the current sample is stored in the experience replay pool for future use. The estimation network updates the network weights according to the gradient descent method. The target network updates the network weights from the estimation network every C steps. When the destination node is found, one routing training ends. Randomly generate source and destination node pairs repeatedly for training until the model training converges; 6.2: In the decision-making stage, according to the decision model converged during training: a. Initialize the routing path; b. Input the source node and destination node of the current service; c. The agent of the current node makes the next-hop decision according to the surrounding environmental state; d. Add the decision node to the routing path; e. Take the next-hop node as the new current node; f. Judge whether the current node is the destination node. If it is the destination node, end the routing to obtain the routing path. If it is not the destination node, jump to step c to continue the decision-making.
2. The low-earth orbit satellite network trusted load balancing routing method according to claim 1, characterized in that: The specific method for constructing the low-earth orbit satellite network communication scenario in Step 1 is as follows: Construct the LEO - SN topology as an undirected graph \(G=(V, E)\), where \(V = \{v_1, v_2,\cdots, v\) n \} represents the set of satellite nodes, \(v\) i is the \(i\)-th satellite node, and \(n\) is the total number of satellite nodes. represents the set of inter - satellite links. represents the inter - satellite link between satellite \(v\) i and satellite \(v\) j . Define \(\tau\) i,j and \(\omega\) i,j to describe the transmission delay and bandwidth of the inter - satellite link respectively. Define \(\xi\) i as the queue utilization on satellite node \(v\) i , which is expressed as: Among them, is the number of current data packets in v i , is the total queue capacity in v i , ξ i reflects the load situation of the satellite node, that is, ξ i the higher it is, the heavier the load and the easier congestion occurs; After selecting the routing path, construct a one-dimensional array P = [v s , …, v m-1 , v m , v m+1 , …, v d to represent the satellite elements of the routing path, where v s represents the source node, v m represents the intermediate node, v d represents the destination node, and represent the links in the path P as L P = [e1, e2, …, e k , …, e K , where e k represents the k-th link that makes up the path. The transmission delay on the routing path is calculated as: Among them, τ k represents the transmission delay of link e k The bandwidth of the path is represented by the bottleneck link bandwidth, and the formula is: Among them, ω k represents the bandwidth of link e k The maximum queue utilization rate and the average queue utilization rate of path P are respectively: where l represents the number of nodes in path P; Use the decentralized DTMS trust evaluation model to evaluate the trust values of each satellite based on the historical behavior of the satellites, and obtain the trust value Γ of each satellite in each state i .
3. A low-earth orbit satellite network trusted load balancing routing method according to claim 2, characterized in that: The specific method for modeling the quality of service constraints of the trusted load balancing routing for low-earth orbit satellites in Step 2 is as follows: For each node v in the routing path i , its trust value should satisfy: Among them, is the trust threshold of the m-th service; Use a penalty strategy based on delay constraints to improve the quality of each service. If the delay of the m-th flow is greater than its QoS threshold then this routing action will be severely penalized, where the reward function results in negative feedback. The QoS threshold is not a fixed constant and varies to ensure different delay requirements. The constraint for variable delay is as follows:
4. A method for trusted load balancing routing in a low-earth orbit satellite network according to claim 3, characterized in that: The specific method for modeling the objective function of the trusted load balancing routing for low-earth orbit satellites in Step 3 to construct the objective function during model solving is as follows: The objective function is expressed as: C1: λ1 + λ2 = 1, λ1, λ2 ≥ 0 where λ1 and λ2 are the weights of the objective function, C1 is the weight constraint, C2 is the delay constraint, and C3 is the trust constraint.
5. A method for trusted load balancing routing in a low-earth orbit satellite network according to claim 4, characterized in that: The specific method for Step 5 is as follows: The multi-agent POMDP model is combined with the D3QN algorithm to design the Trusted Load Balancing Routing (TLBR) for LEO-SN. The agent obtains the current state through interaction with surrounding nodes, and after making a decision, derives the reward value through the state transition function and the reward function; 5.1: State space setting; Satellite node v i The state at time t is: where d is a binary vector of dimension 1×N. If v i is the destination node, the i-th element in d is set to 1 and the other elements are set to 0. The shortest hop count from the current node v i to the destination node is K i , and the shortest hop count from the neighbor node j to the destination node is K j , ξ j is the queue utilization rate of the neighbor node j, indicating the load condition of node j. τ ij and ω ij describe the link delay and bandwidth from v i to the j-th neighbor node respectively, and Γ j is the trust value of the j-th neighbor node; 5.2: Action space setting; The current satellite node v i has an action space represented as a i = {j}, when a i = j, it means that the agent selects the j-th neighbor node as the next-hop node; 5.3: Reward function setting; The specific reward design is discussed as follows: (1) Trust reward: The reward related to trust is designed as: Among them, is a proportional adjustment parameter, is a negative value used to penalize decisions that do not meet the trust requirements; (2) QoS reward: The reward related to the delay quality of service is: Among them, is the ratio adjustment parameter, is a negative value used to penalize decisions that do not meet the Qos requirements, l s is the hop count of the shortest path of the current service; (3) Load reward: The reward related to the load is: Among them, is the ratio adjustment parameter; (4) Path reward: The reward related to the routing is: Among them, is a scale adjustment parameter, and using K i and K j can accelerate the convergence of the routing scheme; when K i - K j > 0, it means that the current decision result is closer to the destination node, and a positive reward is given to this decision. Otherwise, when K i - K j < 0, it means that the current decision result is far from the destination node, and a negative reward is given; Therefore, the overall reward is described as: r i = ρ1r1 + ρ2r2 + ρ3r3 + ρ4r4 where ρ1, ρ2, ρ3, and ρ4 are positive weights, and ρ1 + ρ2 + ρ3 + ρ4 = 1.
6. A system for implementing the trusted load balancing routing method for a low-earth orbit satellite network according to any one of claims 1 to 5, comprising: A low-earth orbit satellite network topology construction module: used to implement the construction of the low-earth orbit satellite network topology model, and construct the low-earth orbit satellite network into a topology graph with topological relationships and network parameters; A multi-agent D3QN training module: used to implement the training of the multi-agent D3QN decision model. According to the set constraints and objective function, use the data set to train the multi-agent D3QN decision model so that it can decide the best routing path according to the state environment of the nodes; A trusted load balancing routing policy decision module: used to implement the trusted load balancing routing policy, and plan the routing paths for services with different types of service requirements according to the trained multi-agent D3QN decision model.
7. A trusted load balancing routing device for a low-earth orbit satellite network, characterized in that: Including: A memory for storing computer programs; A processor for implementing a trusted load balancing routing method for a low-earth orbit satellite network according to any one of claims 1-5 when executing the computer program.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement a trusted load balancing routing method for a low-earth orbit satellite network according to any one of claims 1-5.
Citation Information
Patent Citations
Load balancing routing method based on accurate link state feedback
CN114499644A
Routing method and system for load balancing of low earth orbit satellite network
CN114567365A