DQN-Based Adaptive Link State Update Method and System in LEO Satellite Network
By adopting a distributed adaptive link state update method based on DQN in the LEO satellite network, the problem of poor link state update interval design in the prior art is solved, and more efficient and adaptable routing performance is achieved.
Patent Information
- Application Number
- CN202211689958.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-27
AI Technical Summary
The poor design of link status update intervals in existing LEO satellite networks leads to inaccurate routing tables, affecting network throughput and energy efficiency, and the unified update interval cannot adapt to high dynamic characteristics and unbalanced traffic distribution.
Using a distributed adaptive link state update method based on DQN, a distributed adaptive link state update mechanism is designed, and a deep Q learning algorithm is used to make each satellite independently link state distribution decisions based on the current link state, optimizing information deviation and signaling overhead.
It improves routing performance, ensures the accuracy and adaptability of link state information, reduces signaling overhead and energy consumption, and adapts to the dynamic characteristics of large-scale LEO satellite networks.
Smart Images

Figure QLYQS_5 
Figure QLYQS_6 
Figure QLYQS_8
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite routing in satellite communication, and particularly relates to an adaptive link state update method and system based on DQN in a LEO satellite network. Background Art
[0002] In recent years, with the development of low-cost satellite platforms and satellite communication technologies, low Earth orbit (LEO) satellite networks have become a powerful supplement to terrestrial cellular networks and will play an important role in the upcoming 6G communication system. Due to the highly dynamic topology of satellite networks and the increasing traffic load, dynamic routing is considered the mainstream direction for routing design in large-scale LEO satellite networks. Dynamic routing calculates the routing table based on the collected near-real-time link state information, such as link load, traffic conditions, etc. The link state update directly affects the quality of routing. Therefore, the design of a satellite link state update scheme is one of the important issues in routing for large-scale LEO satellite networks.
[0003] Although there have been many studies in the field of LEO satellite routing design and link state update, these studies follow the mechanism of distributing link state information with a fixed and unified update interval for all satellites, and there are the following problems. First, the setting of the link state update interval is not well designed. A larger link state update period will result in a significant difference between the link state used for routing calculation and the actual state, thus affecting the accuracy of the routing table. Inaccurate routing will lead to a large number of packet losses, resulting in a decrease in network throughput and energy efficiency. On the contrary, a smaller update interval will greatly increase the update frequency, causing a large signaling overhead and energy consumption, and correspondingly also leading to low energy efficiency performance. Second, for the highly dynamic characteristics of LEO satellite networks, the fixed link state update interval has poor flexibility. Finally, due to the unbalanced traffic distribution across the entire constellation, the load changes of different satellites and links are inconsistent. The link states of some satellites change violently, while the link states of some satellites are relatively stable. Then a unified update interval cannot be the optimal setting for each satellite. Therefore, it is necessary to design a distributed adaptive link state update method for routing in large-scale LEO satellite networks, where each satellite independently makes a link state distribution decision based on the current actual link state. Summary of the Invention
[0004] To solve the problems existing in the prior art, the present invention provides a distributed adaptive link state update method based on DQN in routing for large-scale LEO satellite networks. By fully considering the heterogeneity of satellites and the time-varying nature of satellite link states in large-scale low-orbit satellite networks, an adaptive distributed satellite link state information update scheme for large-scale LEO satellite networks is designed, improving the routing performance of the system.
[0005] To achieve the above object, the technical solution adopted by the present invention is: an adaptive link state update method based on DQN in a LEO satellite network, including the following steps:
[0006] Based on the characteristics of a large-scale LEO satellite network and routing requirements, design a distributed adaptive link state update mechanism for large-scale LEO satellites;
[0007] Based on the update mechanism, define information deviation and signaling overhead to characterize the timeliness and cost of link state information update, construct a multi-objective optimization problem that simultaneously minimizes information deviation and minimizes signaling overhead, and use the weighted sum method to transform the multi-objective optimization problem into a single-objective optimization problem, requiring each satellite to make a decision on link state information distribution, and the optimal decision changes in real time;
[0008] Model the link state information distribution decision of a single satellite as a Markov decision process, and use the deep Q-learning algorithm. The satellite learns and optimizes the link state information distribution decision through continuous interaction with the environment.
[0009] The large-scale satellite network considers a polar orbit LEO satellite constellation, which consists of N satellites in total, evenly distributed on M orbital planes. The satellite index set is expressed as Each satellite establishes four inter-satellite links, including two in-orbit ISLs and two inter-orbit ISLs; the connection of the in-orbit ISLs always remains stable. The gap between the first orbit and the last orbit is called the "reverse seam". Two satellites connected by an inter-satellite link are called neighbor satellites to each other. A cache queue is assigned to each ISL, and the queue length of the queue Q(i, j) at time slot t is denoted as q i,j,t , the dynamic satellite network is modeled as a sequence of time-discrete snapshots using a virtual topology strategy. The time is divided into T time slots, and the length of each time slot is τ s , the time slot sequence set is expressed as The network topology is updated every T net time slots.
[0010] The distributed adaptive link state update mechanism is specifically: define a decision interval including T lsi time slots. Each satellite makes an adaptive link state information distribution decision every T lsi time slots. If the satellite chooses to perform link state distribution, it uses the flooding method to distribute the latest link state to the whole network. If it chooses not to distribute, it only needs to wait to see if it receives the link state information of other satellites. If it receives it, it performs routing calculation. When it is not a decision time slot, all satellites do not distribute.
[0011] Express the queue length of the queue Q(i, j) mastered by other satellites at time slot t as Define q i,j,t The distance between is the information deviation, that is
[0012]
[0013] If the routing signaling overhead is represented by the number of link state packets, then the signaling overhead for the satellite to distribute link state information once is expressed as:
[0014]
[0015] Among them, depending on the specific network topology, the first term means that the source satellite generating the link state packet needs to forward the state packet to all four neighbor satellites, the second term means that other relay satellites need to forward the state packet in three directions, and the last term means the sum of the number of inter-satellite links disconnected in the entire satellite network at time slot t, represents the number of inter-satellite links disconnected by satellite i at time slot t, then
[0016] With the goal of simultaneously minimizing the information deviation and minimizing the signaling overhead, the following multi-objective optimization problem is constructed:
[0017]
[0018]
[0019]
[0020] Among them, the optimization variable is the link state information distribution decision matrix. Considering the information deviation and the signaling overhead comprehensively, using the weighted sum method, the MOP is converted into a single-objective optimization problem. After normalizing ID and SO, the optimization problem Q1 can be transformed into:
[0021]
[0022]
[0023] Among them, ID max and SO max respectively represent the maximum values of ID and SO, and ω and 1 - ω respectively represent the weight factors of the information deviation and the signaling overhead.
[0024] Model the link state information distribution decision of a single satellite as a Markov decision process. The settings of its state, action and reward function are specifically as follows:
[0025] ① State, at each decision time slot (t = kT lsi , k ∈ Z +)Update the state of satellite i at the start time, and its state is represented as:
[0026]
[0027] where v i,j,t represents the change intensity of the cache queue length within the previous decision interval, and is calculated as follows:
[0028]
[0029] ② Action. The task of the agent is to make a decision on the distribution of link state information. Denote the action of satellite i at time slot t as:
[0030]
[0031] ③ Reward. Corresponding to the above optimization problem, use the weighted sum of the cumulative information deviation and signaling overhead of the satellite within the decision interval as the reward. The reward function can be expressed as:
[0032]
[0033] where neb(i) represents the set of neighbor satellites of satellite i, T lsi represents the link state information distribution decision period, Δ i,j,m represents the information deviation of the queue Q(i, j) at time slot t, L buffer represents the maximum length of the cache queue, which is the maximum value that Δ i,j,m may take. represents the signaling overhead generated by a satellite distributing link state information once at time slot t, represents the maximum value of, and β1 and β2 respectively represent the weight factors of information deviation and signaling overhead.
[0034] Based on the standard Markov decision process, use the deep Q-learning algorithm to solve the optimization problem. Establish a double-network structure, including the current value network Q(s, a; θ) and the target network Q(s, a; θ′). The network parameters are θ and θ′ respectively. For the current value network, at the start of each decision time slot, the agent takes the current state as the input and selects an action based on the ε-greedy strategy:
[0035]
[0036] The agent distributes link state information according to the selected action, calculates the corresponding reward and the next state at the end of the decision interval, and the agent stores the quadruple experience (s, a, r, s′) in the experience pool Among them, in the training stage, prioritized experience replay is adopted to extract a batch of samples from the experience pool, and the gradient descent algorithm is used to update the parameters of the current value network. The loss function is expressed as:
[0037]
[0038] where γ is the discount factor, Q(s′, a′; θ) is the Q value calculated by the target network. The target network has the same network structure as the estimation network, and the parameters of the current value network are copied every certain number of steps to update its parameters.
[0039] On the other hand, the present invention provides an adaptive link state update system based on DQN in a LEO satellite network, including a link state update mechanism design module, a problem modeling module, and a solution module;
[0040] The link state update mechanism design module is used to design a distributed adaptive satellite link state update mechanism during the routing and data transmission processes based on a large-scale LEO satellite network;
[0041] The problem modeling module is used to jointly consider information deviation and signaling overhead, design a multi-objective optimization problem to minimize both information deviation and signaling overhead, and use the weighted sum method to transform it into a single-objective optimization problem;
[0042] The solution module models the satellite link state information distribution decision problem using a Markov process and adopts a deep Q-learning algorithm to adaptively learn the optimal link state information distribution decision strategy.
[0043] The present invention can also provide a satellite terminal for communication in a large-scale LEO satellite network, including a processor and a memory; the memory is used to store computer-executable programs, and the processor reads part or all of the computer-executable programs from the memory and executes them. When the processor executes part or all of the computer-executable programs, it can implement the adaptive link state update method based on DQN in the LEO satellite network of the present invention.
[0044] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention fully considers the characteristics of a large-scale LEO satellite network and the requirements of dynamic routing, designs a set of link state update mechanisms for routing in a large-scale LEO satellite network, enabling the satellite to independently make adaptive link state information distribution decisions according to the current link state situation, thereby adjusting its own link state information distribution period; jointly considers information deviation and signaling overhead, designs a multi-objective optimization problem to minimize both information deviation and signaling overhead, and uses the weighted sum method to transform it into a single-objective optimization problem; adopts a deep Q-learning algorithm, enabling the satellite to adaptively learn the optimal link state information distribution decision through continuous interaction with the environment, improving the routing performance. Description of the Drawings
[0045] Figure 1 Schematic diagram of the routing calculation process designed for the present invention based on distributed adaptive link state update.
[0046] Figure 2 Schematic diagram of the deep Q - learning network structure established for the present invention. Detailed Embodiment
[0047] The present invention will be elaborated in detail below with reference to the accompanying drawings.
[0048] Consider a polar - orbiting LEO satellite constellation, which consists of N satellites in total, evenly distributed in M orbital planes. The satellite index set is denoted as Each satellite establishes four inter - satellite links (ISLs), including two in - orbit ISLs and two inter - orbit ISLs. The in - orbit ISLs connect adjacent satellites in the same orbit, and the inter - orbit ISLs connect adjacent satellites on adjacent orbits. The connection of the in - orbit ISLs always remains stable, while the inter - orbit ISLs will be temporarily disconnected when entering the high - latitude polar region due to the limitation of antenna tracking technology. The gap between the first orbit and the last orbit is called the "reverse seam". Since the satellite motion directions on both sides of the reverse seam are opposite, it is difficult to maintain a stable connection for the inter - orbit ISLs. Two satellites connected by an inter - satellite link are mutually called neighbor satellites.
[0049] Regarding the dynamic characteristics of the LEO satellite network, a virtual topology strategy is adopted to model the dynamic satellite network as a sequence of time - discrete snapshots. The time is divided into T time slots, and the length of each time slot is τ s , and the time - slot sequence set is denoted as The network topology is updated every T net time slots, and the network topology remains unchanged in other time slots.
[0050] As described above, each satellite can establish at most four inter - satellite links, and a buffer queue is assigned to each ISL to temporarily store the data packets to be sent. The total length of the buffer queue is L buffer . Q(i, j) represents the buffer queue maintained by satellite i for the inter - satellite link between satellite i and satellite j. When a data packet arrives, the satellite queries its own routing table and adds the data packet to the corresponding buffer queue. The queue scheduling policy follows the first - in first - out (FIFO) rule, and the packets exceeding the queue will be discarded. The satellite monitors and records the change in the buffer queue length in each time slot. Assume that the traffic received by each satellite from the ground station follows a Poisson process, and the specific access process between the ground station and the satellite is not considered. Then at the beginning of time slot t, the queue length of queue Q(i, j) is denoted as q i,j,t , which is denoted as
[0051] q i,j,t = min{L buffer , max{0, q i,j,t-1 + I i,j,t-1 - O i,j,t-1}
[0052] where I i,j,t represents the number of data packets arriving at queue Q(i, j) within time slot t, and O i,j,t represents the number of data packets leaving queue Q(i, j) within time slot t. In addition, I i,j,t includes data packets from neighboring satellites and data packets generated by the satellite itself, that is
[0053]
[0054] where neb(i, t) represents the set of neighboring satellites of satellite i at time slot t, G i,j,t and respectively represent the number of data packets generated by satellite i itself and added to queue Q(i, j) and the number of data packets from neighboring satellite h.
[0055] The DQN-based distributed adaptive link state update method for large-scale LEO satellite networks is as follows:
[0056] In the LEO satellite network, the network topology and link state information affect the propagation delay and queuing delay respectively. When the network topology or link state information is updated, the satellite needs to recalculate the routing table. The network topology can be calculated in advance according to the movement law of the constellation. For the distribution of link state information, the present invention adopts the flooding method. Specifically, the satellite first sends the link state packet generated by itself to all neighboring satellites; when the neighboring satellite receives the link state packet, it checks whether the data packet has been received before. If it receives the data packet for the first time, it forwards the data packet to other directions except the direction from which the data packet comes, otherwise it directly discards it. The link state information of satellite i at time slot t includes the set of neighboring satellites of satellite i and the queue length of the corresponding inter-satellite link, that is
[0057] LSI i,t = {neb(i, t), q(i, t)}
[0058] where q(i, t) = {q i,j,t , j ∈ neb(i, t)}. After other satellites receive the link state data packet, they can calculate the queuing delay QD i,j,t of this link, and the calculation is as follows:
[0059]
[0060] where Lpkt represents the data packet size, and C represents the inter-satellite link capacity.
[0061] Different from the existing work, not all satellites distribute link state information based on fixed-period synchronization. In the present invention, each satellite independently makes a decision on link state information distribution according to its actual link state change. The present invention defines a decision interval including T lsi time slots. Each satellite makes a decision on link state information distribution every T lsi time slots. Let represent the link state information distribution decision of satellite i at time slot t, where means that satellite i distributes the updated link state information to the whole network at time slot t, means that satellite i does not perform distribution. When t is not a decision time slot, is always zero.
[0062] In addition, the time required for a satellite to receive status data packets from different satellites is slightly different, that is, a satellite will first receive the status data packets of satellites that are closer, and then the status data packets of satellites that are farther away will arrive. In order to make the routing calculation converge better and save computing resources at the same time, the present invention sets a fixed time period including T d time slots after each decision time slot for link state information distribution between satellites. After T d time slots, each satellite performs unified routing calculation according to all the received link state information packets. The routing calculation process of each satellite based on adaptive distributed link state update is as Figure 1 shown.
[0063] In the low-earth orbit satellite network, the one-hop propagation delay is relatively large, and the link state information distributed is always an outdated information, that is, the latest updated link state information received by other satellites is not exactly the same as its current true state. The present invention represents the queue length of the queue Q(i, j) mastered by other satellites at time slot t as Let q i,j,t and The distance between them is defined as the information deviation. The smaller this value is, the higher the accuracy of the information, which is specifically expressed as:
[0064]
[0065] According to the foregoing link state information distribution scheme, using the number of link state data packets to represent the routing signaling overhead, the signaling overhead for a satellite to distribute link state information once is expressed as
[0066]
[0067] Depending on the specific network topology, specifically, the first term indicates that the source satellite generating the link state packet needs to forward the state packet to all four neighbor satellites. The second term indicates that the other relay satellites need to forward the state data packets in three directions. The last term represents the sum of the number of inter-satellite links that are disconnected in the entire satellite network at time slot t. represents the number of inter-satellite links disconnected by satellite i at time slot t, then
[0068] Adopting the above link state information distribution mechanism, the low Earth orbit satellite network faces the following challenges. On the one hand, the timely update of link state information is beneficial to improving the accuracy of link state information and making positive routing adjustments to the satellites, thus ensuring the accuracy and effectiveness of routing. On the other hand, frequent distribution of link state information will bring a large signaling overhead, causing resource consumption such as bandwidth and on-board computing, which will also in turn affect the routing performance. Accordingly, aiming at minimizing both the information deviation and the signaling overhead simultaneously, the following multi-objective optimization (MOP) problem can be constructed.
[0069]
[0070]
[0071]
[0072] Among them, the optimization variable is the link state information distribution decision matrix. Obviously, these two objectives are contradictory. Considering both the information deviation and the signaling overhead comprehensively and using the weighted sum method, the MOP is converted into a single-objective optimization problem (SOP). In addition, since the orders of magnitude of ID and SO are different, it is necessary to normalize them. Furthermore, the optimization problem Q1 can be transformed into:
[0073]
[0074]
[0075] where ID max and SO max represent the maximum values of ID and SO respectively, and ω and 1 - ω represent the weight factors of the information deviation and the signaling overhead respectively.
[0076] For the above optimization problem, each satellite is required to make a decision on link state information distribution, and the optimal decision is time-varying. Moreover, the link state information distribution decision in the current time slot will affect the link state in subsequent time slots, so it is difficult to solve using traditional optimization theories and methods. Therefore, it is modeled as a Markov Decision Process (MDP), and deep reinforcement learning is used to solve this problem in an intelligent way.
[0077] For a single satellite, whether to distribute its link state information depends on its actual link state, including buffer queue length information and traffic load changes. First, links with large information deviations need to be updated in a timely manner. Second, the change in buffer queue length is a direct indication of traffic load changes. When the buffer queue length changes drastically, the information accuracy of the current link is likely to be poor, and link state information distribution is required to guide route adjustment. On the contrary, when the queue length of the current link changes smoothly, the satellite can choose not to distribute to reduce signaling overhead. The settings of the state, action, and reward functions are introduced in detail below.
[0078] (1) State: At the start time of each decision time slot (t = kT lsi , k ∈ Z + ), the state of satellite i is updated, and its state is expressed as:
[0079]
[0080] where v i,j,t represents the change intensity of the buffer queue length within the previous decision interval, and is calculated as follows:
[0081]
[0082] (2) Action: The task of the agent is to make a decision on link state information distribution. Therefore, the action of satellite i at time slot t is denoted as
[0083]
[0084] (3) Reward: Corresponding to the above optimization problem, the reward function of the satellite is affected by both information deviation and signaling overhead, and the weighted sum of the cumulative information deviation and signaling overhead of the satellite within this decision interval is used as the benefit. At the same time, since the optimization problem is a minimization problem, a negative sign needs to be added to the reward function. Then the reward function can be expressed as
[0085]
[0086] where, neb(i) represents the set of neighbor satellites of satellite i, T lsi represents the link state information distribution decision period, and Δ i,j,mDenote the information deviation of queue Q(i, j) at time slot t, L buffer Denote the maximum length of the buffer queue, which is Δ i,j,m The maximum value that may occur Denote the signaling overhead generated by a satellite distributing link state information once at time slot t Denote The maximum value of, where β1 and β2 denote the weight factors of information deviation and signaling overhead respectively
[0087] Based on the above MDP framework, the DQN algorithm is applied to learn the optimal link state information allocation decision strategy. Specifically, a double-network structure is established in the present invention, as Figure 2 shown, including the current value network Q(s, a; θ) and the target network Q(s, a; θ′), and the network parameters are θ and θ′ respectively. For the current value network, at the beginning of each decision time slot, the agent takes the state as the input and selects an action based on the ε-greedy strategy
[0088]
[0089] The agent distributes link state information according to the selected action, calculates the corresponding reward and the next state at the end of the decision interval, and then the agent stores the quadruple experience (s, a, r, s′) in the experience pool In the training stage, priority experience replay is adopted to extract a batch of samples from the experience pool. Then the gradient descent algorithm is used to update the estimated network parameters, and the loss function is expressed as
[0090]
[0091] where γ is the discount factor and Q(s′, a′; θ) is the Q value calculated by the target network. The target network has the same network structure as the estimation network, and the parameters of the evaluation network are copied every certain number of steps to update its parameters. The specific process of the algorithm is as described in Algorithm 1
[0092]
[0093]
[0094] On the other hand, the present invention provides a distributed adaptive link state update system for a large-scale LEO satellite network, including a link state update mechanism design module, a problem modeling module, and a solution module
[0095] The link state update mechanism design module is used to design a distributed adaptive satellite link state update mechanism during the routing and data transmission processes based on a large-scale LEO satellite network
[0096] The problem modeling module is used to jointly consider information deviation and signaling overhead, design a multi-objective optimization problem to minimize both information deviation and signaling overhead, and use the weighted sum method to transform it into a single-objective optimization problem;
[0097] The solution module models the satellite link state information distribution decision problem using a Markov process and adopts a deep Q-learning algorithm to adaptively learn the optimal link state information distribution decision strategy.
[0098] The present invention can also provide a satellite terminal for communication in a large-scale LEO satellite network, including a processor and a memory; the memory is used to store computer-executable programs, and the processor reads part or all of the computer-executable programs from the memory and executes them. When the processor executes part or all of the computer-executable programs, it can implement the DQN-based adaptive link state update method in the LEO satellite network of the present invention.
[0099] The above content is a detailed description of the present invention. It cannot be determined that the present invention is limited thereto. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope determined by the claims submitted for the present invention.
Claims
1. An adaptive link state update method based on DQN in a LEO satellite network, characterized in that The steps include: Based on the characteristics and routing requirements of large-scale LEO satellite networks, a distributed adaptive link state update mechanism for large-scale LEO satellites is designed; the specific distributed adaptive link state update mechanism is as follows: Define a decision interval including T lsi time slots, and each satellite makes an adaptive link state information distribution decision every T lsi time slots. If the satellite chooses to perform link state distribution, the latest link state is distributed to the entire network using the flooding method. If it chooses not to distribute, it only needs to wait to see if it receives link state information from other satellites. If it receives it, routing calculation is performed. When it is not a decision time slot, all satellites do not distribute; Based on the above update mechanism, define information deviation and signaling overhead to characterize the timeliness and cost of link state information update, construct a multi-objective optimization problem that simultaneously minimizes information deviation and signaling overhead, and use the weighted sum method to transform the multi-objective optimization problem into a single-objective optimization problem, requiring each satellite to make a decision on link state information distribution, and the optimal decision changes in real time; Model the link state information distribution decision of a single satellite as a Markov decision process, and use the deep Q-learning algorithm. The satellite learns and optimizes the link state information distribution decision through continuous interaction with the environment.
2. The adaptive link state update method based on DQN in the LEO satellite network according to claim 1, wherein The large-scale satellite network considers a constellation of LEO satellites in polar orbits, consisting of a total of N satellites evenly distributed on M orbital planes. The satellite index set is denoted as Each satellite establishes four inter-satellite links, including two in-orbit ISLs and two inter-orbit ISLs. The connection of the in-orbit ISLs remains stable all the time. The gap between the first orbit and the last orbit is called the "reverse seam". Two satellites connected by an inter-satellite link are called neighbor satellites to each other. A buffer queue is assigned to each ISL, and the queue length of the queue Q(i, j) at time slot t is denoted as q i,j,t , and the dynamic satellite network is modeled as a sequence of time-discrete snapshots using the virtual topology strategy. The time is divided into T time slots, and the length of each time slot is τ s , and the set of time slot sequences is denoted as The network topology is updated every T net time slots.
3. The adaptive link state update method based on DQN in the LEO satellite network according to claim 1, wherein Denote the queue length of the queue Q(i, j) mastered by other satellites at time slot t as Define q i,j,t The distance between is the information deviation, that is Use the number of link state data packets to represent the routing signaling overhead, then the signaling overhead of a satellite distributing link state information once is expressed as: Among them, Depending on the specific network topology, the first item indicates that the source satellite generating the link state packet needs to forward the state packet to all four neighbor satellites, the second item indicates that the other relay satellites need to forward the state data packet in three directions, and the last item indicates the sum of the number of inter-satellite links disconnected in the entire satellite network at time slot t. represents the number of inter-satellite links disconnected by satellite i at time slot t, then 4. The adaptive link state update method based on DQN in the LEO satellite network according to claim 1, wherein With the goal of simultaneously minimizing information deviation and signaling overhead, construct the following multi-objective optimization problem: Among them, the optimization variable is the link state information distribution decision matrix. Considering information deviation and signaling overhead comprehensively, using the weighted sum method, the MOP is converted into a single-objective optimization problem, and the ID and SO are normalized. Furthermore, the optimization problem Q1 can be transformed into: Among them, ID max and SO max respectively represent the maximum values of ID and SO, and ω and 1 - ω respectively represent the weight factors of information deviation and signaling overhead.
5. The adaptive link state update method based on DQN in the LEO satellite network according to claim 4, wherein Model the link state information distribution decision of a single satellite as a Markov decision process, and the settings of its state, action, and reward function are as follows: ① State, update the state of satellite i at the start of each decision time slot (t = kT lsi , k ∈ Z + ), and its state is represented as: where v i,j,t represents the change intensity of the cache queue length in the previous decision interval, and is calculated as follows: ② Action. The task of the intelligent agent is to make a decision on link state information distribution. Denote the action of satellite i at time slot t as: ③ Reward. Corresponding to the above optimization problem, use the weighted sum of the cumulative information deviation and signaling overhead of the satellite within this decision interval as the benefit, and the reward function can be expressed as: where, neb(i) represents the set of neighbor satellites of satellite i, T lsi represents the link state information distribution decision period, Δ i,j,m represents the information deviation of queue Q(i, j) at time slot t, L buffer represents the maximum length of the cache queue, which is the maximum value that Δ i,j,m may reach, represents the signaling overhead generated by a satellite distributing link state information once at time slot t, represents the maximum value of, where β1 and β2 represent the weight factors of information deviation and signaling overhead respectively.
6. The adaptive link state update method based on DQN in the LEO satellite network according to claim 5, wherein Based on the standard Markov decision process, use the deep Q-learning algorithm to solve the optimization problem, and establish a double-network structure, including the current value network Q(s,a;θ) and the target network Q(s,a; θ′), and the network parameters are θ and θ′ respectively. For the current value network, at the beginning of each decision time slot, the intelligent agent takes the current state as the input and selects an action based on the ε-greedy strategy: The agent distributes link state information according to the selected action, calculates the corresponding reward and the next state at the end of the decision-making interval, and the agent stores the quadruple experience (s, a, r, s′) in the experience pool. In the training phase, prioritized experience replay is adopted to extract a batch of samples from the experience pool, and the gradient descent algorithm is used to update the current value network parameters. The loss function is expressed as: where γ is the discount factor, Q(s′,a′;θ) is the Q value calculated by the target network, and the target network has the same network structure as the estimation network. The parameters of the target network are updated by copying the parameters of the current value network every certain number of steps.
7. A distributed adaptive link state update system based on DQN for large-scale LEO satellite networks, characterized in that It includes a link state update mechanism design module, a problem modeling module, and a solution module; The link state update mechanism design module is used to design a distributed adaptive satellite link state update mechanism during the routing and data transmission processes based on a large-scale LEO satellite network; the specific distributed adaptive link state update mechanism is as follows: Define a decision interval including T lsi time slots, and each satellite makes an adaptive link state information distribution decision every T lsi time slots. If the satellite chooses to perform link state distribution, it uses the flooding method to distribute the latest link state to the entire network. If it chooses not to distribute, it only needs to wait to see if it receives the link state information of other satellites. If it does, it performs routing calculation. When it is not a decision time slot, all satellites do not distribute; The problem modeling module is used to jointly consider information deviation and signaling overhead, design a multi-objective optimization problem, simultaneously minimize information deviation and signaling overhead, and use the weighted sum method to transform it into a single-objective optimization problem; The solution module models the satellite link state information distribution decision problem using a Markov process and uses the deep Q-learning algorithm to adaptively learn the optimal link state information distribution decision strategy.
8. A satellite terminal, characterized in that, Communicating in a large-scale LEO satellite network, including a processor and a memory; the memory is used to store computer-executable programs, the processor reads part or all of the computer-executable programs from the memory and executes them, and when the processor executes part or all of the computer-executable programs, it can implement the DQN-based adaptive link state update in the LEO satellite network described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-layer dynamic self-adapting routing method based on LEO satellite network
CN103685025A
Multi-path optimization algorithm planning method based on medium and low earth orbit satellite network
CN106452555A