Multi-satellite cooperative distributed routing method for multi-agent reinforcement learning

Through the multi-star collaborative distributed routing method of multi-agent reinforcement learning, the problems of dynamic topology changes, user demand growth and uneven traffic distribution in low-orbit satellite networks are solved, and distributed routing decisions and load balancing are realized, ensuring the end-to-end transmission delay of services.

CN119995693AActive Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510451360.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Insufficient routing flexibility, difficulty in scaling and large communication overhead caused by dynamic topology changes, growth in user demand and uneven traffic distribution in low-orbit satellite networks.

Method used

Multi-star collaborative distributed routing method using multi-agent reinforcement learning is adopted. By obtaining satellite network structure data, a static topology model is built, an objective function that minimizes the end-to-end delay of data packets is established, a satellite agent network is built, and a distributed routing decision is realized through interactive experience data.

Benefits of technology

It enhances the decision-making ability and collaboration efficiency of satellite agents, realizes distributed routing decisions, solves the problems of dynamic topology changes, user demand growth and uneven traffic distribution in low-orbit satellite networks, ensures the end-to-end transmission delay of services, and realizes load balancing between satellite nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995693A_ABST
    Figure CN119995693A_ABST
Patent Text Reader

Abstract

The invention provides a multi-satellite cooperative distributed routing method for multi-agent reinforcement learning, and the method comprises the steps: building a satellite network static topology model through a time slicing technology, and building a target function for minimizing the end-to-end time delay of a data packet; constructing a satellite agent network based on the satellite network static topology model, and obtaining satellite interaction experience data; wherein the satellite agent network comprises a multi-satellite cooperative hybrid network and a satellite decision network corresponding to each agent; training a satellite agent network according to the satellite interaction experience data to obtain a trained satellite agent network; and respectively deploying the satellite decision networks in the trained satellite agent network to corresponding agents, so that the agents perform routing decision based on the deployed satellite decision networks. Through cooperative work of the satellite decision network and the multi-satellite cooperative hybrid network, efficient distributed routing decision of a low-orbit satellite constellation can be realized, time delay is reduced, and load is balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of low-orbit satellite routing optimization, and in particular to a multi-satellite collaborative distributed routing method using multi-agent reinforcement learning. Background Art

[0002] In recent years, with the rapid development of large low-orbit satellite constellation networks, the scale and complexity of the network have increased dramatically, accompanied by the continuous growth of business types and demand. The uneven distribution of user traffic in time and space often causes some inter-satellite links to experience sudden traffic peaks, causing local congestion or even link interruption. This not only increases the packet loss rate, but also reduces the quality of service (QoS).

[0003] Traditional centralized routing methods cannot dynamically adjust routing plans according to real-time network conditions, and thus have shortcomings such as insufficient flexibility, difficulty in expansion, and high communication overhead. These limitations make it difficult to meet the needs of different business flows. Summary of the invention

[0004] The purpose of the present invention is to provide a multi-satellite collaborative distributed routing method based on multi-agent reinforcement learning to solve the problems of dynamic changes in topology, continuous growth in user demand, and uneven distribution of network traffic in low-orbit satellite networks.

[0005] The present invention provides a multi-satellite collaborative distributed routing method for multi-agent reinforcement learning, comprising: Obtain the network structure data of the low-orbit satellite constellation and use time slicing technology to build a static topology model of the satellite network; each satellite in the low-orbit satellite constellation is regarded as an intelligent agent; Based on the static topology model of satellite network, the objective function of minimizing the end-to-end delay of data packets is established; A satellite agent network is constructed based on the static topology model of the satellite network, and satellite interaction experience data is obtained; the satellite agent network includes a multi-satellite collaborative hybrid network and a satellite decision network corresponding to each agent. The multi-satellite collaborative hybrid network is used to calculate the joint reward according to the objective function, and the satellite decision network is used to output the routing decision of the agent according to the current local observation information of the agent. The satellite agent network is trained based on the satellite interaction experience data to obtain the trained satellite agent network; each satellite decision network is used to calculate the local value function based on the local observation information, and the multi-satellite collaborative hybrid network is used to generate the global value function based on the global state information and all local value functions. The global value function and the joint reward are used to update the network parameters; The satellite decision networks in the trained satellite agent network are deployed to the corresponding agents respectively, so that the agents make routing decisions based on the deployed satellite decision networks.

[0006] In an optional embodiment, the network structure data includes the number of satellite orbits, the number of satellites on each orbit, the orbit altitude and two rows of orbit data, and the satellite network static topology model includes an undirected graph consisting of a set of satellite nodes and a set of links between satellites in each time slice.

[0007] In an optional embodiment, the step of establishing an objective function for minimizing the end-to-end delay of a data packet based on a static topology model of a satellite network includes: Based on the static topology model of satellite network, the inter-satellite link model and data packet delay model are established by using logical address and graph theory methods. Based on the intersatellite link model and data packet delay model, the satellite routing problem is expressed as a multi-objective optimization problem with the goal of minimizing the end-to-end delay of data packets, and the objective function is obtained.

[0008] In an optional embodiment, the objective function is expressed as: ; ; Where K represents the set of data packets, Indicates the number of packets, D( k ) indicates a data packet k On the path P k The end-to-end delay on , V k Indicates data packet k The set of nodes passed, E k Indicates data packet k The set of edges along the way, V t represents the set of satellite nodes, E t represents the set of links between satellites, H Indicates the maximum lifetime of a data packet, dis i,j Indicates satellite i With satellite j The distance between Indicates satellite i With satellite j The maximum viewing distance between Indicates data packet k exist t Is the satellite in the moment? i In the queue, Indicates data packet k exist t Is the satellite in the moment? g ,i Transmission between Z t,i Indicates satellite i exist t The neighbor set at time, Indicates satellite i The maximum capacity of the queue, Indicates data packet k exist t Is the satellite in the moment? i was discarded.

[0009] In an optional implementation, the step of acquiring satellite interaction experience data includes: Build the satellite agent interaction environment and experience replay pool, and initialize the network parameters of the satellite agent network; The satellite interaction experience data generated by each agent through the interaction between the initialized satellite agent network and the satellite agent interaction environment is obtained, and the satellite interaction experience data is stored in the experience replay pool; wherein, each agent obtains local observation information from the satellite agent interaction environment, and makes routing decisions through the initialized corresponding satellite decision network, the central node in each agent obtains the local observation information and decision action of each agent, and calculates the joint reward through the initialized multi-satellite collaborative hybrid network.

[0010] In an optional implementation, the satellite interaction experience data includes multiple data samples in different time slices, and each data sample includes global state information, joint local observation information, joint actions and joint rewards.

[0011] In an optional implementation, the satellite interaction experience data is stored in an experience playback pool; and the step of training the satellite agent network according to the satellite interaction experience data to obtain the trained satellite agent network includes: Perform non-uniform sampling in the experience replay pool to obtain target sample data; According to the target sample data, the global value function is obtained through the satellite agent network; According to the joint reward in the global value function and the target sample data, the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network are updated through the back propagation algorithm.

[0012] In an optional embodiment, the sampling probability of each data sample in the satellite interaction experience data is p t as follows: ; ; in, ε Indicates the preset value. ε To prevent the sampling probability from being zero, Indicates status Under this condition, the error between the predicted value and the true value of the multi-satellite collaborative hybrid network is s t express t The global state at the moment, a t express t The joint action of the moment, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of Indicates that the multi-satellite cooperative hybrid network is in state ( ), network parameters The value of r t express t Joint rewards of the moment, represents the preset discount factor, and A represents the action space of the satellite.

[0013] In an optional embodiment, the satellite agent network further includes a target hybrid network and a target decision network corresponding to the multi-satellite collaborative hybrid network and the satellite decision network, respectively; at predetermined intervals, the network parameters of the target hybrid network are adjusted to be the same as the network parameters of the multi-satellite collaborative hybrid network, and the network parameters of the target decision network are adjusted to be the same as the network parameters of the corresponding satellite decision network; and the steps of updating the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network by a back propagation algorithm according to the joint reward in the global value function and the target sample data include: According to the joint reward in the target sample data, the target value function is obtained through the target hybrid network; According to the global value function and the target value function, the target loss function is calculated; According to the objective loss function, the gradient descent method is used to update the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network.

[0014] Target value function is defined as: ; The objective loss function is defined as: ; Network parameters of multi-satellite cooperative hybrid network and the network parameters of the satellite decision network The update formula is: ; ; in, rt express t Joint rewards of the moment, represents the preset discount factor, Indicates that the target hybrid network is in state ( ), network parameters The value of s t+1 express t+ The global state at time 1, Expressed as t+ 1 moment is used to maximize the joint action of the local value function, Represents the network parameters of a multi-satellite collaborative hybrid network and satellite i The corresponding network parameters of the satellite decision network The objective loss function under b represents the target sample data sampled from the experience replay pool, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of α t Represents a preset learning rate that decreases over time.

[0015] The multi-satellite collaborative distributed routing method of multi-agent reinforcement learning provided by the present invention can obtain the network structure data of a low-orbit satellite constellation and construct a satellite network static topology model by using a time slicing technology; wherein each satellite in the low-orbit satellite constellation is used as an agent; based on the satellite network static topology model, an objective function of minimizing the end-to-end delay of a data packet is established; based on the satellite network static topology model, a satellite agent network is constructed, and satellite interaction experience data is obtained; wherein the satellite agent network includes a multi-satellite collaborative hybrid network and a satellite decision network corresponding to each agent, the multi-satellite collaborative hybrid network is used to calculate a joint reward according to the objective function, and the satellite decision network is used to output the routing decision of the agent according to the current local observation information of the agent; the satellite agent network is trained according to the satellite interaction experience data to obtain a trained satellite agent network; wherein each satellite decision network is used to calculate a local value function according to local observation information, the multi-satellite collaborative hybrid network is used to generate a global value function according to global state information and all local value functions, and the global value function and the joint reward are used to update network parameters; and the satellite decision networks in the trained satellite agent network are respectively deployed to corresponding agents, so that the agent makes a routing decision based on the deployed satellite decision network. In this way, during the training process, the satellite decision network and the multi-satellite collaborative hybrid network work together to enhance the decision-making ability and collaboration efficiency of satellite intelligent entities. The trained satellite decision network can make routing decisions that meet various business needs in a distributed manner based on real-time local observation information and business preferences, thereby solving the problems of dynamic changes in topology in low-orbit satellite networks, continuous growth in user demand, and uneven distribution of network traffic, ensuring the end-to-end transmission delay of the business and achieving effective load balancing between satellite nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A schematic diagram of a flow chart of a multi-satellite collaborative distributed routing method for multi-agent reinforcement learning provided by an embodiment of the present invention; Figure 2 A schematic diagram of the overall process provided by an embodiment of the present invention; Figure 3 A schematic diagram of a satellite routing scenario provided by an embodiment of the present invention; Figure 4A network structure diagram of a multi-agent distributed routing algorithm provided by an embodiment of the present invention; Figure 5 A diagram showing the working principle of a multi-agent distributed routing algorithm provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] In the large-scale low-orbit constellation routing optimization scenario, in response to the dynamic changes in satellite network topology, the continuous growth of business traffic, and the shortcomings of centralized routing strategies, satellite routing solutions have gradually turned to a distributed routing architecture based on virtual topology. This architecture uses time slicing technology to divide continuous time into multiple time slots. In each time slot, the network topology is considered to be stable and unchanged. It also uses logical topology and graph theory methods to build a static topology model to analyze the connectivity relationship between nodes and the optimal routing path. This distributed routing architecture based on virtual topology effectively reduces the interaction frequency between satellites, reduces computing requirements, and improves the scalability and flexible distributed deployment capabilities of the network.

[0020] As the field of artificial intelligence continues to develop, satellite routing algorithms combined with deep reinforcement learning (DRL) have shown strong adaptability in dealing with dynamically changing network states. These new routing schemes use the continuous interaction between satellite agents and the environment to dynamically learn and optimize routing decisions to adapt to the continuous changes in the satellite network environment, making them more flexible and performant than traditional routing strategies. Given the complexity of satellite routing tasks, multi-agent reinforcement learning technology provides new ideas and tools for designing low-orbit satellite network routing methods.

[0021] Based on this, an embodiment of the present invention provides a multi-satellite collaborative distributed routing method based on multi-agent reinforcement learning, which applies multi-agent reinforcement learning technology to traditional satellite routing design, which not only ensures the end-to-end transmission delay of the service, but also realizes effective load balancing between satellite nodes.

[0022] See also Figure 1 A schematic flow chart of a multi-star collaborative distributed routing method for multi-agent reinforcement learning is shown, and the method mainly includes the following steps S110 to S150: Step S110, obtaining network structure data of the low-orbit satellite constellation, and constructing a static topology model of the satellite network using time slicing technology; wherein each satellite in the low-orbit satellite constellation is regarded as an intelligent agent.

[0023] The network structure data of the Low Earth Orbit Satellite Constellation can include the number of satellite orbits, the number of satellites on each orbit, the orbital altitude, and two-line orbital data (TWO-Line Elementset, referred to as TLE). The static topology model of the satellite network is constructed using time slicing technology. Time slicing technology is used to discretize dynamic systems over a period of time for the convenience of analysis and modeling. The static topology model of the satellite network can include an undirected graph consisting of a set of satellite nodes and a set of links between satellites in each time slice.

[0024] Step S120: establishing an objective function for minimizing the end-to-end delay of data packets based on the static topology model of the satellite network.

[0025] In some possible embodiments, the above step S120 may include: based on the static topology model of the satellite network, using logical addresses and graph theory methods, establishing an inter-satellite link model and a data packet delay model; based on the inter-satellite link model and the data packet delay model, with the goal of minimizing the end-to-end delay of the data packet, expressing the satellite routing problem as a multi-objective optimization problem, and obtaining an objective function.

[0026] Among them, the logical address can be used to identify each node in the satellite network, so that each satellite has a unique identifier to facilitate communication and routing selection. Intersatellite link models can be divided into two categories: links between satellites in the same orbit and links between satellites in different orbits. For links between satellites in the same orbit, the connection between them is usually relatively stable, and the link is only disconnected when the distance between satellites exceeds the maximum communication range. For links between satellites in different orbits, in addition to considering the communication distance, the relative movement speed between satellites must also be considered. Especially when satellites are located in high latitudes or in the reverse seam of the orbit, the high-speed movement between satellites will make it difficult to establish a stable communication link, and these links will usually be in a disconnected state.

[0027] The packet delay model can include three parts: transmission delay, queuing delay and propagation delay. Transmission delay is defined as the time required for a data packet to enter the transmission medium (i.e., channel), which is related to the length of the data packet and the transmission rate. Queuing delay is the time a data packet waits for transmission in the satellite node cache. Propagation delay is the time required for a data packet to be transmitted from the current satellite to the next hop satellite, which is mainly determined by the inter-satellite distance and the signal propagation speed.

[0028] Alternatively, the above objective function can be expressed as: ; ; Where K represents the set of data packets. Indicates the number of packets, D( k ) indicates a data packet k On the path P k The end-to-end delay on , V k Indicates data packet k The set of nodes passed, E k Indicates data packet k The set of edges along the way, V t represents the set of satellite nodes, E t represents the set of links between satellites, H Indicates the maximum lifetime of a data packet, dis i,j Indicates satellite i With satellite j The distance between Indicates satellite i With satellite j The maximum viewing distance between Indicates data packet k exist t Is the satellite in the moment? i In the queue, Indicates data packet k exist t Is the satellite in the moment? g , i Transmission between Z t,i Indicates satellite i exist t The neighbor set at time, Indicates satellite i The maximum capacity of the queue, Indicates data packet k exist t Is the satellite in the moment? i was discarded.

[0029] In the above formula P Indicates the goal, st Represents constraints. Constraint C1 stipulates that the lifetime of a data packet in the network must not exceed its maximum lifetime (Time to Live, referred to as TTL). Constraint C2 indicates that the routing path selected by the satellite must be in communication status at the current moment. Constraint C3 requires that at the moment tThe sum of the number of packets buffered in the satellite's queue and the number of packets received from neighboring satellites at the next moment must not exceed the maximum capacity of the satellite queue. Constraint C4 indicates that the packet remains in one of the following states during the entire forwarding process: either in the satellite's queue or discarded during the routing process.

[0030] Step S130, constructing a satellite intelligent agent network based on the static topology model of the satellite network, and obtaining satellite interaction experience data; wherein the satellite intelligent agent network includes a multi-satellite collaborative hybrid network and a satellite decision network corresponding to each intelligent agent, the multi-satellite collaborative hybrid network is used to calculate the joint reward according to the objective function, and the satellite decision network is used to output the routing decision of the intelligent agent according to the current local observation information of the intelligent agent.

[0031] Building a multi-satellite collaborative hybrid network and Satellite Decision Network ,in, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of s t express t The global state at the moment, a t express t The joint action of the moment, Indicates satellite i The corresponding satellite decision network is in state , network parameters The value of Indicates satellite i exist t The local observation at the moment, the decision action at the previous moment, and the one-hot ID encoding of the agent, a t,i Indicates satellite i exist t Each satellite is regarded as an intelligent agent, and a satellite decision network is deployed on each intelligent agent. , these agents collaborate through a multi-star hybrid network Coordinate with each other to find the global optimal solution in the routing decision process. Among them, the multi-satellite cooperative hybrid network is deployed on the central node of each intelligent agent.

[0032] Local observation information can include tThe data packet sending capacity between the neighboring satellites connected within the time, the propagation delay between the current satellite and the neighboring satellite, the queue occupancy rate of the current satellite and the neighboring satellite, and the number of hops of the transmission data packet from the current satellite and the neighboring satellite to the destination satellite. The joint reward can be calculated based on the local observation information, decision actions and objective functions, where the decision actions cover the influence of queuing delay and propagation delay.

[0033] In some possible embodiments, the above-mentioned step of obtaining satellite interaction experience data may include: constructing a satellite agent interaction environment and an experience replay pool, and initializing network parameters of the satellite agent network; obtaining satellite interaction experience data generated by each agent through the interaction between the initialized satellite agent network and the satellite agent interaction environment, and storing the satellite interaction experience data in the experience replay pool; wherein each agent obtains local observation information from the satellite agent interaction environment, and makes routing decisions through the initialized corresponding satellite decision network, the central node in each agent obtains the local observation information and decision actions of each agent, and calculates the joint reward through the initialized multi-satellite collaborative hybrid network.

[0034] The above-mentioned satellite interaction experience data may include multiple data samples in different time slices, each of which includes global state information, joint local observation information, joint actions and joint rewards. The satellite interaction experience data in the experience replay pool will be extracted by the central node later for the training of the multi-satellite collaborative hybrid network and the satellite decision network. The network can be trained online or offline. During online training, each agent first deploys the initial satellite decision network. After the central node calculates the target loss function of the hybrid network, it is sent to each agent, and each agent updates the network parameters according to the target loss function.

[0035] Step S140, training the satellite agent network according to the satellite interaction experience data to obtain the trained satellite agent network; wherein each satellite decision network is used to calculate the local value function according to the local observation information, and the multi-satellite collaborative hybrid network is used to generate the global value function according to the global state information and all the local value functions, and the global value function and the joint reward are used to update the network parameters.

[0036] In some possible embodiments, the satellite interaction experience data is stored in an experience replay pool; the step of training the satellite agent network based on the satellite interaction experience data to obtain the trained satellite agent network may include: performing non-uniform sampling in the experience replay pool to obtain target sample data; obtaining a global value function through the satellite agent network based on the target sample data; and updating the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network through a back propagation algorithm based on the joint reward in the global value function and the target sample data.

[0037] Optionally, the sampling probability of each data sample in the above satellite interaction experience data is p t It can be as follows: ; ; in, ε Indicates the preset value. ε To prevent the sampling probability from being zero, Indicates status Under this condition, the error between the predicted value and the true value of the multi-satellite collaborative hybrid network is s t express t The global state at the moment, a t express t The joint action of the moment, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of Indicates that the multi-satellite cooperative hybrid network is in state ( ), network parameters The value of r t express t Joint rewards of the moment, represents the preset discount factor, and A represents the action space of the satellite.

[0038] In order to reduce the bias propagation caused by bootstrapping and the overestimation of target values ​​caused by maximization operations, target networks with the same network structure are constructed for the multi-satellite collaborative hybrid network and the satellite decision network, that is, the above-mentioned satellite agent network also includes a target hybrid network and a target decision network corresponding to the multi-satellite collaborative hybrid network and the satellite decision network, respectively; at predetermined intervals, the network parameters of the target hybrid network are adjusted to be the same as the network parameters of the multi-satellite collaborative hybrid network, and the network parameters of the target decision network are adjusted to be the same as the network parameters of the corresponding satellite decision network.

[0039] Based on this, the above-mentioned step of updating the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network through the back propagation algorithm according to the global value function and the joint reward in the target sample data may include: obtaining the target value function through the target hybrid network according to the joint reward in the target sample data; calculating the target loss function according to the global value function and the target value function; and updating the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each satellite decision network according to the target loss function using the gradient descent method.

[0040] Optionally, the above objective value function The definition of can be: ; The definition of the objective loss function can be: ; Network parameters of multi-satellite cooperative hybrid network and the network parameters of the satellite decision network The update formula can be: ; ; in, r t express t Joint rewards of the moment, represents the preset discount factor, Indicates that the target hybrid network is in state ( ), network parameters The value of s t+1 express t+ The global state at time 1, Expressed as t+ 1 moment is used to maximize the joint action of the local value function, Represents the network parameters of a multi-satellite collaborative hybrid network and satellite i The corresponding network parameters of the satellite decision network The objective loss function under b represents the target sample data sampled from the experience replay pool, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of α t Represents a preset learning rate that decreases over time.

[0041] Step S150, deploying the satellite decision networks in the trained satellite agent network to corresponding agents respectively, so that the agents make routing decisions based on the deployed satellite decision networks.

[0042] The trained satellite decision network is deployed to the corresponding satellite agent, and the satellite agent can make routing decisions that meet different business needs based on real-time local observation information and business preference information. The embodiment of the present invention adopts a normalized joint reward scheme that takes into account end-to-end distance, queuing delay and load balancing, allowing the satellite agent to heuristically adjust the reward weight according to business needs. Among them, the business preference information can be a preference for a specific aspect. For example, if the preference is to minimize the propagation delay, the reward weight of the propagation delay in the routing decision can be increased. For example, if the preference is for the overall delay, the reward weight of the overall delay in the routing decision can be increased.

[0043] The multi-satellite collaborative distributed routing method of multi-agent reinforcement learning provided by the embodiment of the present invention can enhance the decision-making ability and collaboration efficiency of satellite agents through the collaborative work of the satellite decision network and the multi-satellite collaborative hybrid network during the training process. The trained satellite decision network can make routing decisions that meet various business needs in a distributed manner based on real-time local observation information and business preferences, thereby solving the problems of dynamic changes in topology in low-orbit satellite networks, continuous growth in user demand, and uneven distribution of network traffic, ensuring the end-to-end transmission delay of the business, and realizing effective load balancing between satellite nodes.

[0044] For ease of understanding, the multi-star collaborative distributed routing method for the above-mentioned multi-agent reinforcement learning is introduced in detail below.

[0045] In view of the dynamic changes in topology, the continuous growth of user demand and the uneven distribution of network traffic in low-orbit satellite networks, the embodiment of the present invention designs a multi-satellite collaborative distributed routing method based on multi-agent reinforcement learning. Specifically, it includes: according to the periodic motion characteristics of the satellite, the satellite topology is divided into time slots using time slicing technology; combining logical addresses and graph theory methods to establish inter-satellite link models and data packet delay models, thereby converting the routing problem into a multi-objective optimization problem; deploying an agent for each satellite, and coordinating routing decisions through a multi-satellite collaborative hybrid network to enhance the collaboration between satellite agents; implementing a normalized joint reward scheme that takes into account end-to-end distance, queuing delay and load balancing, allowing satellite agents to heuristically adjust reward weights according to business needs. By integrating multi-agent reinforcement learning technology into traditional satellite routing design, not only the end-to-end transmission delay of the service is guaranteed, but also effective load balancing between satellite nodes is achieved.

[0046] In the large-scale low-orbit constellation routing optimization scenario, in view of the dynamic changes in satellite network topology, the continuous growth of business traffic and the shortcomings of centralized routing strategies, the embodiment of the present invention designs a multi-satellite collaborative distributed routing algorithm based on multi-agent reinforcement learning to solve the problem of end-to-end path selection for satellite routing. Experimental results show that the algorithm can effectively coordinate routing decision conflicts between satellite agents, ensure the end-to-end transmission delay of the business, and achieve load balancing between satellite nodes. The specific innovations of this solution are as follows: A. A multi-agent Transformer-MIX routing algorithm is proposed, which combines the satellite observation attention mechanism to capture the latent information in the satellite sequence data, and generates a joint action-value function through a Transformer-based parameter loop mechanism to achieve a more stable training process.

[0047] B. Adopt a centralized training and distributed execution strategy. In the training phase, by introducing a multi-satellite collaborative hybrid network, the possible conflicts between local optimality and global optimality among agents are coordinated to promote cooperation among agents. In the execution phase, satellites can rely only on their own observation space to complete routing decisions, greatly reducing the communication overhead between satellites.

[0048] C. A normalized joint reward scheme is designed which comprehensively considers end-to-end distance, queuing delay and load balance, so that the satellite agent can heuristically adjust the weights according to business needs to meet different business needs.

[0049] According to the above ideas and innovations, the distributed routing method proposed in the embodiment of the present invention includes the following steps: Step 1: Determine the network structure of the low-orbit satellite constellation, including the number of satellite orbits, the number of satellites on each orbit, the orbital altitude, and two lines of orbital data, and use time slicing technology to build a static topology model of the satellite network.

[0050] Step 2: Using logical addresses and graph theory methods, establish the intersatellite link model and data packet delay model, minimize the end-to-end delay of data packets, and express the satellite routing problem as a multi-objective optimization problem.

[0051] Among them, inter-satellite link models can be divided into two categories: links between satellites in the same orbit and links between satellites in different orbits. For links between satellites in the same orbit, the connection between them is usually relatively stable, and the link is only disconnected when the distance between satellites exceeds the maximum communication range. For links between satellites in different orbits, in addition to considering the communication distance, the relative movement speed between satellites must also be considered. Especially when satellites are located in high latitudes or in reverse orbital gaps, the high-speed movement between satellites makes it difficult to establish stable communication links, and these links are usually disconnected.

[0052] The packet delay model can include three parts: transmission delay, queuing delay, and propagation delay. Transmission delay is defined as the time required for a packet to enter the transmission medium, which is related to the packet length and transmission rate. Queuing delay is the time a packet waits for transmission in the satellite node cache. Propagation delay is the time required for a packet to be transmitted from the current satellite to the next hop satellite, which is mainly determined by the inter-satellite distance and signal propagation speed.

[0053] Step 3: Build a multi-satellite collaborative hybrid network and Satellite Decision Network In order to reduce the bias propagation caused by bootstrapping and the overestimation of the target value caused by maximization estimation, the target network is introduced into the multi-satellite cooperative hybrid network and the satellite decision network respectively. , The parameters of the original network are updated to its target network at fixed intervals to stabilize the training process and avoid continuous changes in the target update.

[0054] Among them, the satellite decision network takes the local observation information of the current satellite agent and the decision action of the previous moment (the decision action refers to which neighboring satellite the data packet is sent to) as input, generates the routing decision of the current satellite through the observation attention mechanism and the gated recurrent unit (GRU), and calculates the local value function. The multi-satellite collaborative hybrid network takes the local value function and global state information of all satellite decision networks as input, and outputs the global value function, which is used to evaluate the joint decision-making of all agents.

[0055] Step 4: Build the satellite agent interaction environment and experience playback pool. The satellite agent generates a six-tuple of experience data by observing network information and selecting routing decisions based on the ε-greedy strategy. , respectively represent the current (i.e. t moment), the global state information, joint local observation information, joint actions and joint rewards, and the next moment (i.e. t +1 moment) and the joint local observation information. Among them, the ε-greedy strategy is an action selection method used in reinforcement learning to balance exploration and exploitation.

[0056] Step 5: Centralized training of routing algorithms. First, randomly select training samples from the experience playback pool. Then, the satellite decision network uses the observed information to train the routing algorithm. o t Calculate the local value function. At the same time, the multi-satellite collaborative hybrid network uses the global state information s tAnd all local value functions are combined to generate a global value function. Then, the local value function and the global value function are used to execute the back propagation algorithm to update the network parameters until the algorithm parameters converge.

[0057] Among them, both the satellite decision network and the multi-satellite collaborative hybrid network have a replica target network to solve the problem of overestimation of the target value caused by maximization. The central node randomly extracts samples from the experience replay pool to calculate the loss function, which is specifically calculated as the global value function output by the satellite decision network at the current moment The difference between the expected value function at the next moment predicted by the target network can be expressed as: ; ; The central node uses this loss function to update the network parameters through the gradient descent method. In addition, every fixed training cycle, the parameters of the target network Will be replaced by the parameters of the current network , to stabilize the training process.

[0058] Step 6: Deploy the trained satellite decision network to the satellite agent, which can make routing decisions that meet different business needs based on real-time local observation information and business preference information.

[0059] See also Figure 2 The overall process diagram shown in the figure shows that the multi-satellite collaborative distributed routing method of multi-agent reinforcement learning mainly includes the following processes: satellite static topology modeling, multi-objective optimization problem representation, building a satellite agent network, centralized training, agent network parameter update and distributed routing strategy deployment, among which centralized training includes initializing the satellite agent environment and collecting satellite interaction experience data. The specific steps are as follows: Step 1: Construct a satellite static topology model, including steps 101 and 102.

[0060] Step 101: When considering a low-orbit satellite routing scenario, Figure 3 As shown, first of all, the period is M The topological state of the low-orbit satellite network (all satellites have the same orbital height and period) is discretized in the time dimension. Specifically, the length of each time slot is , thus obtaining discrete time slices In any time slice , the network topology is considered to be stable and unchanged, then the satellite network can be modeled as an undirected graph ,in V t represents the set of satellite nodes, i.e. ,and Et represents the set of links between satellites and is defined as .

[0061] Step 102: Further optimize the intersatellite link configuration. In each time slice, the intersatellite link adopts a mesh topology structure. Each satellite establishes communication links with up to four adjacent satellites, including two intra-plane links and two inter-plane links. The satellite's TLE data and other satellite network information (such as the number of satellite orbits, the number of satellites on each orbit, and the orbital altitude) are used to calculate the satellite's spatial coordinates in each time slice, and these coordinates are used to construct intersatellite links between adjacent satellites. Specifically, if the distance between two satellites meets the conditions for establishing a communication link and they are not located in a high-latitude area, there is an effective intersatellite communication link between the two satellites; conversely, if these conditions are not met, there will be no communication link between the two satellites.

[0062] Step 2: Multi-objective optimization problem representation, including steps 201 to 203.

[0063] Step 201: Establish an inter-satellite link model. Inter-satellite link models are mainly divided into two categories. The first category is the link between satellites in the same orbit. These links are usually relatively stable and are only disconnected when the distance between satellites exceeds the maximum communication distance. Assume that satellites in the same orbit i With satellite j The distance between them can be calculated by the following formula: ; in, is the relative angle between the two satellites. In addition, the establishment of inter-satellite links in the same orbit must also meet the condition of not being blocked by the earth: ; in, is the radius of the Earth, is the orbital radius of the satellite.

[0064] The second type is satellite communication links between different orbits. Since the satellites are located in different orbital planes, the angles between the two satellites and the center of the earth are It can be calculated by the following formula: ; in, , Satellite i The longitude and latitude of , Indicates satellite j longitude and latitude.

[0065] Therefore, the distance between satellites It can be expressed as: .

[0066] When the line between the two satellites is tangent to the Earth's surface, the maximum visible distance between them is reached. If the distance continues to increase, the inter-orbit communication link cannot be established due to the blockage of the earth. Therefore, the distance between satellites is dis i,j Must be less than the maximum viewing distance: .

[0067] Step 202: Establish a data packet delay model, which includes transmission delay, queuing delay and propagation delay. In the satellite network, define the data packet set as K. For each data packet , whose forwarding path from the source satellite to the destination is expressed as , where V k is the set of nodes passed, E k is the set of edges along the way. k On the path P k The end-to-end delay on can be expressed as: .

[0068] Transmission delay It is the time required for a data packet to completely enter the transmission medium, which depends on the packet length pkt and the data transmission rate .in, is the channel bandwidth of the intersatellite link between satellites, is the signal-to-noise ratio between nodes. The data transmission rate can be defined by the following formula: ; Therefore, the transmission delay can be defined as: .

[0069] Queuing delay It is the time it takes for a data packet to be cached in a satellite node and wait to be sent again. Since a satellite node in the network can establish intersatellite links with up to four adjacent satellites within the communication range, when a data packet arrives at the satellite, it often takes a while to be forwarded to the next hop satellite, and the length of the waiting time is related to the number of data packets in the current satellite node's sending queue. Therefore, the queuing delay can be expressed as: ; in, is the update interval of the satellite routing table, Represented as a satellite i Queue at time tThe number of packets in a queue can be described by considering the dynamics of packets arriving and leaving the queue: .

[0070] Propagation Delay refers to the data packet k Via communication link The time required for the communication link to be dis i,j and speed of information dissemination c Therefore, the propagation delay can be defined as: .

[0071] Step 203: Multi-objective optimization problem representation.

[0072] Based on the above model, a joint optimization problem can be established to minimize the end-to-end delay D( k ), expressed as: ; ; Among them, constraint C1 stipulates that the lifetime of the data packet in the network must not exceed its maximum lifetime. Constraint C2 means that the routing path selected by the satellite must be in communication state at the current time. Constraint C3 requires that t The sum of the number of packets buffered in the satellite's queue and the number of packets received from neighboring satellites at the next moment must not exceed the maximum capacity of the satellite queue. Constraint C4 indicates that the packet remains in one of the following states during the entire forwarding process: either in the satellite's queue or discarded during the routing process.

[0073] Step 3: Build a multi-satellite collaborative hybrid network and a satellite decision network. Each satellite is regarded as an intelligent agent, and a satellite decision network is deployed on each intelligent agent. These agents collaborate through a multi-star hybrid network Coordinate with each other to find the global optimal solution in the routing decision process. It includes steps 301 and 302.

[0074] Step 301: The satellite decision network receives local observation information of each satellite agent and the decision action at the previous moment As input. Through the observation attention mechanism and GRU processing, the network outputs the routing decision of the current satellite. The multi-satellite collaborative hybrid network receives the output local value function and global state information of all satellite decision networks. s t , and output the global value function , which is used to evaluate the joint decision-making effect of all agents.

[0075] Step 302: In order to reduce the bias propagation caused by bootstrapping and the overestimation of the target value caused by the maximization operation, target networks with the same network structure are constructed for the multi-satellite cooperative hybrid network and the satellite decision network. and By updating the parameters of the original network to the target network at fixed intervals, the stability of the target during training can be ensured, avoiding the problem of the target constantly changing during the update process.

[0076] Step 4: Build the satellite agent interaction environment and experience playback pool, and initialize network parameters. Each agent obtains its own local network information from the environment. , and make routing decisions based on local observations. Then, the joint decisions of all agents are executed uniformly and the joint rewards are calculated. At each time slot, the global state, local observations, joint actions, and joint rewards are stored as experience in the experience replay pool. This data is then extracted by the central node for network training.

[0077] Step 5: The routing algorithm is centrally trained. The central node calculates the loss function by extracting empirical data and uses the gradient descent method to update the network parameters, including steps 501 to 502.

[0078] Step 501: Prioritize experience playback. Define a probability of being sampled for each data sample. p t , allowing the central node to perform non-uniform sampling in the experience replay pool, giving priority to those experiences with higher expected learning value. The sampling probability p t It can be expressed as: ; in, ε is a very small number used to prevent the sampling probability from approaching zero and ensure that all samples can be drawn with non-zero probability. Is the current state Under this condition, the error between the value predicted by the multi-satellite collaborative hybrid network and the true value is defined as: ; in, A The satellite’s decision space Step 502: Update the agent network parameters. The central node randomly selects data based on the sampling probability of the experience sample to calculate the loss function, and optimizes the target loss function and updates the network weights through the gradient descent method. The target loss function is defined as: ; in, bRepresents training data sampled from the experience replay pool. The target value function calculated for the target network of the multi-satellite collaborative hybrid network is defined as: ; in, is the discount factor for future rewards, ranging from 0 to 1; Expressed as t+ 1 moment is used to maximize the joint action of the local value function. Specifically, for each satellite agent i Local action It can be expressed as: .

[0079] Finally, the central node uses the gradient descent method to update the network parameters according to the loss function. The specific update process can be expressed as: ; ; in, α t Indicates a preset learning rate that decreases over time, for example, between 0.01 and 0.005.

[0080] At the same time, after a fixed number of training rounds, the parameters of the target network Will be replaced by the original network parameters , in order to ensure the stability of training.

[0081] Step 6: Deploy the trained satellite decision network to the satellite agent, which can make routing decisions that meet different business needs based on real-time local observation information and business preference information.

[0082] See also Figure 4 The network structure diagram of the multi-agent distributed routing algorithm shown in Figure 2 is for the proposed satellite decision network ( Figure 4 In (a), each satellite agent i Maintaining the action-value function , where the input of the satellite decision network includes local observations , the decision action at the last moment and the one-hot ID encoding of the agent The proposed satellite decision network integrates an attention participation unit (SelfAtt) and a gated recurrent unit (GRU), which work in parallel to assist the agent in making routing decisions. The parallel results are then passed through two fully connected layers (MLP) and ε -greedy action sampling strategy obtains the satellite's decision action; among them, the hidden layer state , are the input and output of GRU respectively. For the proposed multi-star collaborative hybrid network ( Figure 4 In (c), the multi-satellite collaborative hybrid network uses a Transformer-based parameter recursive mechanism to solve the temporal correlation in the training data. This mechanism is used to combine the action value functions of all satellite agents into a joint value function. By using the Transformer module hybrid network, the temporal dependency and long-range correlation in the data can be effectively captured, so that the parameters of the joint function can be dynamically generated. are the weights generated by the hybrid network through the Transformer-based parameter recursive mechanism, TB represents the processing performed by the Transformer module, and Represents the global state processed by the fully connected layer s t The environmental feature vector of Corresponding to the time step t -1, and by the equation Update the weights, where relu Represents the use of relu activation function to process parameters, Indicates that the parameter is processed using absolute value operation.

[0083] See also Figure 5 The working principle diagram of the multi-agent distributed routing algorithm shown in the figure, during the training phase, each satellite agent maintains a queue to effectively manage data packets. By obtaining local observation information of neighboring satellites and itself, each agent uses the satellite decision network to independently make the next hop decision of the data packet, and then jointly executes the decision actions of all agents and calculates the joint reward. r t Second, the global state, action, joint reward, and observation are combined as , and at time t Stored in the experience replay pool. Finally, the multi-satellite assisted hybrid network samples these stored experience data to update its parameters, thereby coordinating and optimizing the decision-making process of all satellite agents.

[0084] In summary, the embodiment of the present invention proposes a multi-satellite collaborative distributed routing method based on multi-agent reinforcement learning, which aims to solve the end-to-end path selection problem of large-scale low-orbit constellation networks. In this method, the satellite decision network works with the multi-satellite collaborative hybrid network to enhance the decision-making ability and collaboration efficiency of the satellite agent. Each satellite decision network integrates the attention mechanism and the gated recurrent unit, so that the agent can not only effectively process and extract relevant features in the decision data, but also make more reasonable routing choices. The multi-satellite collaborative hybrid network adopts a Transformer-based parameter circulation mechanism to generate nonlinear parameters to fuse the local action value functions of all agents into a global value function, thereby achieving the stability of the training process. After the training is completed, these optimized satellite decision networks will be deployed to each satellite agent, enabling it to make routing decisions that meet various business needs in a distributed manner based on real-time local observation information and business preferences.

[0085] In all examples shown and described herein, any specific values ​​should be interpreted as merely exemplary and not as limiting, and thus other examples of the exemplary embodiments may have different values.

[0086] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-star collaborative distributed routing method for multi-agent reinforcement learning, characterized in that: include: Acquire network structure data of a low-orbit satellite constellation, and construct a static topology model of the satellite network using time slicing technology; wherein each satellite in the low-orbit satellite constellation is regarded as an intelligent agent; Based on the satellite network static topology model, an objective function for minimizing the end-to-end delay of data packets is established; A satellite agent network is constructed based on the satellite network static topology model, and satellite interaction experience data is obtained; wherein the satellite agent network includes a multi-satellite collaborative hybrid network and a satellite decision network corresponding to each of the agents, the multi-satellite collaborative hybrid network is used to calculate the joint reward according to the objective function, and the satellite decision network is used to output the routing decision of the agent according to the current local observation information of the agent; The satellite agent network is trained according to the satellite interaction experience data to obtain a trained satellite agent network; wherein each satellite decision network is used to calculate a local value function according to local observation information, the multi-satellite collaborative hybrid network is used to generate a global value function according to global state information and all the local value functions, and the global value function and the joint reward are used to update network parameters; The satellite decision networks in the trained satellite agent network are respectively deployed on the corresponding agents, so that the agents make routing decisions based on the deployed satellite decision networks.

2. The method according to claim 1, characterized in that The network structure data includes the number of satellite orbits, the number of satellites on each orbit, the orbit height and two rows of orbit data. The satellite network static topology model includes an undirected graph consisting of a satellite node set and a link set between satellites in each time slice.

3. The method according to claim 1, characterized in that The step of establishing an objective function for minimizing the end-to-end delay of a data packet based on the satellite network static topology model comprises: Based on the satellite network static topology model, an intersatellite link model and a data packet delay model are established by using logical addresses and graph theory methods; Based on the intersatellite link model and the data packet delay model, the satellite routing problem is expressed as a multi-objective optimization problem with the goal of minimizing the end-to-end delay of the data packet, and the objective function is obtained.

4. The method according to claim 3, characterized in that: The objective function is expressed as: ; ; Where K represents the set of data packets, Indicates the number of packets, D( k ) indicates a data packet k On the path P k The end-to-end delay on , V k Indicates data packet k The set of nodes passed, E k Indicates data packet k The set of edges along the way, V t represents the set of satellite nodes, E t represents the set of links between satellites, H Indicates the maximum lifetime of a data packet, dis i,j Indicates satellite i With satellite j The distance between Indicates satellite i With satellite j The maximum viewing distance between Indicates data packet k exist t Is the satellite in the moment? i In the queue, Indicates data packet k exist t Is the satellite in the moment? g , i Transmission between Z t,i Indicates satellite i exist t The neighbor set at time, Indicates satellite i The maximum capacity of the queue, Indicates data packet k exist t Is the satellite in the moment? i was discarded.

5. The method according to claim 1, characterized in that The steps to obtain satellite interaction experience data include: Constructing a satellite agent interaction environment and an experience replay pool, and initializing network parameters of the satellite agent network; Obtain satellite interaction experience data generated by each of the agents through the interaction of the initialized satellite agent network with the satellite agent interaction environment, and store the satellite interaction experience data in the experience replay pool; wherein each of the agents obtains local observation information from the satellite agent interaction environment, and makes routing decisions through the initialized corresponding satellite decision network, the central node in each of the agents obtains the local observation information and decision actions of each of the agents, and calculates the joint reward through the initialized multi-satellite collaborative hybrid network.

6. The method according to claim 5, characterized in that The satellite interaction experience data includes multiple data samples in different time slices, and each of the data samples includes global state information, joint local observation information, joint actions and joint rewards.

7. The method according to claim 1, characterized in that The satellite interaction experience data is stored in an experience playback pool; The step of training the satellite agent network according to the satellite interaction experience data to obtain the trained satellite agent network includes: Perform non-uniform sampling in the experience replay pool to obtain target sample data; According to the target sample data, obtaining the global value function through the satellite agent network; According to the global value function and the joint reward in the target sample data, the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each of the satellite decision networks are updated through a back propagation algorithm.

8. The method according to claim 7, characterized in that The sampling probability of each data sample in the satellite interaction experience data is p t as follows: ; ; in, ε Indicates the preset value. ε To prevent the sampling probability from being zero, Indicates status The error between the value predicted by the multi-satellite collaborative hybrid network and the true value is s t express t The global state at the moment, a t express t The joint action of the moment, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of Indicates that the multi-satellite cooperative hybrid network is in state ( ), network parameters The value of r t express t Joint rewards of the moment, represents the preset discount factor, and A represents the action space of the satellite.

9. The method according to claim 7, characterized in that: The satellite agent network also includes a target hybrid network and a target decision network corresponding to the multi-satellite collaborative hybrid network and the satellite decision network respectively; at predetermined intervals, the network parameters of the target hybrid network are adjusted to be the same as the network parameters of the multi-satellite collaborative hybrid network, and the network parameters of the target decision network are adjusted to be the same as the network parameters of the corresponding satellite decision network; according to the global value function and the joint reward in the target sample data, the steps of updating the network parameters of the multi-satellite collaborative hybrid network and the network parameters of each of the satellite decision networks by a back propagation algorithm include: According to the joint reward in the target sample data, obtaining a target value function through the target hybrid network; Calculate a target loss function based on the global value function and the target value function; According to the target loss function, the network parameters of the multi-satellite cooperative hybrid network and the network parameters of each satellite decision network are updated using the gradient descent method.

10. The method according to claim 9, characterized in that The target value function is defined as: ; The objective loss function is defined as: ; Network parameters of the multi-satellite cooperative hybrid network and network parameters of the satellite decision network The update formula is: ; ; in, r t express t Joint rewards of the moment, represents the preset discount factor, Indicates that the target hybrid network is in state ( ), network parameters The value of s t+1 express t+ The global state at time 1, Expressed as t+ 1 moment is used to maximize the joint action of the local value function, Represents the network parameters of the multi-satellite cooperative hybrid network and satellite i The corresponding network parameters of the satellite decision network The objective loss function under b represents the target sample data sampled from the experience replay pool, Indicates that the multi-satellite cooperative hybrid network is in state , network parameters The value of α t Represents a preset learning rate that decreases over time.

Citation Information

Patent Citations

  • Low earth orbit satellite network trusted load balancing routing method, system, device and medium

    CN116390164A

  • Low earth orbit satellite network flow routing method based on multi-agent reinforcement learning

    CN117041129A

  • Large-scale low-orbit constellation low-delay distributed intelligent routing algorithm

    CN117560053A

  • Multi-agent-based low-orbit satellite network routing decision-making method and device

    CN117614882A

  • Method, device, and computer readable storage medium for communication

    WO2024248648A1

Cited By

  • Satellite network distributed routing method, system and device and storage medium

    CN120474973A

  • Satellite Internet of Things distributed node parallel transmission acceleration method

    CN120601954A

  • Sectional satellite routing method and system based on multi-agent reinforcement learning

    CN121308819A

  • A segmented satellite routing method and system based on multi-agent reinforcement learning

    CN121308819B

  • Anti-interference waveform decision-making method based on multi-agent reinforcement learning

    CN121751204A