A method and system for balancing routing loads in a giant constellation satellite network
Through the multi-agent deep reinforcement learning model, combined with the giant constellation clustering mechanism and the cluster state compression mechanism, the inefficiency and congestion problems of the giant constellation satellite network in routing decision-making and load balancing are solved, distributed routing decision-making and load balancing are realized, and transmission delay and congestion probability are reduced.
Patent Information
- Application Number
- CN202210783945.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-05
AI Technical Summary
Giant constellation satellite networks have inefficiency and congestion problems in routing decision-making and load balancing, especially in the absence of global topological information.
The multi-agent deep reinforcement learning model is adopted to realize distributed routing decisions and load balancing by establishing a giant constellation clustering mechanism and a cluster state compression mechanism. This model includes an agent network and a hybrid network. The agent network is composed of a deep recursive Q network. The hybrid network is responsible for the collaboration between each agent and maximizes the global reward function.
It effectively reduces the delay and congestion probability of packet transmission, realizes load balancing and distributed routing decisions for giant constellation satellite networks, and extends the working life of the network.
Smart Images

Figure CN115297508B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a method and system for routing load balancing in a giant constellation satellite network. Background Art
[0002] With the advancement of low-cost small satellite platforms and advanced satellite communication equipment, giant constellation networks have developed greatly in recent years. The giant constellation network places tens of thousands of satellites in low-Earth orbit (LEO), which can provide low-latency broadband communications and global coverage for ground users and become an important supplement to the ground network. Some companies have already started plans for giant constellation networks, such as Starlink, OneWeb and Kuiper.
[0003] Although many routing algorithms have been proposed for low-orbit satellite networks, these routing algorithms are often very inefficient when applied to giant constellations due to the scale of satellites in giant constellations. Traditional routing algorithms, such as the Dijkstra-based routing algorithm, require centralized topology information collection for path calculation, which is often difficult to achieve in giant constellations. Traditional flooding-based routing protocols will cause large signaling overhead, especially in giant constellations. Distributed giant constellation routing algorithms that utilize the predictability and regularity of network topology can minimize the impact of link or node failures, and local detour mechanisms can avoid large signaling overheads, but due to the lack of global topology information, they will cause additional delays and congestion to data packets. Therefore, the routing problem of low-orbit constellation satellite networks must consider distributed routing mechanisms and congestion avoidance to reduce data packet transmission delays and congestion probabilities.
[0004] Reinforcement Learning (RL) is developed from theories such as animal learning and adaptive control. It emphasizes that agents gain experience in the process of interacting with the environment to make optimal strategy choices. Reinforcement learning does not rely on a complete environmental model in the learning process. It has a certain understanding of the environment and can automatically approach the optimal strategy. The agent analyzes the current state of the environment and the reward function for each action taken at previous moments. It selects actions based on the principle of maximum reward. The agent then calculates the reward value of the current action based on the feedback information from the environment and stores it to complete a learning cycle. The learning process of reinforcement learning is to complete the mapping of state space and action under the premise of maximizing rewards.
[0005] In order to maximize the global reward, multi-agent deep reinforcement learning was proposed, in which a group of agents coordinate their action strategies to maximize the global reward. Multi-agent reinforcement learning usually trains agents in a centralized manner in a simulated environment, where the joint agents can obtain global state information. Multi-agent deep reinforcement learning introduces a hybrid network that estimates the value of the joint behavior of the agents as a complex nonlinear combination, and the reward value of each individual depends only on the observation of the local environment. At the same time, the structure forces the joint action value of each agent to be monotonic, ensuring the consistency of the strategy during centralized training and distributed execution.
[0006] Multi-agent deep reinforcement learning can complete the strategy coordination among a group of agents to maximize the global reward value. Applying it to the giant constellation routing problem can solve the local observation problem of giant constellation routing and realize the collaborative work among satellites to reduce the transmission delay and congestion probability of data packets. Summary of the invention
[0007] The technical problem to be solved by the present invention is to provide a giant constellation satellite network routing load balancing method and system in view of the deficiencies in the above-mentioned prior art, solve the load balancing problem in the operation of the low-orbit giant constellation satellite network, and realize the distributed routing decision and congestion avoidance strategy of the low-orbit giant constellation.
[0008] The present invention adopts the following technical solutions:
[0009] A method for routing load balancing of a giant constellation satellite network comprises the following steps:
[0010] S1. Establish a giant satellite constellation network and generate the topology of the giant satellite constellation network;
[0011] S2. Establishing a giant constellation clustering mechanism and collecting intra-cluster information of the giant constellation clusters;
[0012] S3, establishing a cluster state compression mechanism, based on the cluster information collected in step S2, using an automatic encoder to compress the cluster load to obtain a feature vector, and using the feature vector to express the load information of each satellite node in the cluster;
[0013] S4, constructing a multi-agent deep reinforcement learning model based on the giant constellation satellite network topology established in step S1 and the state compression mechanism in step S3;
[0014] S5. The satellite node periodically sends a Hello message to the neighboring node to determine whether a link is established with the neighboring node;
[0015] S6, based on the connection information with the neighboring nodes obtained in step S5, the agent on the satellite node in the multi-agent deep reinforcement learning model constructed in step S4 makes the next hop decision according to the current observation space, generates experience and transmits it to the cluster head;
[0016] S7, each cluster head regularly collects the experience value and load information generated by each satellite in step S6, completes state compression according to the cluster state compression mechanism in step S3, and sends the experience value of each satellite node at each time and the compressed state feature vector to the ground control center;
[0017] S8. The ground control center completes the multi-agent deep reinforcement learning training based on the experience data and state vectors sent by each cluster head in step S7, and regularly completes the update of Eval-Net;
[0018] S9. After the update of Eval-Net in step S8 is completed, the ground control center sends the deep recursive Q network parameters to all satellite nodes, the satellite node agents complete the strategy update, and each satellite agent completes the routing decision based on the newly sent parameters.
[0019] Specifically, in step S1, in the giant constellation satellite network topology, each satellite is a topological node, and the inter-satellite link is a topological edge; the intra-orbit inter-satellite link does not change with time; and the inter-orbit inter-satellite link changes with the movement of the satellite.
[0020] Specifically, in step S2, the cluster information of the giant constellation clustering uses balanced clustering, including cluster heads and cluster members. The cluster head is responsible for collecting information of each satellite node in the cluster, including the task status of onboard data packet transmission and the remaining energy of the satellite. In the clustering mechanism, cluster heads exchange information with each other to complete information collection and routing strategy issuance, and each member in the cluster transmits information to the cluster head to obtain the latest routing strategy. After collecting the information in the cluster, each cluster head transmits it back to the control center to complete the training of the multi-agent deep reinforcement learning model. Each cluster periodically completes the re-selection of the cluster head. The cluster head re-selection mechanism is that the current cluster head is responsible for collecting the remaining energy information of each cluster member in the cluster, completing the cluster head election calculation, selecting the satellite node with the largest remaining working time of each satellite node as the new cluster head, and issuing information to all members in the cluster to complete the cluster head update.
[0021] Furthermore, the remaining working time T(i) of each satellite node is as follows:
[0022]
[0023] Among them, E r (i) is the remaining energy of each satellite node, E av (i) is the average residual energy of satellite nodes in the cluster, is the number of hops from the current node to the remaining satellite nodes, a i is the harmonic coefficient.
[0024] Specifically, in step S3, the autoencoder uses multi-layer compression, and each layer is connected by a fully connected layer, and finally an output vector is obtained through an activation function; when the autoencoder is trained, the input is a cluster load vector, and a compressed vector is obtained through a fully connected layer and an activation function, and then decompressed; the neural network used for decompression is completely symmetrical with the autoencoder to obtain a decoding vector, and a loss function is obtained based on the decoding vector and the original input vector, and back propagation is performed to correct the weights and biases; only encoding is performed during execution.
[0025] Specifically, in step S4, the multi-agent deep reinforcement learning model includes an agent network and a hybrid network. The agent network is composed of a deep recursive Q network. The deep recursive Q network is placed on the onboard agent to complete real-time routing decisions; the hybrid network is a super network responsible for the coordination between the agents. The hybrid network is placed on the ground station and is sent to each satellite node after completing the centralized training to achieve coordination of the transmission strategies of each satellite node.
[0026] Furthermore, the construction of the intelligent agent network is as follows:
[0027] Complete the mapping of various parameters of the agent network to the actual problem, including observation space o, action a, and reward r; observation space o is the transmission task of the current satellite node; action a is the next hop transmission direction of the task, including forward, backward, left, and right, corresponding to the four intersatellite links of the current satellite node; during execution, the input layer is the observation space o, and after passing through the fully connected layer, gated recurrent unit, and activation function in sequence, the output action a is generated, and the reward function r is generated, and then it goes to the next state o next .
[0028] Furthermore, the construction of a hybrid network is as follows:
[0029] Complete the input of the hybrid network and the mapping of the state space. The input of the hybrid network is the reward value r of each agent, the state space s is the global state information, the state information is mapped to the network load information, and the automatic encoder is used to complete the load compression of each satellite node in the cluster; after the network is executed, the comprehensive output and cost function are obtained. According to the cost function, back propagation is performed and the weights and biases of the deep recursive Q network of Target-Net and the super network are corrected; Eval-Net is responsible for real-time routing decisions, Target-Net is responsible for parameter updates, and network parameters are regularly updated to Eval-Net.
[0030] Specifically, in step S5, when the Hello message from the neighboring node cannot be received, the corresponding link is disconnected.
[0031] In a second aspect, an embodiment of the present invention provides a giant constellation satellite network routing load balancing system, including:
[0032] A topology generation module is used to establish a giant satellite constellation network and generate the topology of the giant satellite constellation network;
[0033] Clustering mechanism module, which establishes a giant constellation clustering mechanism and collects intra-cluster information of giant constellation clusters;
[0034] The state compression module establishes a cluster state compression mechanism. Based on the cluster information collected by the clustering mechanism module, it uses an automatic encoder to compress the cluster load and uses a feature vector to describe the load information of each satellite node in the cluster.
[0035] The model building module builds a multi-agent deep reinforcement learning model based on the giant constellation satellite network topology established by the topology generation module and the state compression mechanism of the compression state module;
[0036] Link judgment module: the satellite node periodically sends Hello messages to neighboring nodes to determine whether a link is established with the neighboring nodes;
[0037] The routing decision module, based on the connection information with neighboring nodes obtained by the link judgment module, uses the agent on the satellite node in the multi-agent deep reinforcement learning model built by the model construction module to make the next hop decision according to the current observation space, generate experience and transmit it to the cluster head;
[0038] Experience sending module: each cluster head regularly collects the experience value and load information generated by each satellite in the routing decision module, completes the state compression according to the cluster state compression mechanism of the model building module, and sends the experience value and compressed state vector of each satellite node at each time to the ground control center;
[0039] Network training module: The ground control center completes multi-agent deep reinforcement learning training based on the experience data and state vectors sent by each cluster head in the sending module, and regularly updates Eval-Net;
[0040] Network execution module, after the training module updates Eval-Net, the ground control center sends the deep recursive Q network parameters to all satellite nodes, the satellite node agent completes the strategy update, and each satellite agent completes the routing decision based on the newly sent parameters.
[0041] Compared with the prior art, the present invention has at least the following beneficial effects:
[0042] A method for routing load balancing of giant constellation satellite networks uses multi-agent deep reinforcement learning to balance the routing decision-making process of each satellite and minimize the global transmission delay. It uses a centralized training and distributed execution routing strategy to effectively reduce the computing power of on-board equipment and extend the working life of the network. It establishes a giant network clustering mechanism to complete network information collection and avoid the additional network overhead caused by the flooding mechanism. It proposes a cluster state compression mechanism to improve the network training speed and computing overhead by reducing the data dimension.
[0043] Furthermore, in the giant constellation satellite network topology, each satellite is a topological node and the inter-satellite links are topological edges; the intra-orbit inter-satellite links do not change with time; the inter-orbit inter-satellite links change with the movement of the satellite, and the network topology is simplified.
[0044] Furthermore, in the clustering mechanism, cluster heads exchange information with each other and issue routing decisions to avoid the additional network overhead caused by the flooding mechanism, effectively extending the network life; regular re-election of cluster heads can avoid the problem of excessive energy consumption of a single satellite caused by a single cluster head, effectively extending the satellite working time. Please supplement the description of the balanced clustering used in the cluster information of the giant constellation clustering according to the content of claim 3, including cluster heads and cluster members, and the cluster heads are responsible for collecting information of each satellite node in the cluster, including the task status of on-board data packet transmission and the remaining energy of the satellite; in the clustering mechanism, cluster heads exchange information with each other to complete information collection and routing strategy issuance, and each member in the cluster transmits information to the cluster head to obtain the latest routing strategy; each cluster head collects information within the cluster and then transmits it back to the control center to complete the training of the multi-agent deep reinforcement learning model; each cluster periodically completes the re-selection of cluster heads, and the cluster head re-selection mechanism is responsible for collecting the remaining energy information of each cluster member in the cluster, completing the cluster head election calculation, selecting the satellite node with the largest remaining working time of each satellite node as the new cluster head, and issuing information to all members in the cluster to complete the purpose or benefit of cluster head update setting, and give a principle analysis explanation.
[0045] Furthermore, the cluster head election mechanism selects the cluster head by comprehensively considering the remaining energy of satellites within the cluster and the number of hops to each satellite in the cluster, effectively balancing the energy consumption of satellites within the cluster and extending the satellite life.
[0046] Furthermore, the cluster state compression mechanism completes the compression of the load information within the cluster, reduces the dimension of the load information, and then reduces the dimension of the hypernetwork input parameters, effectively reducing the overhead of network training.
[0047] Furthermore, a distributed execution centralized training method is adopted. The on-board intelligent agent only deploys a deep recursive Q network to complete real-time routing decisions, and the ground control center completes the training of the deep recursive Q network and the hybrid network, effectively reducing on-board overhead.
[0048] Furthermore, the onboard intelligent agent completes real-time routing decisions based on the transmission tasks and surrounding link information to ensure the real-time nature of the routing process.
[0049] Furthermore, we establish Target-Net, only update its parameters during training, and periodically pass the parameters to Eval-Net to prevent the network from falling into the local optimum.
[0050] Furthermore, the satellite network uses the hello message mechanism to dynamically maintain links to surrounding satellites, ensuring the effectiveness of the next hop of the transmission task.
[0051] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0052] In summary, the present invention effectively reduces network congestion and end-to-end transmission delay; the routing strategy of distributed execution of centralized training effectively reduces the computing workload of on-board equipment and extends the working life of the network; the satellite network clustering mechanism completes network information collection and avoids the additional network overhead caused by the flooding mechanism; the cluster state compression mechanism improves the network training speed and computing overhead by reducing the data dimension.
[0053] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a diagram of the giant satellite constellation network structure of the present invention;
[0055] Figure 2 A topological diagram of a giant satellite constellation network of the present invention;
[0056] Figure 3 This is a schematic diagram of the multi-agent deep reinforcement learning architecture of the present invention;
[0057] Figure 4 This is a flow chart of the routing strategy based on the multi-agent deep reinforcement learning model of the present invention;
[0058] Figure 5 The following is a comparison chart of the delivery success probability of each routing method when the constellation scale is 6*6, 12*12 and 24*24 respectively. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0061] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0062] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship.
[0063] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0064] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0065] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0066] The present invention provides a method for routing load balancing of a giant constellation satellite network. The research object is a giant constellation low-orbit satellite network. Aiming at the load balancing problem in the operation process of the low-orbit giant constellation satellite network, the distributed routing decision and congestion avoidance strategy of the low-orbit giant constellation are realized. By generating a walker-delta constellation, a satellite network topology is established; satellite clusters are generated and cluster head election is completed; a multi-agent deep reinforcement learning model is established to complete the mapping of state space, observation space and action space; the on-board agent completes real-time routing decision based on observation information and generates experience values; the cluster head regularly collects satellite experience values and load information, completes load information compression and transmits it to a ground control center; the ground control center completes multi-agent deep reinforcement learning model training based on the experience values collected by the cluster head, regularly completes Eval-Net network update, and sends transmission strategies to each satellite; each satellite completes routing decision based on the sent strategy.
[0067] Satellite nodes are considered as intelligent agents, which complete space observations according to the mission status and make real-time routing decisions to forward data packets. The network status information is collected through giant constellation clustering and concentrated to the ground control center for training. After the ground control center completes the training, it sends the latest transmission strategy to each satellite. Each satellite node intelligent agent is only responsible for execution, which reduces the complexity of on-board processing. Through the collaboration between satellites, congestion avoidance of transmission tasks is completed and transmission delay is reduced.
[0068] See also Figure 4 The present invention provides a method for routing load balancing of a giant constellation satellite network, comprising the following steps:
[0069] S1. Establish a giant satellite constellation network and generate the topology of the giant satellite constellation network;
[0070] See also Figure 1 The application scenario of the present invention considers a regular walker-delta constellation, and each satellite is regarded as a communication node; according to the characteristics of giant constellation satellite equipment, each satellite only establishes inter-satellite links with four surrounding adjacent satellites, namely, intra-orbit inter-satellite links within two orbits and two inter-orbit inter-satellite links.
[0071] See also Figure 2 , each satellite is regarded as a topological node and the inter-satellite links are regarded as topological edges.
[0072] The in-orbit intersatellite links do not change over time and can exist permanently.
[0073] The inter-orbital and inter-satellite links gradually change with the movement of the satellite, requiring antenna tracking, and sometimes the inter-orbital and inter-satellite links will be shut down in time to save energy due to poor link quality.
[0074] S2. Establishing a giant constellation clustering mechanism and collecting intra-cluster information of the giant constellation clusters;
[0075] The clustering of giant constellations uses balanced clustering, that is, when the number of satellites in each cluster is equal, the clustering overhead can be significantly reduced.
[0076] The clustering of giant constellations includes cluster heads and cluster members. The cluster head is responsible for collecting information of satellite nodes in the cluster, including the status of onboard data packet transmission tasks and the remaining energy of the satellite. In the clustering mechanism, cluster heads need to exchange information with each other to complete information collection and routing strategy delivery, and each member in the cluster needs to transmit information with the cluster head to obtain the latest routing strategy, both of which will cause additional overhead.
[0077] Each cluster head collects information within the cluster and transmits it back to the control center to complete the training of the multi-agent deep reinforcement learning model.
[0078] Each cluster periodically completes the re-selection of the cluster head to avoid excessive occupation of the same device and extend the working life of the network.
[0079] The cluster head re-selection mechanism is that the current cluster head is responsible for collecting the residual energy information of each cluster member in the cluster and completing the cluster head election calculation.
[0080] The update formula used for re-election is:
[0081]
[0082] Among them, T(i) is the remaining working time of each satellite node, E r (i) is the remaining energy of each satellite node, E av (i) is the average residual energy of satellite nodes in the cluster, is the number of hops from the current node to the remaining satellite nodes, a i is the harmonic coefficient.
[0083] Select the satellite node with the largest T(i) as the new cluster head, send the information to all cluster members, and complete the cluster head update.
[0084] S3, cluster load compression;
[0085] When the load information of each satellite is used as the state information, the state space is too large. An automatic encoder is used to compress the cluster load, that is, the feature vector is used to describe the load information of each satellite node in the cluster.
[0086] The autoencoder uses multi-layer compression, and each layer is connected by a fully connected layer. Finally, the output vector is obtained through an activation function. Generally, a two-layer neural network can be used for data compression when the cluster size is relatively small.
[0087] When training the autoencoder, the input is the cluster load vector, which is compressed by the fully connected layer and activation function and then decompressed. Decompression is the inverse process of compression. The neural network used is completely symmetrical with the autoencoder to obtain the decoded vector. The loss function is obtained based on the decoded vector and the original input vector, and back propagation is performed to correct the weights and biases.
[0088] During execution, only encoding is performed without decoding or training operations, which reduces the computational complexity during execution.
[0089] S4. Build a multi-agent deep reinforcement learning model
[0090] Based on the satellite network topology established in step S1 and the state compression mechanism in step S3, a multi-agent deep reinforcement learning model is constructed.
[0091] See also Figure 3 , the multi-agent deep reinforcement learning model is divided into agent network and hybrid network. The agent network is composed of a deep recursive Q network to complete the action decision of each agent. The hybrid network is a super network responsible for the coordination between agents to maximize the global reward function. The deep recursive Q network is placed in the on-board agent to complete the real-time routing decision. The hybrid network is placed in the ground station and sent to each satellite node after completing the centralized training to achieve the coordination of the transmission strategy of each satellite node.
[0092] The process of establishing the agent network is as follows:
[0093] First, complete the mapping of various parameters of the agent network to the actual problem, including observation space o, action a, and reward r.
[0094] The observation space o is the transmission task of the current satellite node, o(t) = {p s , o s , p d , o d}. Among them, p s is the source satellite node orbit number of the current mission, o s The satellite number in the orbit of the source satellite node of the current mission. d is the target satellite node orbit number of the current mission, o d The satellite number in orbit of the target satellite node of the current mission.
[0095] Action a is the next-hop transmission direction of the task, including forward, backward, left, and right, corresponding to the four intersatellite links of the current satellite node.
[0096] The reward function r includes two parts: the transmission distance and the remaining energy, which is defined as:
[0097] r=ω 1 diff+ω 2E c
[0098]
[0099] Among them, diff is related to the number of hops from the current satellite node to the target node after the current action selection is executed. When the current satellite node is far away from the target node, the difference of diff in each direction is small, avoiding the congestion problem of the only path transmission under the shortest path. If the next hop transmission direction of the current node is far away from the target node, it is set to the penalty value -rp. E c is the remaining energy of the next-hop satellite node, ω 1 and ω 2 is a hyperparameter responsible for the trade-off between transmission delay and satellite network uptime.
[0100] During execution, the input layer is the observation space o, which passes through the fully connected layer, the gated recurrent unit, and the activation function in sequence to generate the output action a, generate the reward function r, and go to the next state o next .
[0101] The process of establishing a hybrid network is as follows:
[0102] First, complete the input and state space mapping of the hybrid network.
[0103] The input of the hybrid network is the reward value r of each agent.
[0104] The state space s is the global state information. To achieve the load balancing of the giant constellation satellite network, the state information is mapped to the network load information. The autoencoder is used to complete the load compression of each satellite node in the cluster to reduce the dimension of the state space.
[0105] After the network is executed, the comprehensive output y is obtained tot as follows:
[0106]
[0107] The cost function is:
[0108]
[0109] Based on the cost function, backpropagation is performed and the weights and biases of the deep recursive Q network of Target-Net and the hypernetwork are corrected.
[0110] Eval-Net is responsible for real-time routing decisions, while Target-Net is responsible for parameter updates. The network parameters are updated to Eval-Net regularly to avoid excessive correlation and overfitting problems caused by real-time updates.
[0111] S5. The satellite node periodically sends a Hello message to the neighbor node to determine whether a link is established with the neighbor node. If the Hello message from the neighbor node cannot be received, the corresponding link is disconnected and this link is not considered as the next hop selection;
[0112] S6, the satellite node on-board agent based on the current observation space o(t) = {p s , o s , p d , o d}, make the next hop decision, generate experience e(t) = {o(t), a(t), r(t)}, and transmit it to the cluster head;
[0113] S7, each cluster head regularly summarizes the experience e(t) of the satellite nodes in the cluster, completes state compression according to the automatic encoder in step S4, and sends the experience value of each satellite node at each time and the compressed state vector to the ground control center;
[0114] When the cluster head collects the experience e(t) within the cluster and performs state compression, it will cause additional overhead, which is mainly determined by the cluster type and cluster size. The additional overhead is the lowest when the cluster is balanced, but the cluster size also affects the information collection load within the cluster and the training overhead of the multi-agent model. The larger the cluster, the higher the information collection overhead within the cluster but the lower the model training overhead. The actual cluster size can be comprehensively weighed by the network equipment energy and ground equipment resources.
[0115] S8. The ground control center completes the multi-agent deep reinforcement learning training based on the experience data and state vector sent by the cluster head, and regularly updates Eval-Net.
[0116] S9. The ground control center sends the deep recursive Q network parameters to all satellite nodes. The satellite node agent completes the strategy update and completes the routing decision based on the sent strategy.
[0117] In yet another embodiment of the present invention, a giant constellation satellite network routing load balancing system is provided, which can be used to implement the above-mentioned giant constellation satellite network routing load balancing method. Specifically, the giant constellation satellite network routing load balancing system includes a topology generation module, a clustering mechanism module, a state compression module, a model building module, a link judgment module, a routing decision module, an experience sending module, a network training module and a network execution module.
[0118] Among them, the topology generation module establishes a giant satellite constellation network and generates the topology of the giant satellite constellation network;
[0119] Clustering mechanism module, which establishes a giant constellation clustering mechanism and collects intra-cluster information of giant constellation clusters;
[0120] The state compression module establishes a cluster state compression mechanism. Based on the cluster information collected by the clustering mechanism module, it uses an automatic encoder to compress the cluster load and uses a feature vector to describe the load information of each satellite node in the cluster.
[0121] The model building module builds a multi-agent deep reinforcement learning model based on the giant constellation satellite network topology established by the topology generation module and the state compression mechanism of the compression state module;
[0122] Link judgment module: the satellite node periodically sends Hello messages to neighboring nodes to determine whether a link is established with the neighboring nodes;
[0123] The routing decision module, based on the connection information with neighboring nodes obtained by the link judgment module, uses the agent on the satellite node in the multi-agent deep reinforcement learning model built by the model construction module to make the next hop decision according to the current observation space, generate experience and transmit it to the cluster head;
[0124] Experience sending module: each cluster head regularly collects the experience value and load information generated by each satellite in the routing decision module, completes the state compression according to the cluster state compression mechanism of the model building module, and sends the experience value and compressed state vector of each satellite node at each time to the ground control center;
[0125] Network training module: The ground control center completes multi-agent deep reinforcement learning training based on the experience data and state vectors sent by each cluster head in the sending module, and regularly updates Eval-Net;
[0126] Network execution module, after the training module updates Eval-Net, the ground control center sends the deep recursive Q network parameters to all satellite nodes, the satellite node agent completes the strategy update, and each satellite agent completes the routing decision based on the newly sent parameters.
[0127] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can usually be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0128] See also Figure 5, and a comparison of the delivery success probability of each routing method when the constellation size is 6*6, 12*12 and 24*24 is given. The load balancing method based on multi-agent deep reinforcement learning has an improvement of more than 50% in the delivery success probability compared with the single-agent reinforcement learning method. With the expansion of the constellation scale, direct routing based on numbering can no longer complete data transmission, and most of the data is congested. The routing load balancing model based on multi-agent deep reinforcement learning completes the coordination of each satellite transmission strategy, effectively improving the delivery success probability of network transmission tasks.
[0129] When the satellite network is actually running, each satellite node obtains the observation space o(t)={p s , o s , p d , o d}, make routing decisions in real time. After the data packet is transmitted to the next hop, it is independent of the current satellite node, and continues to acquire the observation space and make real-time routing decisions until it reaches the target node. The cluster head collects the experience value and transmits it to the ground control center for network training to complete the adaptation to the dynamic changes of the satellite network environment. Therefore, the giant constellation routing strategy based on multi-agent deep reinforcement learning of the present invention is a dynamic and intelligent path planning method.
[0130] In summary, the present invention provides a method and system for routing load balancing of a giant constellation satellite network, which has the following characteristics:
[0131] (1) A multi-agent deep reinforcement learning model was used to complete the distributed routing decision-making and load balancing of the giant satellite network, solving the huge overhead problem caused by the centralized management of the giant constellation routing strategy. At the same time, it achieved collaboration among satellite nodes, reduced the congestion probability of transmission tasks, and reduced the transmission delay of data packets.
[0132] (2) The strategy of centralized training and distributed execution makes full use of the ground control center with relatively abundant computing resources for training, and takes into account the limited resources on board the giant constellation satellite network. The onboard equipment is only responsible for execution and occupies fewer resources.
[0133] (3) The multi-agent deep reinforcement learning model has a strong ability to adapt to the environment. When the spatial links change drastically, the reinforcement learning model can adjust the strategy in time, complete the reconstruction of the network topology, and improve the transmission efficiency and stability.
[0134] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0136] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0138] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A method for routing load balancing in a giant constellation satellite network. It is characterized in that The following steps are involved: S1. Establish a giant satellite constellation network and generate the topology of the giant satellite constellation network; S2. Establish a giant constellation clustering mechanism to collect the intra-cluster information of the giant constellation clusters. The intra-cluster information of the giant constellation clusters uses balanced clustering, including cluster heads and cluster members. The cluster head is responsible for collecting information of each satellite node in the cluster, including the task status of on-board data packet transmission and the remaining energy of the satellite. In the clustering mechanism, cluster heads exchange information with each other to complete information collection and routing strategy delivery, and each member in the cluster transmits information with the cluster head to obtain the latest routing strategy. Each cluster head collects information within the cluster and transmits it back to the control center to complete the training of the multi-agent deep reinforcement learning model; Each cluster periodically completes the re-election of the cluster head. The cluster head re-election mechanism is responsible for the current cluster head to collect the remaining energy information of each cluster member within the cluster, complete the cluster head election calculation, select the satellite node with the longest remaining working time among all satellite nodes as the new cluster head, send information to all cluster members within the cluster, and complete the update of the cluster head. The remaining working time of each satellite node is as follows: in, is the remaining energy of each satellite node, is the average residual energy of satellite nodes in the cluster, is the number of hops from the current node to the remaining satellite nodes, is the harmonic coefficient; S3, based on the cluster information collected in step S2, use an automatic encoder to compress the cluster load to obtain a feature vector, and use the feature vector to express the load information of each satellite node in the cluster; S4. Based on the giant constellation satellite network topology established in step S1 and the state compression mechanism of step S3, a multi-agent deep reinforcement learning model is constructed. The multi-agent deep reinforcement learning model includes an agent network and a hybrid network. The agent network is composed of a deep recursive Q network. The deep recursive Q network is placed in the on-board agent to complete real-time routing decisions; the hybrid network is a super network responsible for the coordination between the agents. The hybrid network is placed in the ground station and sent to each satellite node after completing the centralized training to achieve the coordination of the transmission strategies of each satellite node. The construction of the hybrid network is specifically as follows: Complete the input of the hybrid network and the mapping of the state space. The input of the hybrid network is the reward value of each agent. , state space The global state information is mapped to the network load information, and the autoencoder is used to complete the load compression of each satellite node in the cluster; after the network is executed, the comprehensive output and cost function are obtained, and according to the cost function, the back propagation is performed and the weights and biases of the deep recursive Q network of Target-Net and the super network are corrected; Eval-Net is responsible for real-time routing decisions, and Target-Net is responsible for parameter updates, and the network parameters are regularly updated to Eval-Net; S5. The satellite node periodically sends a Hello message to the neighboring node to determine whether a link is established with the neighboring node; S6, based on the connection information with the neighboring nodes obtained in step S5, the agent on the satellite node in the multi-agent deep reinforcement learning model constructed in step S4 makes the next hop decision according to the current observation space, generates experience and transmits it to the cluster head; S7, each cluster head regularly collects the experience value and load information generated by each satellite in step S6, completes state compression according to step S3, and sends the experience value of each satellite node at each time and the compressed state feature vector to the ground control center; S8. The ground control center completes the multi-agent deep reinforcement learning training based on the experience data and state vectors sent by each cluster head in step S7, and regularly updates the Eval-Net responsible for real-time routing decision-making; S9. After the update of Eval-Net in step S8 is completed, the ground control center sends the deep recursive Q network parameters to all satellite nodes, the satellite node agents complete the strategy update, and each satellite agent completes the routing decision based on the newly sent parameters.
2. The method for routing load balancing of a giant constellation satellite network according to claim 1, It is characterized in that In step S1, in the giant constellation satellite network topology, each satellite is a topological node, and the inter-satellite link is a topological edge; the intra-orbit inter-satellite link does not change with time; and the inter-orbit inter-satellite link changes with the movement of the satellite.
3. The method for routing load balancing of a giant constellation satellite network according to claim 1, It is characterized in that In step S3, the autoencoder uses multi-layer compression, and each layer is connected by a fully connected layer, and finally the output vector is obtained through the activation function; when the autoencoder is trained, the input is the cluster load vector, and the compressed vector is obtained through the fully connected layer and the activation function, and then decompressed; the neural network used for decompression is completely symmetrical with the autoencoder to obtain a decoding vector, and the loss function is obtained based on the decoding vector and the original input vector, and back propagation is performed to correct the weights and biases; only encoding is performed during execution.
4. The method for routing load balancing of a giant constellation satellite network according to claim 1, It is characterized in that In step S4, constructing the agent network is specifically as follows: Complete the mapping of the parameters of the agent network to the actual problem, including the observation space ,action , Reward Value ; Observation space The transmission task of the current satellite node; action is the next hop transmission direction of the task, including front, back, left, and right, corresponding to the four intersatellite links of the current satellite node; during execution, the input layer is the observation space , after passing through the fully connected layer, gated recurrent unit, and activation function, the output action is generated , and generate a reward value , and go to the next state .
5. The method for routing load balancing of a giant constellation satellite network according to claim 1, It is characterized in that In step S5, when the Hello message from the neighboring node cannot be received, the corresponding link is disconnected.
6. A giant constellation satellite network routing load balancing system, It is characterized in that include: A topology generation module is used to establish a giant satellite constellation network and generate the topology of the giant satellite constellation network; Clustering mechanism module, establishes giant constellation clustering mechanism, collects cluster information of giant constellation clusters, and uses balanced clustering for cluster information, including cluster heads and cluster members. Cluster heads are responsible for collecting information of satellite nodes in the cluster, including the task status of on-board data packet transmission and the remaining energy of satellites. In the clustering mechanism, cluster heads exchange information with each other to complete information collection and routing strategy delivery, and cluster members transmit information with cluster heads to obtain the latest routing strategy. Each cluster head collects information within the cluster and transmits it back to the control center to complete the training of the multi-agent deep reinforcement learning model; Each cluster periodically completes the re-election of the cluster head. The cluster head re-election mechanism is that the current cluster head is responsible for collecting the remaining energy information of each cluster member in the cluster, completing the cluster head election calculation, selecting the satellite node with the largest remaining working time of each satellite node as the new cluster head, sending the information to all cluster members, completing the cluster head update, and the remaining working time of each satellite node as follows: in, is the remaining energy of each satellite node, is the average residual energy of satellite nodes in the cluster, is the number of hops from the current node to the remaining satellite nodes, is the harmonic coefficient; The state compression module uses an automatic encoder to compress cluster loads based on the cluster information collected by the clustering mechanism module, and uses feature vectors to express the load information of each satellite node in the cluster; The model building module builds a multi-agent deep reinforcement learning model based on the giant constellation satellite network topology established by the topology generation module and the state compression mechanism of the compression state module. The multi-agent deep reinforcement learning model includes an agent network and a hybrid network. The agent network is composed of a deep recursive Q network. The deep recursive Q network is placed on the on-board agent to complete real-time routing decisions; the hybrid network is a super network responsible for the coordination between the agents. The hybrid network is placed on the ground station and sent to each satellite node after completing the centralized training to achieve the coordination of the transmission strategies of each satellite node. The construction of the hybrid network is specifically as follows: Complete the input of the hybrid network and the mapping of the state space. The input of the hybrid network is the reward value of each agent. , state space The global state information is mapped to the network load information, and the autoencoder is used to complete the load compression of each satellite node in the cluster; after the network is executed, the comprehensive output and cost function are obtained, and according to the cost function, the back propagation is performed and the weights and biases of the deep recursive Q network of Target-Net and the super network are corrected; Eval-Net is responsible for real-time routing decisions, and Target-Net is responsible for parameter updates, and the network parameters are regularly updated to Eval-Net; Link judgment module: the satellite node periodically sends Hello messages to neighboring nodes to determine whether a link is established with the neighboring nodes; The routing decision module, based on the connection information with neighboring nodes obtained by the link judgment module, uses the agent on the satellite node in the multi-agent deep reinforcement learning model built by the model construction module to make the next hop decision according to the current observation space, generate experience and transmit it to the cluster head; Experience sending module: each cluster head regularly collects the experience value and load information generated by each satellite in the routing decision module, completes the state compression according to the cluster state compression mechanism of the model building module, and sends the experience value and compressed state vector of each satellite node at each time to the ground control center; Network training module: The ground control center completes multi-agent deep reinforcement learning training based on the experience data and state vectors sent by each cluster head in the sending module, and regularly updates Eval-Net; Network execution module, after the training module updates Eval-Net, the ground control center sends the deep recursive Q network parameters to all satellite nodes, the satellite node agent completes the strategy update, and each satellite agent completes the routing decision based on the newly sent parameters.
Citation Information
Patent Citations
Low-orbit satellite routing strategy method based on deep reinforcement learning architecture
CN110012516A
Hybrid routing method based on clustering and reinforcement learning and ocean communication system
CN111510956A