Metareinforcement learning-based satellite-ground convergence network routing method and system

By applying meta-reinforcement learning methods in the satellite-earth fusion network, a multi-task experience pool and neural network structure is constructed, which solves the problems of slow training speed of routing strategy and poor adaptability in multiple environments in the existing technology, and achieves more efficient and intelligent routing decisions.

CN120034931APending Publication Date: 2025-05-23SHANGHAI SPACEFLIGHT INST OF TT&C & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233982.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art faces the problems of slow training speed, poor multi-environment adaptability and degraded routing decision performance when dealing with optimal routing strategies for satellite-ground converged networks.

Method used

Using a routing method based on meta-reinforcement learning, we use the online neural network and target neural network to conduct meta-reinforcement learning training to realize intelligent routing strategies by building a multi-task experience pool of the star-ground fusion network, defining evaluation indicators, and building state space, action space and reward functions.

Benefits of technology

It improves the performance and resource utilization of the computing network, realizes a more efficient and intelligent routing strategy, and is suitable for space-based, air-based and ground-based multitasking environments in the satellite-ground fusion network, making up for the limitations of existing routing algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034931A_ABST
    Figure CN120034931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of satellite network communication, and provides a satellite-ground convergence network routing method and system based on meta reinforcement learning, and the method comprises the steps: constructing and generating a network topology structure through employing an STK tool kit; creating a multi-task experience pool for storing training samples; defining an evaluation index for determining an optimal route, and constructing a state space, an action space and a reward function of the satellite-ground fusion network according to a Markov decision process; determining a current state Q value by the online neural network, updating the state Q value by a Bellman equation in a target neural network, training by adopting an experience playback and greedy search method, and selecting a corresponding agent action; the updating frequency and the updating step number are initialized, and the optimal routing strategy is achieved through parallel training of multiple network structures. The method is suitable for a multi-task application environment in a satellite-ground fusion network, an optimal routing strategy is realized in a new network environment through less training, network resources are reasonably utilized, and the problems of low training speed and multi-environment adaptability of an existing reinforcement learning method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite network communication technology, and in particular to a satellite-ground fusion network routing method and system based on meta-reinforcement learning. Background Art

[0002] Building an integrated space-ground network is an important direction for the development of future wireless communication systems. It deeply combines space-based networks, air-based networks, and ground-based networks, gives full play to the dimensional advantages of different networks, provides high-bandwidth, large-connection global consistent communication services, and realizes wide-area full-coverage interconnection. Among them, the satellite-ground integrated fusion networking routing technology provides the optimal path for the transmission of data packets in the network, which is one of the core functions of the integrated space-ground communication network. Reasonable routing decisions can make more reasonable use of network resources and improve network performance. The satellite-ground fusion network has complex topology, large transmission delay, and high deployment cost, which seriously restricts the performance of the space-ground integrated network. Therefore, the main challenge faced by the satellite-ground integrated fusion network with inter-satellite links is to establish the optimal routing strategy. In order to solve the optimal routing of the satellite-ground fusion network, the following four problems need to be solved urgently: (1) The high-speed movement of low-orbit satellites will cause frequent dynamic changes in network topology; (2) Relative positioning instability and complex space structure will lead to potential link failure; (3) Satellite power, memory, and bandwidth are limited, which will lead to a decrease in communication overhead; (4) Ultra-large-scale fusion networks make it relatively difficult to maintain high-efficiency data packet forwarding.

[0003] Most of the mainstream dynamic routing methods rely on the shortest path algorithm, which may cause data flows to be frequently transmitted on the same source and destination nodes, and a large amount of data will flow into the shortest forwarding path. Routing algorithms based on the shortest path are prone to load imbalance and local network congestion, which in turn cause data frame loss, delay increase and delay jitter. In addition, there are also deep reinforcement learning methods to deal with the problems of packet loss and uneven link load during routing link switching. However, the following problems still exist: (1) Reinforcement learning methods have the problem of slow training speed when dealing with large-scale networks; (2) The environments of space-based satellites and ground base stations are very different, and the routing algorithms applicable to ground networks may not be applicable to space-based satellite networks; (3) The optimal routing decisions trained in advance are only applicable to the current network, and the performance may degrade after the topology changes. Summary of the invention

[0004] In view of the above problems, the present invention proposes a satellite-ground fusion network routing method and system based on meta-reinforcement learning. Through the meta-reinforcement learning intelligent routing strategy, the problems of slow training speed and multi-environment adaptability of the existing reinforcement learning method are solved. It is suitable for the space-based, air-based and ground-based multi-task environments in the satellite-ground fusion network, and makes up for the limitations of the existing routing algorithm in the satellite-ground fusion network scenario, improves the performance and resource utilization of the computing network, and realizes a more efficient and intelligent routing strategy. The above purpose of the present invention is achieved through the following technical solutions: The present invention provides a satellite-ground fusion network routing method based on meta-reinforcement learning, comprising: Step S1, using the STK toolkit to build a satellite-ground fusion network database and generate a network topology structure; Step S2: Create a multi-task experience pool for storing meta-reinforcement learning training samples; Step S3: Based on the topological structure of the satellite-ground fusion network, define the evaluation index for determining the optimal route, and construct the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; Step S4: Based on the state space, action space and reward function, an online neural network and a target neural network are constructed in the Markov decision process, the state Q value of the current state is determined according to the online neural network, and the state Q value is updated in the target neural network by the Bellman equation, and meta-reinforcement learning training is performed using experience replay and greedy search methods to select corresponding agent actions; Step S5, initialize the meta-reinforcement learning update frequency and update steps, implement the optimal routing strategy by parallel training in multiple network structures, test the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.

[0005] Furthermore, in step S1, the satellite-ground fusion network database is constructed using the STK toolkit, and the satellite-ground fusion network topology is generated, including: The STK toolkit is used to build a satellite-ground fusion network database, which contains data on space-based satellite constellations, air-based network base stations, and ground-based network base stations; and the satellite operation data, relay station, and ground-based network positions of each time slot are obtained according to the satellite-ground fusion network database, where the satellite operation data includes orbital data, geodetic polar coordinates, and angular velocity; Build the corresponding network topology based on the satellite operation data, relay station and ground-based network location in each time slot ;in, It is the set of all links in the satellite-ground fusion network; is the node set in the network topology, including the space-based satellite constellation set , Space-based network base station collection and the set of ground-based network base stations is .

[0006] Further, in step S2, a multi-task experience pool storing training sample data of the meta-reinforcement learning algorithm is constructed; including: Build The optimal routing task of each satellite-ground fusion network is defined as a Markov decision process ,in is the state space matrix, is the action space matrix, The state transition matrix, Reward Matrix; Perform optimal routing tasks in the satellite-ground fusion network During the process, the satellite-ground fusion network interacts with the environment to generate empirical data , and stored in the experience pool The experience pool of all optimal routing tasks The combination constitutes a multi-task experience pool D.

[0007] Further, in step S3, the evaluation index includes path bandwidth, path delay and / or path packet loss rate, and the path is the path from the source to the destination of the satellite-ground fusion network, and the formula is, Path bandwidth: ; Path Delay: ; Path packet loss rate: ; For Link bandwidth, For Link The delay, For Link Packet loss rate; link The bandwidth is ; link The delay is ; link The packet loss rate is ; in, For Link The capacity, For Link The throughput, The time The number of bytes received during the interval; and They are the actual arrival time and theoretical arrival time of link data respectively; for The amount of data sent at any time, for The amount of data received at any given moment.

[0008] Furthermore, the state space is the state set observed by the agent, each state in the state set corresponds to the source-destination node pair of the satellite-ground fusion network; based on the network topology structure established in step S1, the size of the state space is ; Action Space is the set of actions taken on the states in the state space, given the current state in the state set , each action Corresponding to a specific end-to-end path ; The reward value in the reward function R is calculated based on the evaluation index of the path and gives the cost of the potential path in the action space. The formula is, Normalized path bandwidth: Normalized path delay: Normalized path packet loss rate: in, is the minimum and maximum value of each indicator in the data set; In the reward function, the reward value is inversely proportional to the path bandwidth and proportional to the path delay and packet loss rate, and is specifically defined as: ; in is the reward function weight.

[0009] Further, step S4 includes, Construct a target neural network and an online neural network for meta-reinforcement learning, and estimate the current state through the online neural network Q value , the target neural network updates the next state through the Bellman equation of Value; the formula is, ; And train the online neural network to reduce the loss function at each learning step , the formula is, ; Among them, at the beginning of the learning process, the weights of the online neural network and the target neural network are the same. During the training phase, the weights of the target neural network are periodically updated through predefined learning steps to match the online neural network.

[0010] Furthermore, the online neural network and the target neural network have the same structure, including an input layer, a hidden layer, and an output layer. For each state in the state space, the agent encodes each source and destination pair as a state. The input layer has a neuron for receiving the state as the input of the neural network, and the output layer has neurons, used to output the action space actions, each neuron in the output layer estimates the Related value.

[0011] Furthermore, step S4 further includes: The agent uses past decision-making experience as The data is stored in the dataset in the form of batch sampling and offline training is performed on the observed data; The agent uses a decaying greedy search method, through an adjustable parameter Determine the agent with probability Search by expression , throughout the learning process The value is usually from the maximum Start with a decay rate Linearly decreases to the minimum value ; The agent chooses the next action in a specific state according to the following formula; .

[0012] Further, in step S5, it includes: Step S51: Set the number of batch training samples based on the multi-task experience pool and in The number of training update steps in different satellite-ground fusion networks is Initialize experience pool , randomly generate online neural network weights , target network weight ; Step S52: Input the network topology of the satellite-ground fusion network into the online neural network, and randomly sample the optimal route execution action to obtain the satellite-ground fusion network. Then get the next moment status , the reward is calculated according to the reward function in step 3 , and the decision data Store to experience pool .

[0013] Step S53: When the number of experiences in the experience pool is greater than the number of batch training samples When , we randomly extract experience samples as training data for the meta-reinforcement learning algorithm, and during training, we use the loss function of the online neural network and the target neural network as The learning rate is The gradient descent of The meta-reinforcement learning algorithm is trained under different network topologies in the satellite-ground fusion network to obtain The network weight , ; Step S54, determining whether the set number of training update steps has been reached, if so, executing the meta-reinforcement learning steps of the target neural network and the online neural network; otherwise, executing steps S51 to S53; Step S55: Meta-learning updates the network weights obtained from different flight environments; Meta-reinforcement learning-based updates require maximization of decision-making strategies Rewards in different environments: The meta-reinforcement learning update process is as follows: in, represents the weight of the meta-strategy network updated by meta-learning, Indicates the task The network weights learned using the gradient descent algorithm, Represents the meta-learning update learning rate.

[0014] Based on the same inventive concept, the present invention provides a satellite-ground fusion network routing system based on meta-reinforcement learning, which adopts the satellite-ground fusion network routing method as described above, including: The network topology building module is used to build a satellite-ground fusion network database using the STK toolkit and generate a network topology structure; A multi-task experience pool building module, which is used to create a multi-task experience pool for storing meta-reinforcement learning training samples; The data processing module is used to define the evaluation index for determining the optimal route based on the satellite-ground fusion network topology, and to construct the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; The meta-reinforcement learning optimization module is used to construct an online neural network and a target neural network in the Markov decision process based on the state space, action space, and reward function, determine the state Q-value of the current state according to the online neural network, and update the state Q-value in the target neural network using the Bellman equation. It performs meta-reinforcement learning training and selects the corresponding agent actions using the experience replay and greedy search methods; The decision execution module is used to initialize the meta-reinforcement learning update frequency and update steps, implement the optimal routing strategy through parallel training of multiple network structures, test the satellite-ground integrated network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.

[0015] Compared with the prior art, the present invention has at least one of the following beneficial effects: The method of this patent is applicable to the multi-task application environments of space-based, air-based, and ground-based in the satellite-ground integrated network. It can achieve the optimal routing strategy in a new network environment after less training, reasonably utilizes network resources, solves the problems of slow training speed and poor multi-environment adaptability of existing reinforcement learning methods, improves the performance and resource utilization rate of the computing network, and realizes a more efficient and intelligent routing strategy. (1) Through the intelligent routing strategy of meta-reinforcement learning, it solves the problem of environmental adaptability of existing reinforcement learning methods, is applicable to the multi-task environments of space-based, air-based, and ground-based in the satellite-ground integrated network, and only requires less training to achieve the optimal routing strategy in a new network environment, making up for the limitations of existing routing algorithms in the satellite-ground integrated network scenario.

[0016] (2) It can intelligently optimize routing decisions in a time-varying topological structure, improve the performance and resource utilization rate of the computing network through the parallel training method in a multi-task environment, effectively solve the problems of packet loss and uneven link load, and realize a more efficient and intelligent routing strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the steps of the satellite-ground integrated network routing method based on meta-reinforcement learning of the present invention; Figure 2 It is a schematic diagram of the training process of parallel training of satellite-ground integrated network routing meta-reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0019] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0020] First embodiment like Figure 1 , 2 As shown, the present invention provides a satellite-ground fusion network routing method based on meta-reinforcement learning, which aims to solve the problems of slow training speed and multi-environment adaptability of existing reinforcement learning methods through meta-reinforcement learning intelligent routing strategies. By constructing a satellite-ground fusion network database according to the STK toolkit, and constructing a satellite-ground fusion network topology structure; constructing a multi-task experience pool for storing meta-reinforcement learning algorithm training sample data; defining the evaluation index of the optimal routing, and constructing the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; constructing a target neural network and an online neural network, updating the state Q value according to the Bellman equation, using experience replay and greedy search methods to perform meta-reinforcement learning training and select agent actions; initializing the meta-reinforcement learning update frequency and update steps, realizing the optimal routing strategy through parallel training of multiple network structures, testing the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluating the routing decision performance. The specific implementation methods are as follows, including, Step S1, using the Satellite Tool Kit (STK) to build a satellite-ground fusion network database and generate a network topology structure; Step S2: Create a multi-task experience pool for storing meta-reinforcement learning training samples; Step S3: Based on the topological structure of the satellite-ground fusion network, define the evaluation index for determining the optimal route, and construct the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; Step S4: Based on the state space, action space and reward function, an online neural network and a target neural network are constructed in the Markov decision process, the state Q value of the current state is determined according to the online neural network, and the state Q value is updated in the target neural network by the Bellman equation, and meta-reinforcement learning training is performed using experience replay and greedy search methods to select corresponding agent actions; Step S5, initialize the meta-reinforcement learning update frequency and update steps, implement the optimal routing strategy by parallel training in multiple network structures, test the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.

[0021] Furthermore, in step S1, a satellite tool kit (STK) is used to construct a satellite-ground fusion network database and generate a satellite-ground fusion network topology structure, including: The STK toolkit is used to build a satellite-ground fusion network database, which contains data on space-based satellite constellations, air-based network base stations, and ground-based network base stations; and the satellite operation data, relay station, and ground-based network positions of each time slot are obtained according to the satellite-ground fusion network database. Among them, the satellite operation data mainly includes orbital data, geodetic polar coordinates, and angular velocity; Build the corresponding network topology based on the satellite operation data, relay station and ground-based network location in each time slot ;in, It is the set of all links in the satellite-ground fusion network; the link set is all the communication links in the network, that is, the connection relationship between the satellite and the relay station and the ground base station; is a node set in the network topology, which includes a space-based satellite constellation set, an air-based network base station set, and a ground-based network base station set, such as, Space-based satellite constellation collection , G 1 ,G 2 ,...,G n is a satellite node, L 1 ,L 2 ,...,L m Backup satellites or satellites with different orbital parameters are used to enhance network reliability and coverage; Space-based network base station collection , U 1 , U 2 , ..., U P ,These are different air-based network nodes; The set of ground-based network base stations is ,D 1 , D 2 , ..., D q It is a ground network base station node, responsible for ground communication, data transmission and other tasks.

[0022] Further, in step S2, a multi-task experience pool storing training sample data of the meta-reinforcement learning algorithm is constructed; including: Build The optimal routing task of each satellite-ground fusion network is defined as a Markov decision process ,in is the state space matrix, is the action space matrix, The state transition matrix, Reward Matrix; Perform optimal routing tasks in the satellite-ground fusion network During the process, the satellite-ground fusion network interacts with the environment to generate empirical data , and stored in the experience pool The experience pool of all optimal routing tasks The combination constitutes a multi-task experience pool D.

[0023] Further, in step S3, the evaluation index includes path bandwidth, path delay and / or path packet loss rate, and the path is the path from the source to the destination of the satellite-ground fusion network, and the formula is, Path bandwidth: ; Path Delay: ; Path packet loss rate: ; Each path consists of multiple links. For Link bandwidth, For Link The delay, For Link Packet loss rate; link The bandwidth is ; link The delay is ; link The packet loss rate is ; in, For Link The capacity, For Link The throughput, The time The number of bytes received during the interval; and They are the actual arrival time and theoretical arrival time of link data respectively; for The amount of data sent at any time, for The amount of data received at any given moment.

[0024] The meta-reinforcement learning evaluation index is to establish a satellite-ground fusion network routing decision process, that is, to construct a path state index to explore and learn the paths between all sources and destinations. The indicators used in this patent are path bandwidth, path delay, and path packet loss rate.

[0025] Furthermore, the Markov decision process is established from the four parts of reinforcement learning: state space, action space, reward function, target network and online neural network. State Space is the state set observed by the agent, each state in the state set corresponds to the source-destination node pair of the satellite-ground fusion network; based on the network topology structure established in step S1, there are a total of Node, the size of the state space is ; Action Space is the set of actions taken on the states in the state space, given the current state in the state set , each action Corresponding to a specific end-to-end path ; The agent can connect the source and destination nodes Meta-reinforcement learning defines the path from each source node to the destination. The shortest path.

[0026] The reward value in the reward function R is calculated based on the evaluation index of the path, and the cost of the potential path in the action space is given. In order to avoid the inconsistent influence of different path state indicators in meta-reinforcement learning, this patent adopts normalization processing to rescale the indicator range of path bandwidth, path delay, and path packet loss rate to range, that is, Normalized path bandwidth: Normalized path delay: Normalized path packet loss rate: in, is the minimum and maximum value of each indicator in the data set; In the reward function, the reward value is inversely proportional to the path bandwidth and proportional to the path delay and packet loss rate, and is specifically defined as: ; in is the reward function weight.

[0027] Further, step S4 includes, Construct a target neural network and an online neural network for meta-reinforcement learning. During the meta-reinforcement learning training process, the current state is estimated through the online neural network. Q value , the target neural network updates the next state through the Bellman equation of Value; the formula is, ; And train the online neural network to reduce the loss function at each learning step , the formula is, ; Among them, at the beginning of the learning process, the weights of the online neural network and the target neural network are the same. During the training phase, the weights of the target neural network are periodically updated through predefined learning steps to match the online neural network.

[0028] Furthermore, the online neural network and the target neural network have the same structure, including an input layer, a hidden layer, and an output layer. For each state in the state space, the agent encodes each source and destination pair as a state. The input layer has a neuron for receiving the state as the input of the neural network, and the output layer has neurons, used to output the action space actions, each neuron in the output layer estimates the Related Value. The number of hidden layers is defined by the test.

[0029] Furthermore, step S4 further includes: Experience replay: The agent replays past decision-making experience The data set is stored in the form of, experience is sampled in batches and trained offline on the observed data; the data set is used to reduce the number of interactions required for the agent to learn, and small batches can be sampled to reduce the variance of the learning updates; Search method: The agent uses a decaying greedy search method, through an adjustable parameter Determine the agent with probability Search by expression , throughout the learning process The value is usually from the maximum Start with a decay rate Linearly decreases to the minimum value ; The agent chooses the next action in a specific state according to the following formula; .

[0030] Further, in step S5, it includes: Step S51: Set the number of batch training samples based on the multi-task experience pool and in The number of training update steps in different satellite-ground fusion networks is Initialize experience pool , randomly generate online neural network weights , target network weight ; Step S52: Input the network topology of the satellite-ground fusion network into the online neural network, and randomly sample the optimal route execution action to obtain the satellite-ground fusion network. Then get the next moment status , the reward is calculated according to the reward function in step 3 , and the decision data Store to experience pool .

[0031] Step S53: When the number of experiences in the experience pool is greater than the number of batch training samples When , we randomly extract experience samples as training data for the meta-reinforcement learning algorithm, and during training, we use the loss function of the online neural network and the target neural network as The learning rate is The gradient descent of The meta-reinforcement learning algorithm is trained under different network topologies in the satellite-ground fusion network to obtain The network weight , ; Step S54, determining whether the set number of training update steps has been reached, if so, executing the meta-reinforcement learning steps of the target neural network and the online neural network; otherwise, executing steps S51 to S53; Step S55: Meta-learning updates the network weights obtained from different flight environments; Meta-reinforcement learning-based updates require maximization of decision-making strategies Rewards in different environments: The meta-reinforcement learning update process is as follows: in, represents the weight of the meta-strategy network updated by meta-learning, Indicates the task The network weights learned using the gradient descent algorithm, Represents the meta-learning update learning rate.

[0032] In summary, this embodiment is aimed at the time-varying satellite-ground fusion network topology structure. This patent solves the environmental adaptability problem of the existing reinforcement learning method through the meta-reinforcement learning intelligent routing strategy. It is suitable for the space-based, air-based and ground-based multi-task environment in the satellite-ground fusion network. Only a small amount of training is required to achieve the optimal routing strategy in the new network environment, which makes up for the limitations of the existing routing algorithm in the satellite-ground fusion network scenario. In the time-varying topology structure, the routing decision can be intelligently optimized, and the performance and resource utilization of the computing network are improved through the multi-task environment parallel training method, which effectively solves the problems caused by data packet loss and uneven link load, and realizes a more efficient and intelligent routing strategy.

[0033] Second embodiment Based on the same inventive concept, the present invention provides a satellite-ground fusion network routing system based on meta-reinforcement learning, which adopts the satellite-ground fusion network routing method as described above, including: The network topology building module is used to build a satellite-ground fusion network database using the STK toolkit and generate a network topology structure; A multi-task experience pool building module, which is used to create a multi-task experience pool for storing meta-reinforcement learning training samples; The data processing module is used to define the evaluation index for determining the optimal route based on the satellite-ground fusion network topology, and to construct the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; The meta-reinforcement learning optimization module is used to construct an online neural network and a target neural network in the Markov decision process based on the state space, action space and reward function, determine the state Q value of the current state according to the online neural network, and update the state Q value in the target neural network using the Bellman equation. It uses experience replay and greedy search methods to perform meta-reinforcement learning training and select corresponding agent actions; The decision execution module is used to initialize the update frequency and number of steps of meta-reinforcement learning, implement the optimal routing strategy through parallel training in multiple network structures, test the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.

[0034] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

[0035] It should be noted that the above embodiments can be freely combined as needed. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention.

Claims

1. A satellite-ground fusion network routing method based on meta-reinforcement learning, characterized in that: include, Step S1, using the STK toolkit to build a satellite-ground fusion network database and generate a network topology structure; Step S2: Create a multi-task experience pool for storing meta-reinforcement learning training samples; Step S3: Based on the satellite-ground fusion network topology, define the evaluation index for determining the optimal route, and construct the state space, action space and reward function of the satellite-ground fusion network according to the Markov decision process; Step S4: Based on the state space, action space and reward function, an online neural network and a target neural network are constructed in the Markov decision process, a state Q value of the current state is determined according to the online neural network, and the state Q value is updated by the Bellman equation in the target neural network, and the meta-reinforcement learning training is performed by using experience replay and greedy search methods to select corresponding agent actions; Step S5: Initialize the meta-reinforcement learning update frequency and update steps, implement the optimal routing strategy by parallel training on multiple network structures, test the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.

2. The satellite-ground fusion network routing method according to claim 1, characterized in that: In step S1, the satellite-ground fusion network database is constructed using the STK toolkit, and the satellite-ground fusion network topology is generated, further including: The STK toolkit is used to construct a satellite-ground fusion network database, which contains data of space-based satellite constellations, air-based network base stations, and ground-based network base stations; and the satellite operation data, relay station, and ground-based network position of each time slot are obtained according to the satellite-ground fusion network database, wherein the satellite operation data includes orbital data, geodetic polar coordinates, and angular velocity; Based on the satellite operation data, relay station and ground-based network location of each time slot, the corresponding network topology is constructed ;in, It is the set of all links in the satellite-ground fusion network; is a set of nodes in the network topology, including the set of space-based satellite constellations , the air-based network base station set And the set of the ground-based network base stations is .

3. The satellite-ground fusion network routing method according to claim 1, characterized in that: In step S2, a multi-task experience pool is constructed to store the training sample data of the meta-reinforcement learning algorithm; further include, Build The satellite-ground fusion network, and the optimal routing task of each satellite-ground fusion network is defined as a Markov decision process ,in is the state space matrix, is the action space matrix, The state transition matrix, Reward Matrix; Execute the optimal routing task in the satellite-ground fusion network During the process, the satellite-ground fusion network interacts with the environment to generate empirical data , and stored in the experience pool The experience pool of all the optimal routing tasks The combination constitutes the multi-task experience pool D.

4. The satellite-ground fusion network routing method according to claim 1, characterized in that: In step S3, the evaluation index includes the path bandwidth, path delay and / or path packet loss rate, the path is the path from the source to the destination of the satellite-ground fusion network, and the formula is, The path bandwidth: ; The path delay is: ; The packet loss rate of the path: ; For Link bandwidth, For the link The delay, For the link Packet loss rate; The link The bandwidth is ; The link The delay is ; The link The packet loss rate is ; in, For the link The capacity, For the link The throughput, The time The number of bytes received during the interval; and are respectively the actual arrival time and theoretical arrival time of the link data; for The amount of data sent at any time, for The amount of data received at any given moment.

5. The satellite-ground fusion network routing method according to claim 4, characterized in that: The state space is a set of states observed by the agent, each state in the state set corresponds to a source-destination node pair of the satellite-ground fusion network; Based on the network topology established in step S1, the size of the state space is ; The action space is the set of actions taken on the states in the state space, for a given current state in the state set , each action Corresponding to a specific end-to-end path ; The reward value in the reward function R is calculated based on the evaluation index of the path, and the cost of the potential path in the action space is given by the formula: Normalized path bandwidth: Normalized path delay: Normalized path packet loss rate: in, is the minimum and maximum value of each indicator in the data set; In the reward function, the reward value is inversely proportional to the path bandwidth and directly proportional to the path delay and the packet loss rate, and is specifically defined as: ; in is the reward function weight.

6. The satellite-ground fusion network routing method according to claim 5, characterized in that: Step S4 further comprises, Constructing the target neural network and the online neural network for the meta-reinforcement learning, and estimating the current state through the online neural network The Q value of , the target neural network updates the next state through the Bellman equation of Value; the formula is, ; And at each learning step, train the online neural network to reduce the loss function , the formula is, ; Wherein, at the beginning of the learning process, the weights of the online neural network and the target neural network are the same, and during the training phase, the weights of the target neural network are periodically updated through predefined learning steps to match the online neural network.

7. The satellite-ground fusion network routing method according to claim 6, characterized in that: The online neural network and the target neural network have the same structure, including an input layer, a hidden layer and an output layer. For each state in the state space, the agent encodes each source and destination pair as a state. The input layer has a neuron for receiving the state as an input of the neural network. The output layer has neurons, used to output the The action, each neuron in the output layer estimates the action Related value.

8. The satellite-ground fusion network routing method according to claim 7, characterized in that: Step S4 also includes, The agent uses past decision-making experience as The data set is stored in the form of, the experience is sampled in batches and offline training is performed on the observed data; The agent adopts the decaying greedy search method, through an adjustable parameter Determine the agent with probability Search by expression , throughout the learning process The value is usually from the maximum Start with a decay rate Linearly decreases to the minimum value ; The agent selects the next action in a specific state according to the following formula; 。 9. The satellite-ground fusion network routing method according to claim 3 or 8, characterized in that: In step S5, further comprising: Step S51: setting the number of batch training samples based on the multi-task experience pool and in different training update steps in the satellite-ground fusion network, for the satellite-ground fusion network Initialize the experience pool , randomly generate online neural network weights , target network weight ; Step S52: input the network topology of the satellite-ground fusion network into the online neural network, randomly sample and obtain the optimal route execution action of the satellite-ground fusion network. Then get the next moment status , the reward is calculated according to the reward function in step 3 , and the decision data Stored to the experience pool ; Step S53: When the number of experiences in the experience pool is greater than the number of batch training samples , randomly extracting experience samples as training data for the meta-reinforcement learning algorithm, and during training, the loss function of the online neural network and the target neural network is The learning rate is The gradient descent of The meta-reinforcement learning algorithm is trained under the network topology structures in the satellite-ground fusion network to obtain The network weight , ; Step S54, determining whether the set number of training update steps is reached, and if the set number of training update steps is reached, executing the meta-reinforcement learning step of the target neural network and the online neural network; Otherwise, execute step S51 to step S53; Step S55: Meta-learning updates the network weights obtained in different flight environments; Based on the meta-reinforcement learning update, the decision strategy needs to be maximized Rewards in different environments: The meta-reinforcement learning update process is as follows: in, represents the weight of the meta-strategy network updated by meta-learning, Indicates the task The network weights learned using the gradient descent algorithm, Represents the meta-learning update learning rate.

10. A satellite-ground fusion network routing system based on meta-reinforcement learning, using the satellite-ground fusion network routing method according to any one of claims 1 to 9, characterized in that: include, The network topology building module is used to build a satellite-ground fusion network database using the STK toolkit and generate a network topology structure; A multi-task experience pool building module, which is used to create a multi-task experience pool for storing meta-reinforcement learning training samples; A data processing module is used to define evaluation indicators for determining the optimal route based on the satellite-ground fusion network topology, and to construct a state space, action space and reward function of the satellite-ground fusion network according to a Markov decision process; A meta-reinforcement learning optimization module is used to construct an online neural network and a target neural network in the Markov decision process based on the state space, action space and reward function, determine the state Q value of the current state according to the online neural network, and update the state Q value in the target neural network using the Bellman equation, and use experience replay and greedy search methods to perform the meta-reinforcement learning training and select corresponding agent actions; The decision execution module is used to initialize the update frequency and update steps of the meta-reinforcement learning, implement the optimal routing strategy by parallel training on multiple network structures, test the satellite-ground fusion network routing model based on the meta-reinforcement learning algorithm, and evaluate the routing decision performance.