A space-ground integrated satellite internet routing method and system

Through deep reinforcement learning and multi-agent collaboration mechanisms, real-time perception and optimization of the routing method of the integrated satellite Internet of the space and the earth has been solved, and the routing interruption and delay problems caused by the dynamics of satellite nodes and link instability have been achieved, and efficient utilization of network resources and improved user experience have been achieved.

CN120128522BActive Publication Date: 2025-09-02ZHONG KE XING GUANG XIN XI JI SHU YOU XIAN GONG SI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510602594.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-02
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing integrated satellite Internet routing method of space and earth is difficult to adapt to the dynamics of satellite nodes and link instability, resulting in increased routing interruptions and data transmission delays, and it is difficult to meet the quality needs of different services, affecting the efficient utilization of network resources and user experience.

Method used

Adopting deep reinforcement learning model and multi-agent collaboration mechanism, the intelligent routing agent perceives network status in real time, independently learns and optimizes routing strategies, and combines resource scheduling and fault diagnosis to achieve adaptive routing decisions and resource allocation.

Benefits of technology

It improves the stability of routing and the continuity of data transmission, optimizes resource allocation, meets the quality needs of different services, realizes efficient utilization of network resources, reduces routing interruptions and delays, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128522B_ABST
    Figure CN120128522B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of routing and scheduling of satellite communication networks, and discloses a routing method and system for an air-space-ground integrated satellite internet network, wherein an air-space-ground integrated satellite internet network routing method includes the following steps: first, accurately obtain network status information through multi-source heterogeneous information fusion and other means to fully grasp the node status, link quality and network topology changes. Then, based on the obtained information, a scientific and reasonable routing strategy is generated with the help of a deep reinforcement learning model. Subsequently, resource allocation is flexibly adjusted according to the strategy to ensure efficient resource utilization. At the same time, meta-learning and transfer learning mechanisms are introduced to continuously optimize the strategy and improve the adaptability of the strategy. Finally, stable and efficient data transmission is achieved based on the optimized strategy. The present invention can effectively improve the transmission performance and key business service quality of the air-space-ground integrated satellite internet network, and meet the needs of complex and changing network environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of routing and scheduling of satellite communication networks, and more particularly to a routing method and system for an air-ground-integrated satellite internet network. Background Art

[0002] With the rapid development of modern communications technology, integrated space-ground satellite interconnection networks have been widely used in global communications, navigation, remote sensing, and other fields. However, current routing methods for integrated space-ground satellite interconnection networks still face many challenges. For example, in the complex and ever-changing space environment, the dynamic nature of satellite nodes and the instability of links pose significant challenges to route selection. Specifically, the high-speed movement of satellites causes frequent changes in network topology. Traditional routing algorithms based on fixed topologies are unable to adapt to these dynamic changes, easily resulting in routing interruptions or significantly increased data transmission delays.

[0003] Taking the global logistics information network as an example, during the transportation of goods, information such as the location and status of the goods must be transmitted in real time. If routing interruptions or delays prevent this information from being transmitted accurately and promptly, the accuracy and timeliness of logistics scheduling may be affected, leading to delivery delays and increased logistics costs. Furthermore, different types of services (such as voice, data, and video) have significantly different network quality of service (QoS) requirements. Existing routing methods struggle to meet the QoS requirements of various services while also efficiently utilizing network resources. For example, video services have high bandwidth and latency requirements, and traditional routing methods may be unable to dynamically allocate resources based on service needs, resulting in video transmission issues such as lag and reduced image quality, impacting the user experience. Summary of the Invention

[0004] The present invention provides an air-space-ground integrated satellite internet routing method, comprising:

[0005] Obtaining network status information, which includes the working status data of the satellite node obtained by the intelligent routing agent from the status monitoring module of the satellite node itself, the link quality related data obtained by the link quality detection device, and the topology structure information of the satellite network;

[0006] Routing decision-making, intelligent routing agents are based on deep reinforcement learning models to autonomously learn and optimize routing strategies based on network status and transportation demand;

[0007] Multi-agent collaboration: Through the federated learning mechanism, multiple intelligent routing agents exchange routing strategy knowledge to achieve global collaborative optimization;

[0008] Resource scheduling optimization: the central controller optimizes and schedules network resources based on the decisions of the intelligent routing agent and logistics transportation needs;

[0009] Policy update: The intelligent routing agent updates the routing policy using the policy gradient algorithm based on the reward feedback obtained from executing actions;

[0010] The method further includes:

[0011] Adaptive network state awareness, adaptive intelligent routing agent AIRA;

[0012] Intelligent routing strategy adaptation: AIRA is based on a deep reinforcement learning model and can autonomously learn and optimize routing strategies based on the perceived network status.

[0013] Distributed multi-agent collaboration: AIRA builds a distributed trust mechanism through blockchain technology to achieve secure information sharing and collaborative decision-making;

[0014] Autonomous fault diagnosis and recovery: AIRA uses anomaly detection and fault diagnosis algorithms to monitor and locate faults and anomalies in the network in real time;

[0015] Policy evaluation and optimization,AIRA evaluates the effectiveness and optimization of intelligent routing policies based on,business requirements and service quality indicators.

[0016] In a preferred embodiment, the acquired network status information is cleaned and filtered to remove outliers; then the data is standardized; a graph neural network is used to extract and aggregate features from the standardized data, and feature representations of nodes and edges are learned through graph convolution operations to generate input data for routing strategy generation.

[0017] In a preferred embodiment, the deep reinforcement learning model in routing decision-making adopts the ActorCritic framework, which includes a policy network and value network ; The policy network is based on the state Select routing action The probability distribution of the value network estimate state The expected long-term return.

[0018] In a preferred embodiment, in multi-agent collaboration, each agent uploads local strategy parameters To the central controller; the central controller aggregates the strategy parameters of each agent and updates the global strategy , and then distribute it to each agent to update the local strategy.

[0019] In a preferred embodiment, in resource scheduling optimization, the goal of resource scheduling optimization is to minimize network load imbalance and maximize the service quality of key services.

[0020] In a preferred embodiment, a Kalman filter algorithm is used to dynamically estimate and predict the network status to improve the accuracy and real-time performance of perception.

[0021] In a preferred embodiment, a deep reinforcement learning model is used to adapt to the continuous state and action space, and meta-learning and transfer learning mechanisms are introduced to accelerate the strategy learning and adaptation process.

[0022] In a preferred embodiment, a fault propagation model based on a graph neural network is used to learn the propagation law of faults in the network topology and predict the scope and extent of fault impact.

[0023] In a preferred embodiment, an air-ground integrated satellite internet routing system includes:

[0024] State perception module: responsible for sensing satellite node status and link quality, thereby obtaining network status information, and cleaning, filtering and standardizing the obtained network status information to provide accurate and reliable input data for routing strategy generation;

[0025] Routing strategy generation module: Based on the network status information obtained by the state perception module, the network status information is processed and input into the deep reinforcement learning model. According to the network status and transportation demand, the routing strategy is autonomously learned and optimized;

[0026] Resource allocation module: Dynamically adjusts network resource allocation based on the decisions of intelligent routing agents and logistics transportation needs;

[0027] Policy Optimization Module: Utilizes intelligent learning algorithms to optimize routing strategies and improve system adaptability. The policy optimization module includes an intelligent learning unit that uses reinforcement learning and deep learning algorithms to learn and optimize routing strategies. The policy update unit updates the optimized routing strategies to the routing strategy generation module to ensure real-time effectiveness of the routing strategies.

[0028] Data transmission module: completes data transmission tasks based on optimized routing strategies.

[0029] The beneficial effects of the present invention are:

[0030] Improve routing stability: By real-time sensing of satellite node status and link quality, combined with a dynamic topology prediction algorithm, it can quickly adapt to changes in network topology, effectively reduce routing interruptions, and ensure the continuity and stability of data transmission.

[0031] Optimize resource allocation: To meet the QoS requirements of different services, a resource allocation strategy based on service classification is adopted. This can dynamically adjust network resource allocation according to service type, achieving efficient utilization of network resources while meeting the QoS requirements of various services. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flow chart of a space-ground integrated satellite internet routing method of the present invention;

[0033] Figure 2 is an example diagram of network status information of the present invention;

[0034] Figure 3 It is the probability distribution diagram of the routing action of the present invention;

[0035] Figure 4 This is an example diagram of the technical effects of the present invention. DETAILED DESCRIPTION

[0036] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0037] At least one embodiment of the present invention discloses a space-ground integrated satellite internet routing method, such as Figure 1 As shown, the following steps are included:

[0038] Step 100, obtaining network status information, which includes the working status data of the satellite node obtained by the intelligent routing agent from the status monitoring module of the satellite node itself, the link quality related data obtained by the link quality detection device, and the topology information of the satellite network;

[0039] In one embodiment of the present invention, the working status data of the satellite node includes: the power of the satellite, signal strength, orbit parameters, etc., which can reflect whether the satellite node is working normally and its current operating status;

[0040] Link quality data includes information such as link bandwidth, latency, and packet loss rate, which are crucial for evaluating the communication performance of satellite links.

[0041] The topology information of the satellite network includes the connection relationship between each satellite node.

[0042] In one embodiment of the present invention, the acquired network status information is cleaned and filtered to remove outliers. The data is then normalized, for example, by normalizing the battery percentage to the range [0, 1], standardizing link bandwidth to Mbps, and using Z-scores to normalize latency data. A graph neural network (GNN) is then used to extract and aggregate features from the normalized data. Graph convolution operations are used to learn feature representations for nodes and edges, generating input data that can be used to generate routing policies.

[0043] The acquired network status information is processed through a graph neural network to obtain the locally observed network status information;

[0044] In one embodiment of the present invention, a graph neural network uses graph convolution and other operations to aggregate node and edge features and learn network topology and attribute representation. represents the network topology, where is a collection of nodes, is the edge set. For node , whose characteristics are expressed as . Calculated by graph convolution formula:

[0045] ;

[0046] in It is Layer Node Hidden representation of ; It is Layer Node Hidden representation of ; is a node The set of neighbor nodes of It is The weight matrix of the layer; is the bias vector; is a non-linear activation function.

[0047] Step 200, routing decision, the intelligent routing agent uses a deep reinforcement learning model to autonomously learn and optimize routing strategies based on network status and transportation demand;

[0048] In one embodiment of the present invention, the deep reinforcement learning model adopts the ActorCritic framework, which includes a policy network and value network ;

[0049] The policy network is based on the state Select routing action The probability distribution of the value network estimate state The expected long-term return under the condition of ; the decision process can be modeled as MDP (Markov decision process), that is, the five-tuple , in the quintuple represents the state space, which contains all possible states of the network, represents the action space, which contains all routing decisions that the agent can perform, Represents the state transition probability function, which defines the probability of transferring to the next state s' after executing action a in the current state s. Represents the reward function, which defines the immediate reward obtained by performing action a in state s. Represents a discount factor, which is used to balance the importance of immediate rewards and future rewards. The update formula of the policy network is:

[0050] ;

[0051] in, Representation Policy Network The policy network is based on the state Select routing action The probability distribution of is the first learning rate; Indicates the parameters Find the gradient; It is in strategy Next, action State-based The logarithm of the probability of selection; is the action-value function; Representing the Value Network Status An estimate of the expected long-term return;

[0052] The update formula of the value network is:

[0053] ;

[0054] in, Representing the Value Network Parameters; is the second learning rate; It’s an immediate reward; represents the second discount factor, The value is 0-1; Represents the value network for the next state estimates of expected long-term returns; Represents the value network's response to the current state estimates of expected long-term returns; Indicates the parameters Find the gradient; is the next state, that is, the agent is in the current state Take action The new state to which it is transferred.

[0055] Step 300: Multi-agent collaboration. Through the federated learning mechanism, multiple intelligent routing agents exchange routing strategy knowledge to achieve global collaborative optimization.

[0056] In one embodiment of the present invention, each agent uploads local policy parameters To the central controller; local policy parameters include: policy network weight matrix, bias vector, convolution kernel parameters, etc.; the central controller aggregates the policy parameters of each agent and updates the global policy , and then distribute it to each agent to update the local strategy. Assume that there are The formula for the central controller aggregation strategy parameters is: .

[0057] Step 400: Resource scheduling optimization. The central controller optimizes and schedules network resources based on the decisions of the intelligent routing agent and logistics transportation requirements.

[0058] In one embodiment of the present invention, the goal of resource scheduling optimization is to minimize the network load imbalance. ;in For Link The utilization rate, represents the set of all links in the network, Maximize the service quality of key businesses ;in For business flow The weight of Represents the set of all business flows in the network, metric for its service quality.

[0059] Step 500, policy update: the intelligent routing agent updates the routing policy using the policy gradient algorithm based on the reward feedback obtained from executing the action;

[0060] The policy gradient is defined as follows:

[0061] ;

[0062] in, represents the mathematical expectation, represents the policy performance objective function, Indicates the parameters Find the gradient, that is, in the strategy Next, action State-based Take the logarithm of the probability of selection and find the parameter The gradient of , then multiplied by the action value function , and finally the results of the strategy Find the expectation; update the policy network parameters through optimization algorithms such as stochastic gradient ascent The specific update formula is: ,in is the third learning rate.

[0063] In one embodiment of the present invention, an example of the aforementioned air-space-ground integrated satellite internet routing method is provided:

[0064] Assume that in an integrated space-ground satellite internet scenario, there are 5 satellite nodes and a collection of 8 links ;

[0065] State awareness:

[0066] like Figure 2 As shown, it includes node, node status characteristics, and link quality example data;

[0067] Example of graph convolution operation:

[0068] Assumptions , for , , is the ReLU function, for the node , its neighbor nodes , , , .

[0069] ;

[0070] Action selection:

[0071] Set current state is the network topology and the status of each node, and the transportation demand is Transfer data to . Policy Network Give different routing actions (such as Then to ,go through Then to The probability distribution of Figure 3 As shown;

[0072] Reward calculation:

[0073] Value Network Estimate current state The expected long-term return under is 0.7; let the learning rate , For action, it is 0.8;

[0074] Strategic Network Updates: ;

[0075] Value Network Update: ;

[0076] Load balancing optimization:

[0077] Assume there are 5 links ( ), link utilization , , , , , average utilization ;

[0078] Optimizing the quality of key business services:

[0079] There are 3 business flows ( ), business flow weight , , , service quality metrics , , ;

[0080] Multi-objective optimization:

[0081] Network load imbalance:

[0082] ;

[0083] Key business service quality:

[0084] ;

[0085] Using the weighted method, , the single-objective optimization problem is:

[0086] ;

[0087] Federated learning parameter aggregation:

[0088] There are three intelligent agents ( ), whose local policy parameters , , Central controller aggregation strategy parameters:

[0089] ;

[0090] Examples of technical performance data include Figure 4 As shown;

[0091] In one embodiment of the present invention, the satellite internet environment is complex and changeable. To better adapt to the complex and changeable environment of the satellite internet, a space-ground integrated satellite internet routing method further includes:

[0092] s101, Adaptive Network Status Awareness, Adaptive Intelligent Routing Agent (AIRA) integrates multi-source heterogeneous information to perceive node status, link quality, and network topology changes in real time;

[0093] In one embodiment of the present invention, an algorithm such as Kalman filtering is used to dynamically estimate and predict the network state to improve the accuracy and real-time performance of perception; the state transition model is:

[0094] ;

[0095] in, express The state of the moment, which is determined by State of the moment 、 Control input at any time and process noise Through nonlinear functions Got it.

[0096] The observation model is:

[0097] ;

[0098] in represents the observed value, represents the observation noise, the function represents a nonlinear function.

[0099] The prediction steps of Kalman filter are:

[0100] ;

[0101] in express The predicted state at the moment, express The predicted state at the moment, express The control input at the moment, represents a nonlinear function.

[0102] ;

[0103] in express The forecast covariance at time t, represents the state transition matrix, express The estimated covariance at time t, represents the transposed matrix of the state transition matrix, represents the process noise covariance.

[0104] The update steps are:

[0105] ;

[0106] in represents the Kalman gain, represents the observation matrix, represents the transpose of the observation matrix, represents the observation noise covariance.

[0107] ;

[0108] in, Indicates a more accurate update The estimated state at the moment, Indicates a more accurate update The estimated state at the moment; represents the Kalman gain, represents the observed value;

[0109] ;

[0110] in, represents the updated estimated covariance, represents the identity matrix, represents the Kalman gain, represents the observation matrix, express The forecast covariance at time t.

[0111] s102, intelligent routing strategy adaptation, AIRA is based on a deep reinforcement learning model and autonomously learns and optimizes routing strategies based on the perceived network status;

[0112] In one embodiment of the present invention, a reinforcement learning method is used to adapt to the continuous state and action space, and meta-learning and transfer learning mechanisms are introduced to accelerate the policy learning and adaptation process;

[0113] The value function is defined as:

[0114] ;

[0115] in, Indicates The status at the moment is And take action The value function when represents the mathematical expectation; for Instant rewards of the moment; is the third discount factor, whose value is between arrive between, The closer , indicating that the more importance is attached to future rewards, The closer , the more you focus on immediate rewards; yes The state of the moment, yes Actions that can be taken at any time, Indicates The maximum Q value that can be obtained from all possible actions at the moment.

[0116] In Deep Deterministic Policy Gradient (DDPG), the policy network The update formula is:

[0117] .

[0118] in, It is a strategic network Parameters; is the fourth learning rate; is the value function About parameters gradient.

[0119] Value Network The update formula is as follows:

[0120] ;

[0121] in, It is a value network Parameters; is the fifth learning rate; For immediate rewards; is the fourth discount factor; Indicates the next state, is the target policy network In state The following actions are taken, are the target policy network parameters; is the target value network in state And take action The value of time; Is the current value network in state And take action The value of time; is the value function About parameters gradient.

[0122] s103, distributed multi-agent collaboration, AIRA uses blockchain technology to build a distributed trust mechanism to achieve secure and reliable information sharing and collaborative decision-making;

[0123] Adopt smart contracts and consensus algorithms, such as Proof of Work (PoW), to incentivize cooperation and consensus among AIRAs;

[0124] Introducing a federated learning framework, AIRA trains model parameters locally, and the central controller aggregates the model parameters of each AIRA, updates the global model, and then distributes it to AIRA to update the local model;

[0125] In one embodiment of the present invention, the distributed multi-agent collaboration steps are as follows:

[0126] The global goal is to minimize the weighted loss function:

[0127] ;

[0128] in, represents the global model parameters, Indicates the number of AIRA, Indicates the The importance weight of each AIRA; Indicates the The loss function of AIRA.

[0129] The node calculates the hash value:

[0130] ; ;

[0131] in, Indicates that through Hash value obtained by hash calculation; Represents the block header information, which contains the metadata of the block, such as version number, timestamp, hash value of the previous block, etc. Represents the target hash value, which defines a difficulty threshold, and only the calculated hash value Only when the hash value is less than the target value is the node considered to have successfully completed the proof of work and obtained the right to record accounts.

[0132] s104, autonomous fault diagnosis and recovery, AIRA uses anomaly detection and fault diagnosis algorithms to monitor and locate faults and anomalies in the network in real time;

[0133] In one embodiment of the present invention, a fault propagation model based on a graph neural network is used to learn the propagation law of faults in the network topology and predict the scope and extent of fault impact;

[0134] The fault diagnosis model is defined as:

[0135] ;

[0136] in, In a given network topology Under the conditions, the fault state The conditional probability distribution of ; Represents a network topology graph, which describes the connection relationship between each node in the network; Represents the diagnostic function of GNN parameterization, which is composed of the parameters of the graph neural network Decision, used according to the network topology diagram Predicting failure status The probability distribution of

[0137] Fault propagation follows the linear propagation model and the calculation formula is as follows:

[0138] ;

[0139] in, Representation node Fault status; Representation node and The propagation weight between them; Representation node The set of neighbor nodes of Represents the bias parameter.

[0140] s105, policy evaluation and optimization, AIRA evaluates the effectiveness and optimization of intelligent routing policies based on business requirements and service quality indicators;

[0141] In one embodiment of the present invention, policy evaluation and optimization uses a multi-objective reinforcement learning algorithm to balance multiple performance indicators such as latency, throughput, and reliability;

[0142] At the same time, game theory and mechanism design theory are introduced to optimize the incentive and constraint mechanisms in the multi-agent collaboration process;

[0143] The immediate reward function calculation formula is as follows:

[0144] ;

[0145] in, Indicates Instant rewards at all times, 、 、 Represent the first, second and third weight coefficients respectively, express The delay indicator at the moment, express Throughput indicator at the moment, express Reliability index at the moment.

[0146] Solve multi-objective optimization problems, use Pareto optimality method, and find all non-dominated solutions In multi-objective optimization problems, a non-dominated solution is one that does not have any other solution that is strictly superior to it in all objectives. , does not exist Strictly superior to or Strictly superior to This means that in the set Any two solutions in cannot surpass each other in all performance indicators at the same time. These non-dominated solutions constitute the Pareto front, which represents the optimal set of choices when balancing multiple objectives.

[0147] In one embodiment of the present invention, in order to further optimize the routing performance of the integrated space-ground-satellite internet network and improve the overall efficiency of the network, a routing method for the integrated space-ground-satellite internet network further includes the following steps:

[0148] s201, the adaptive intelligent routing agent AIRA, is deployed on each network node and achieves adaptive network status awareness by fusing multi-source heterogeneous information.

[0149] In one embodiment of the present invention, in a deep reinforcement learning model: It is a strategic network Parameters, Represents the state, Indicates routing action; It is a value network Parameters, The decision process is still modeled as a Markov decision process (MDP), that is, ;

[0150] The calculation formula for multi-source heterogeneous information fusion is as follows:

[0151]

[0152] in, represents the network state after fusion, is a nonlinear activation function, is the number of information sources, is the weight coefficient of the kth information source, Indicates the kth information source Feature extraction function;

[0153] Policy network update formula:

[0154] ;

[0155] in, and Represent the policy network parameters before and after the update, is the sixth learning rate, is the policy gradient;

[0156] Fault propagation still follows the linear propagation model;

[0157] ;

[0158] in is a node Fault status, is a node and The propagation weight between Representation node The set of neighbor nodes of is the bias.

[0159] s202, the central controller, communicates with each AIRA; aggregates AIRA status and decision information, and implements distributed multi-agent collaboration through mechanisms such as federated learning;

[0160] In one embodiment of the present invention, each agent uploads local policy parameters To the central controller, the central controller aggregates the policy parameters of each agent and updates the global policy The global goal is to minimize the weighted loss function.

[0161] Global model parameter update formula:

[0162]

[0163] in, represents the updated global model parameters, Indicates the The amount of AIRA data, Indicates the total data volume of all AIRAs, Indicates the AIRA local model parameters;

[0164] Global weighted loss function:

[0165]

[0166] in, represents the global model parameters, Indicates the total data volume of all AIRAs, Indicates the AIRA weights, Indicates the AIRA is based on global model parameters The loss function is Indicates the parameters Perform minimal optimization;

[0167] s203, blockchain network, which is used as a secure and reliable distributed communication and storage platform between AIRAs;

[0168] In PoW, It is through The hash value obtained by hash calculation. It is the block header information; is the target hash value, the node calculates the hash value Need to meet Only then can the proof of work be completed and the right to keep accounts be obtained.

[0169] Blockchain consensus difficulty adjustment formula:

[0170]

[0171] in, is the target hash value, Indicates the target hash value of the previous block, Indicates the actual generation time of the previous block, Indicates the expected block generation time;

[0172] s204, strategy evaluation and optimization module;

[0173] Based on the service quality and network performance indicators, the immediate reward function is still defined as:

[0174] ;

[0175] in, Indicates Instant rewards at all times, 、 、 Represent the first, second and third weight coefficients respectively, express The delay indicator at the moment, express Throughput indicator at the moment, express Reliability index at each moment;

[0176] Cumulative reward calculation formula:

[0177]

[0178] in, represents the cumulative reward starting from time t, is the fifth discount factor, represents the immediate reward at time t+k;

[0179] In order to solve the multi-objective optimization problem, the Pareto optimal method is used to find all non-dominated solution sets. ,in represents the set of all non-dominated solutions. For any two solutions , does not exist Strictly superior to or Strictly superior to We use multi-objective reinforcement learning and game theory to evaluate and optimize AIRA's intelligent routing strategy. We balance multiple performance objectives to improve strategy effectiveness and global optimality.

[0180] s205, the fault diagnosis and recovery module, builds a fault propagation model based on graph neural networks;

[0181] Learn the causal relationships and propagation patterns of faults in the network. This allows for fault tracing, impact prediction, and recovery decision-making. This minimizes service interruptions and performance losses caused by faults.

[0182] Graph neural network layer update formula:

[0183]

[0184] in, Indicates that node i is The hidden state of the layer, is a nonlinear activation function, represents the set of neighbor nodes of node i, is the normalization constant, and Respectively The layer's weight matrix and bias vector, Indicates that node j is in The hidden state of the layer;

[0185] Fault impact range prediction formula:

[0186]

[0187] in, So f represents the influence range, represents the probability of node i failing under the condition of failure f, represents the importance weight of node i, and N represents the total number of nodes in the network;

[0188] s206, a human-computer interaction interface, provides an intuitive and friendly visual interface; displays key information such as the global network status, routing strategies, and performance indicators; supports manual intervention and strategy adjustments to achieve the explainability and controllability of the intelligent routing system.

[0189] Explainability score calculation formula:

[0190]

[0191] in, represents the explainability score of action a, 、 、 Represent the fourth, fifth, and sixth weight coefficients respectively, Indicates the clarity of the causal relationship of the action a, Indicates the transparency of action a, Indicates the degree of visualization of action a;

[0192] The modules achieve interconnection and data sharing through standardized interfaces, and work collaboratively; forming an adaptive, intelligent, secure and efficient integrated satellite interconnection networking routing solution to support the intelligent operation and optimized management of the global logistics network.

[0193] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A satellite internet routing method for integrated space-ground-air communication, characterized in that: include: Obtaining network status information, which includes the working status data of the satellite node obtained by the intelligent routing agent from the status monitoring module of the satellite node itself, the link quality related data obtained by the link quality detection device, and the topology structure information of the satellite network; Routing decision-making, intelligent routing agents are based on deep reinforcement learning models to autonomously learn and optimize routing strategies based on network status and transportation demand; Multi-agent collaboration: Through the federated learning mechanism, multiple intelligent routing agents exchange routing strategy knowledge to achieve global collaborative optimization; Resource scheduling optimization: the central controller optimizes and schedules network resources based on the decisions of the intelligent routing agent and logistics transportation needs; Policy update: The intelligent routing agent updates the routing policy using the policy gradient algorithm based on the reward feedback obtained from executing actions; The method further includes: Adaptive network state awareness, adaptive intelligent routing agent AIRA; Intelligent routing strategy adaptation: AIRA is based on a deep reinforcement learning model and can autonomously learn and optimize routing strategies based on the perceived network status. Distributed multi-agent collaboration: AIRA builds a distributed trust mechanism through blockchain technology to achieve secure information sharing and collaborative decision-making; Autonomous fault diagnosis and recovery: AIRA uses anomaly detection and fault diagnosis algorithms to monitor and locate faults and anomalies in the network in real time; Policy evaluation and optimization,AIRA evaluates the effectiveness and optimization of intelligent routing policies based on,business requirements and service quality indicators.

2. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: The acquired network status information is cleaned and filtered to remove outliers; the data is then standardized; a graph neural network is used to extract and aggregate features from the standardized data, and the feature representations of nodes and edges are learned through graph convolution operations to generate input data for routing strategy generation.

3. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: The deep reinforcement learning model in routing decision-making adopts the ActorCritic framework, which includes a policy network and value network ; The policy network is based on the state Select routing action The probability distribution of the value network estimate state The expected long-term return.

4. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: In multi-agent collaboration, each agent uploads local strategy parameters To the central controller; the central controller aggregates the strategy parameters of each agent and updates the global strategy , and then distribute it to each agent to update the local strategy.

5. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: In resource scheduling optimization, resource scheduling optimization is to minimize network load imbalance and maximize the service quality of key businesses.

6. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: The Kalman filter algorithm is used to dynamically estimate and predict the network status to improve the accuracy and real-time performance of perception.

7. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: A deep reinforcement learning model is adopted to adapt to the continuous state and action space, and meta-learning and transfer learning mechanisms are introduced to accelerate the strategy learning and adaptation process.

8. The method for routing satellite internet networks based on space-ground integration according to claim 1, wherein: A fault propagation model based on graph neural network is used to learn the propagation law of faults in network topology and predict the scope and extent of fault impact.

9. An air-ground integrated satellite internet routing system, characterized in that: A method for executing an air-space-ground integrated satellite internet routing method according to any one of claims 1 to 8, comprising: State perception module: responsible for sensing satellite node status and link quality, thereby obtaining network status information, and cleaning, filtering and standardizing the obtained network status information to provide accurate and reliable input data for routing strategy generation; Routing strategy generation module: Based on the network status information obtained by the state perception module, the network status information is processed and input into the deep reinforcement learning model. According to the network status and transportation demand, the routing strategy is autonomously learned and optimized; Resource allocation module: Dynamically adjusts network resource allocation based on the decisions of intelligent routing agents and logistics transportation needs; Policy Optimization Module: Utilizes intelligent learning algorithms to optimize routing strategies and improve system adaptability. The policy optimization module includes an intelligent learning unit that uses reinforcement learning and deep learning algorithms to learn and optimize routing strategies. The policy update unit updates the optimized routing strategies to the routing strategy generation module to ensure real-time effectiveness of the routing strategies. Data transmission module: completes data transmission tasks based on optimized routing strategies.

Citation Information

Patent Citations

  • Low earth orbit satellite constellation routing method, system and device and storage medium

    CN119966885A