Space-air-ground integrated satellite internet network routing method and space-air-ground integrated satellite internet network routing system
Through the intelligent routing agent, the routing strategy is optimized based on the deep reinforcement learning model and the federated learning mechanism realizes multi-agent collaboration, which solves the problem that the integrated satellite Internet routing method of the space-space and earth is difficult to adapt to complex environments and meets diversified business needs, and improves routing stability and resource utilization efficiency.
Patent Information
- Application Number
- CN202510602594.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The current routing method of the integrated satellite Internet of the space and the earth is difficult to adapt to the complex and changeable space environment, resulting in a significant increase in routing interruption or data transmission delay, and it is difficult to meet the diversified needs of different services for network service quality, resulting in low resource utilization.
The intelligent routing agent is adopted based on the deep reinforcement learning model, obtains network status information and transportation needs, independently learns and optimizes routing strategies, and realizes multi-agent collaboration through federated learning mechanisms, and dynamically adjusts network resource allocation.
It improves routing stability, reduces the occurrence of routing interruptions, ensures the continuity and stability of data transmission, and meets the QoS needs of different services through dynamic resource allocation, and realizes efficient utilization of network resources.
Smart Images

Figure CN120128522A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of routing and scheduling in satellite communication networks, and more specifically, it relates to a routing method and system for an integrated space-air-ground satellite interconnection network. Background Art
[0002] With the rapid development of modern communication technologies, integrated space-air-ground satellite interconnection networks have been widely used in global communication, navigation, remote sensing and other fields. However, the current routing methods for integrated space-air-ground satellite interconnection networks still face many challenges. For example, in the complex and changing space environment, the dynamic nature of satellite nodes and the instability of links pose great difficulties to routing selection. Specifically, the high-speed movement of satellites causes frequent changes in the network topology, and traditional routing algorithms based on fixed topologies are difficult to adapt to this dynamic change, easily resulting in routing interruptions or a significant increase in data transmission delays.
[0003] Taking the global logistics information network as an example, during the transportation of goods, it is necessary to transmit real-time information such as the location and status of goods. If the information cannot be transmitted in a timely and accurate manner due to routing interruptions or delays, it may affect the accuracy and timeliness of logistics scheduling, thereby causing delays in goods delivery and increasing logistics costs. In addition, different types of services (such as voice, data, video, etc.) have quite different requirements for network service quality (QoS). Existing routing methods are difficult to achieve efficient utilization of network resources while meeting the QoS requirements of multiple services. For example, video services have high requirements for bandwidth and latency, while traditional routing methods may not be able to dynamically allocate resources according to service requirements, resulting in problems such as stuttering and decreased video quality during video transmission, affecting the user experience. Summary of the Invention
[0004] The present invention provides a routing method for an integrated space-air-ground satellite interconnection network, including:
[0005] Obtaining network status information, where the network status information includes the working status data of satellite nodes obtained by the intelligent routing agent from the status monitoring module of the satellite nodes themselves, the link quality-related data obtained through the link quality detection device, and the topological structure information of the satellite network;
[0006] Routing decision-making, where the intelligent routing agent, based on a deep reinforcement learning model, autonomously learns and optimizes routing strategies according to the network status and transportation requirements;
[0007] Multi-agent cooperation, where multiple intelligent routing agents exchange routing strategy knowledge through a federated learning mechanism to achieve global collaborative optimization;
[0008] Resource scheduling optimization, where the central controller optimally schedules network resources according to the decisions of the intelligent routing agent and the logistics transportation requirements;
[0009] Policy update: The intelligent routing agent updates the routing policy through the policy gradient algorithm based on the reward feedback obtained from the execution actions.
[0010] In a preferred embodiment, the obtained network state information is cleaned and filtered to remove outliers; then the data is normalized; a graph neural network is used to extract and aggregate features from the normalized data, and through graph convolution operations, the feature representations of nodes and edges are learned to generate the input data for routing policy generation.
[0011] In a preferred embodiment, the deep reinforcement learning model in the routing decision uses the Actor-Critic framework, which includes a policy network and a value network ; the policy network selects the probability distribution of routing actions according to the state ; the value network estimates the expected long-term return under the state .
[0012] In a preferred embodiment, in multi-agent cooperation, each agent uploads its local policy parameters to the central controller; the central controller aggregates the policy parameters of each agent, updates the global policy , and then distributes it to each agent to update its local policy.
[0013] In a preferred embodiment, in resource scheduling optimization, the goal of resource scheduling optimization is to minimize the network load imbalance and maximize the quality of service of critical services.
[0014] In a preferred embodiment, the satellite interconnection network environment is complex and changeable. To better adapt to the complex and changeable environment of the satellite interconnection network, an air-space-ground integrated satellite interconnection network routing method further includes:
[0015] Adaptive network state perception, the adaptive intelligent routing agent AIRA;
[0016] Intelligent routing policy adaptation, AIRA is based on a deep reinforcement learning model and autonomously learns and optimizes the routing policy according to the perceived network state;
[0017] Distributed multi-agent cooperation, AIRA constructs a distributed trust mechanism through blockchain technology to achieve secure information sharing and collaborative decision-making;
[0018] Autonomous fault diagnosis and recovery, AIRA uses anomaly detection and fault diagnosis algorithms to continuously monitor and locate faults and anomalies in the network;
[0019] Policy evaluation and optimization, AIRA evaluates the effectiveness and optimization degree of the intelligent routing policy according to business requirements and quality of service indicators.
[0020] In a preferred embodiment, the Kalman filtering algorithm is used to dynamically estimate and predict the network state, improving the accuracy and real-time performance of perception.
[0021] In a preferred embodiment, a deep reinforcement learning model is adopted to adapt to continuous state and action spaces, and meta-learning and transfer learning mechanisms are introduced to accelerate the policy learning and adaptation processes.
[0022] In a preferred embodiment, a fault propagation model based on graph neural networks is adopted to learn the propagation law of faults in the network topology and predict the scope and degree of fault impact.
[0023] In a preferred embodiment, an integrated space-air-ground satellite interconnection network routing system includes:
[0024] A state perception module: responsible for perceiving the states of satellite nodes and link qualities, thereby obtaining network state information, and performing cleaning, screening, and standardization processing on the obtained network state information to provide accurate and reliable input data for routing policy generation;
[0025] A routing policy generation module: based on the network state information obtained by the state perception module, processes the network state information and inputs it into a deep reinforcement learning model, and autonomously learns and optimizes routing policies according to the network state and transportation requirements;
[0026] A resource allocation module: dynamically adjusts network resource allocation according to the decisions of intelligent routing agents and logistics transportation requirements;
[0027] A policy optimization module: uses intelligent learning algorithms to optimize routing policies and improve system adaptability; the policy optimization module includes an intelligent learning unit that learns and optimizes routing policies using reinforcement learning and deep learning algorithms; a policy update unit updates the optimized routing policies to the routing policy generation module to ensure the real-time effectiveness of routing policies;
[0028] A data transmission module: completes data transmission tasks according to the optimized routing policies.
[0029] The beneficial effects of the present invention are as follows:
[0030] Improve routing stability: By real-time perceiving the states of satellite nodes and link qualities and combining with a dynamic topology prediction algorithm, it can quickly adapt to changes in the network topology, effectively reduce the occurrence of routing interruptions, and ensure the continuity and stability of data transmission.
[0031] Optimize resource allocation: For the QoS requirements of different services, adopt a resource allocation strategy based on service classification, which can dynamically adjust network resource allocation according to service types, and achieve efficient utilization of network resources while meeting the QoS requirements of various services. Description of the Drawings
[0032] Figure 1 is a flowchart of a space-air-ground integrated satellite interconnection network routing method of the present invention;
[0033] Figure 2 is an example diagram of network status information of the present invention;
[0034] Figure 3 is a probability distribution diagram of routing selection actions of the present invention;
[0035] Figure 4 is an example diagram of the technical effects of the present invention. Detailed Embodiments
[0036] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0037] In at least one embodiment of the present invention, a space-air-ground integrated satellite interconnection network routing method is disclosed, as Figure 1 shown, including the following steps:
[0038] Step100, obtain network status information, where the network status information includes the working status data of the satellite node obtained by the intelligent routing agent from the status monitoring module of the satellite node itself, the link quality-related data obtained through the link quality detection device, and the topological structure information of the satellite network;
[0039] In an embodiment of the present invention, the working status data of the satellite node includes: the power of the satellite, signal strength, operating orbit parameters, etc., which can reflect whether the satellite node is working properly and its current operating condition;
[0040] The link quality-related data includes: information such as the bandwidth, delay, and packet loss rate of the link, which are crucial for evaluating the communication performance of the satellite link;
[0041] The topological structure information of the satellite network includes the connection relationships between each satellite node.
[0042] In an embodiment of the present invention, the obtained network status information is cleaned and filtered to remove outliers; then the data is normalized, for example, the battery percentage is normalized to the interval [0, 1], the link bandwidth is unified to the Mbps unit, and the Z-score is used to normalize the delay data. Subsequently, a graph neural network (GNN) is used to extract and aggregate features from the normalized data, and the feature representations of nodes and edges are learned through graph convolution operations to generate input data that can be used for generating routing policies.
[0043] Process the obtained network status information through a graph neural network to obtain the locally observed network status information;
[0044] In an embodiment of the present invention, the graph neural network uses operations such as graph convolution to aggregate the features of nodes and edges, and learns the network topology structure and attribute representations; let represent the network topology, where is the set of nodes, is the set of edges. For node , its feature representation is . Through the graph convolution calculation formula:
[0045] ;
[0046] where is the hidden representation of node in the th layer; is the hidden representation of node in the th layer; is the set of neighbor nodes of node ; is the weight matrix of the th layer; is the bias vector; is the non-linear activation function.
[0047] Step200, Routing decision, The intelligent routing agent autonomously learns and optimizes the routing policy based on the deep reinforcement learning model according to the network status and transportation requirements;
[0048] In an embodiment of the present invention, the deep reinforcement learning model adopts the Actor-Critic framework, which includes a policy network and a value network ;
[0049] The policy network selects the probability distribution of the routing action according to the state ; the value network estimates the state The expected long-term return under; the decision-making process can be modeled as an MDP (Markov Decision Process), that is, a five-tuple Among the five-tuples represents the state space, which contains all possible states in the network, represents the action space, which contains all routing decisions that the agent can execute, represents the state transition probability function, which defines the probability of transitioning to the next state s' after executing action a in the current state s, represents the reward function, which defines the immediate reward obtained by executing action a in state s, represents the discount factor, which is used to balance the importance of immediate rewards and future rewards; among them, the update formula of the policy network is:
[0050] ;
[0051] Among them, represents the parameters of the policy network The policy network selects routing actions according to the state The probability distribution of ; is the first learning rate; represents taking the gradient of the parameter ; is under the policy For the action Based on the state The logarithm of the selection probability of is the action value function; represents the value network For the state The estimated value of the expected long-term return under;
[0052] The update formula of the value network is:
[0053] ;
[0054] Among them, represents the parameters of the value network ; is the second learning rate; is the immediate reward; represents the second discount factor, The value range is 0-1; represents the estimated value of the expected long-term return of the value network for the next state ; represents the estimated value of the expected long-term return of the value network for the current state ; represents taking the gradient of the parameter ; is the next state, i.e., the state where the agent is in the current state takes an action and then transfers to the new state.
[0055] Step300, multi-agent collaboration. Through the federated learning mechanism, multiple intelligent routing agents exchange routing policy knowledge to achieve global collaborative optimization;
[0056] In an embodiment of the present invention, each agent uploads its local policy parameters to the central controller; the local policy parameters include: policy network weight matrix, bias vector, convolution kernel parameters, etc.; the central controller aggregates the policy parameters of each agent, updates the global policy , and then distributes it to each agent to update its local policy. Suppose there are agents in total. The formula for the central controller to aggregate the policy parameters is: .
[0057] Step400, resource scheduling optimization. The central controller optimizes the scheduling of network resources according to the decisions of the intelligent routing agents and the logistics transportation requirements;
[0058] In an embodiment of the present invention, the goal of resource scheduling optimization is to minimize the network load imbalance ; where is the utilization rate of link , represents the set of all links in the network, is the average utilization rate. At the same time, maximize the quality of service of critical services ; where is the weight of traffic flow , represents the set of all traffic flows in the network, is its quality of service metric.
[0059] Step500, policy update. The intelligent routing agent updates the routing policy through the policy gradient algorithm according to the reward feedback obtained by executing the action;
[0060] The policy gradient is defined as follows:
[0061] ;
[0062] where, represents the mathematical expectation, represents the policy performance objective function, represents the gradient with respect to the parameter , that is, under the policy , the logarithm of the selection probability of the action based on the state is taken and then the gradient with respect to the parameter is calculated The gradient is then multiplied by the action-value function , and finally the expectation of the result is taken under the policy ; The policy network parameters are updated through optimization algorithms such as stochastic gradient ascent ; The specific update formula is , where is the third learning rate.
[0063] In an embodiment of the present invention, an example of the above-mentioned space-air-ground integrated satellite interconnection network routing method is provided:
[0064] Suppose in a space-air-ground integrated satellite interconnection network scenario, there are 5 satellite nodes and a set of 8 links ;
[0065] State awareness:
[0066] As Figure 2 shown, which includes nodes, node state characteristics, and example data of link quality;
[0067] Example of graph convolution operation:
[0068] Suppose , is , , is the ReLU function. For node , its neighbor nodes , , , .
[0069] ;
[0070] Action selection:
[0071] Let the current state be the network topology and the states of each node. The transportation demand is to transmit data from to . The policy network gives the probability distribution of different routing actions (such as passing through and then to , passing through and then to , etc.) as Figure 3 shown;
[0072] Reward calculation:
[0073] The value network estimates that the expected long-term return in the current state is 0.7; Let the learning rate , For the action of 0.8;
[0074] Policy network update: ;
[0075] Value network update: ;
[0076] Load balancing optimization:
[0077] Assume there are 5 links ( ), link utilization rate , , , , , average utilization rate ;
[0078] Key business service quality optimization:
[0079] There are 3 traffic flows ( ), traffic flow weight , , , service quality metric , , ;
[0080] Multi-objective optimization:
[0081] Network load imbalance degree:
[0082] ;
[0083] Key business service quality:
[0084] ;
[0085] Using the weighted method, let , the single-objective optimization problem is:
[0086] ;
[0087] Federated learning parameter aggregation:
[0088] Suppose there are 3 agents ( ), their local policy parameters , , . The central controller aggregates the policy parameters:
[0089] ;
[0090] Examples of technical effect data are shown as Figure 4 shown;
[0091] In an embodiment of the present invention, the satellite interconnection network environment is complex and variable. To better adapt to the complex and variable environment of the satellite interconnection network, a space-air-ground integrated satellite interconnection network routing method further includes:
[0092] S101, adaptive network state perception. The Adaptive Intelligent Routing Agent (AIRA) perceives the node state, link quality, and network topology changes in real time through multi-source heterogeneous information fusion;
[0093] In an embodiment of the present invention, algorithms such as Kalman filtering are used to dynamically estimate and predict the network state to improve the accuracy and real-time performance of perception; the state transition model is:
[0094] ;
[0095] where represents the state at time which is obtained from the state at time the control input at time and the process noise through the non-linear function .
[0096] The observation model is:
[0097] ;
[0098] where represents the observation value, represents the observation noise, and the function represents the non-linear function.
[0099] The prediction step of Kalman filtering is:
[0100] ;
[0101] where represents the predicted state at time represents the predicted state at time represents the control input at time represents the non-linear function.
[0102] ;
[0103] where represents The predicted covariance at a moment represents the state transition matrix denotes the estimated covariance at a moment represents the transpose matrix of the state transition matrix represents the process noise covariance
[0104] The update step is as follows
[0105] ;
[0106] where represents the Kalman gain represents the observation matrix represents the transpose of the observation matrix represents the observation noise covariance
[0107] ;
[0108] where represents the more accurate estimated state at a moment obtained after update represents the more accurate estimated state at a moment obtained after update represents the Kalman gain represents the observed value
[0109] ;
[0110] where represents the updated estimated covariance represents the identity matrix represents the Kalman gain represents the observation matrix denotes the predicted covariance at a moment
[0111] s102, intelligent routing policy adaptation, AIRA is based on a deep reinforcement learning model, and autonomously learns and optimizes the routing policy according to the perceived network state
[0112] In an embodiment of the present invention, a reinforcement learning method is adopted to adapt to continuous state and action spaces, and meta-learning and transfer learning mechanisms are introduced to accelerate the policy learning and adaptation process
[0113] The value function is defined as
[0114] ;
[0115] where represents at The value function at the moment when the state is and the action taken is ; denotes the mathematical expectation; is the immediate reward at the moment; is the third discount factor, whose value ranges from to . The closer it is to , the more the future rewards are emphasized. The closer it is to , the more the immediate rewards are concerned; is the state at the moment, is the possible action at the moment, is the maximum Q value that can be obtained among all possible actions at the moment. denotes the maximum Q value that can be obtained among all possible actions at the moment. ;
[0116] In the Deep Deterministic Policy Gradient (DDPG), the update formula of the policy network is:
[0117] .
[0118] Among them, is the parameter of the policy network ; is the fourth learning rate; is the gradient of the value function with respect to the parameter .
[0119] The update formula of the value network is as follows:
[0120] ;
[0121] Among them, is the parameter of the value network ; is the fifth learning rate; is the immediate reward; is the fourth discount factor; denotes the next state, is the action taken by the target policy network in the state , is the parameter of the target policy network; is the value of the target value network in the state and the action taken is ; is the value of the current value network in state and takes action ; is the value function with respect to the parameter .
[0122] s103, Distributed multi-agent collaboration. AIRA constructs a distributed trust mechanism through blockchain technology to achieve secure and reliable information sharing and collaborative decision-making;
[0123] Adopt smart contracts and consensus algorithms, such as Proof of Work (PoW), to incentivize cooperation and reach an agreement among AIRAs;
[0124] Introduce the federated learning framework. AIRA trains model parameters locally, and the central controller aggregates the model parameters of each AIRA, updates the global model, and then distributes it to AIRA to update the local model;
[0125] In an embodiment of the present invention, the steps of distributed multi-agent collaboration are as follows:
[0126] The global objective is to minimize the weighted loss function:
[0127] ;
[0128] where, represents the global model parameters, represents the number of AIRAs, represents the th importance weight of the th AIRA; represents the loss function of the
[0129] th AIRA.
[0130] ; ;
[0131] where, represents the hash value obtained by performing a hash calculation on ; represents the block header information, which contains the metadata of the block, such as version number, timestamp, hash value of the previous block, etc.; represents the target hash value, which defines a difficulty threshold. Only when the calculated hash value is less than this target hash value, the node is considered to have successfully completed the proof of work and obtained the right to record accounts.
[0132] s104, Autonomous Fault Diagnosis and Recovery. AIRA monitors and locates faults and anomalies in the network in real time through anomaly detection and fault diagnosis algorithms;
[0133] In one embodiment of the present invention, a fault propagation model based on a graph neural network is adopted to learn the propagation law of faults in the network topology and predict the scope and degree of fault impact;
[0134] The fault diagnosis model is defined as:
[0135] ;
[0136] Where, represents the conditional probability distribution of the fault state under the condition of a given network topology graph ; represents the network topology graph, which describes the connection relationship between each node in the network; represents the diagnostic function parameterized by GNN, which is determined by the parameters of the graph neural network and is used to predict the probability distribution of the fault state based on the network topology graph ;
[0137] The fault propagation follows a linear propagation model, and the calculation formula is as follows:
[0138] ;
[0139] Where, represents the fault state of node ; represents the propagation weight between nodes and ; represents the set of neighbor nodes of node ; represents the bias parameter.
[0140] s105, Policy Evaluation and Optimization. AIRA evaluates the effectiveness and optimization degree of the intelligent routing policy according to business requirements and service quality indicators;
[0141] In one embodiment of the present invention, multi-objective reinforcement learning algorithm is adopted for policy evaluation and optimization to balance multiple performance indicators such as delay, throughput, and reliability;
[0142] Meanwhile, game theory and mechanism design theory are introduced to optimize the incentive and constraint mechanisms in the multi-agent cooperation process;
[0143] The calculation formula of the immediate reward function is as follows:
[0144] ;
[0145] Among them, represents the immediate reward at moment, , , respectively represent the first, second, and third weight coefficients, represents the delay index at represents the throughput index at represents the reliability index at
[0146] To solve the multi-objective optimization problem, the Pareto optimal method is adopted to find all non-dominated solution sets . In the multi-objective optimization problem, a non-dominated solution refers to a solution for which there is no other solution that strictly dominates it in all objectives. For any two solutions , there does not exist that strictly dominates or that strictly dominates . This means that for any two solutions in the set , neither can simultaneously outperform the other in all performance metrics. These non-dominated solutions form the Pareto front, representing the set of optimal choices when making trade-offs among multiple objectives.
[0147] In an embodiment of the present invention, in order to further optimize the routing performance of the space-air-ground integrated satellite interconnection network and improve the overall efficiency of the network, a routing method for the space-air-ground integrated satellite interconnection network further includes the following steps:
[0148] s201, the Adaptive Intelligent Routing Agent AIRA is deployed on each network node; adaptive network state awareness is achieved through multi-source heterogeneous information fusion;
[0149] In an embodiment of the present invention, in the deep reinforcement learning model: is the parameter of the policy network , represents the state, represents the routing action; is the parameter of the value network , is the state. The decision-making process is still modeled as a Markov decision process (MDP), that is ;
[0150] The multi-source heterogeneous information fusion calculation formula is as follows:
[0151]
[0152] Among them, represents the fused network state representation, is a non-linear activation function, is the number of information sources, is the weight coefficient of the k-th information source, represents the k-th information source feature extraction function;
[0153] Policy network update formula:
[0154] ;
[0155] Among them, and respectively represent the updated and pre-updated policy network parameters, is the sixth learning rate, is the policy gradient;
[0156] The fault propagation still follows the linear propagation model;
[0157] ;
[0158] Among them is the fault state of node , is the node and propagation weight between, represents the set of neighbor nodes of node ; is the bias.
[0159] S202, the central controller, communicates with each AIRA; aggregates the status information and decision-making information of AIRA, and realizes distributed multi-agent collaboration through mechanisms such as federated learning;
[0160] In an embodiment of the present invention, each agent uploads its local policy parameters to the central controller, and the central controller aggregates the policy parameters of each agent to update the global policy . The global goal is to minimize the weighted loss function.
[0161] Global model parameter update formula:
[0162]
[0163] Among them, represents the updated global model parameters, represents the th AIRA data volume, represents the total data volume of all AIRA, Represents the local model parameters of the th AIRA;
[0164] Global weighted loss function:
[0165]
[0166] Among them, Represents the global model parameters, Represents the total data volume of all AIRAs, Represents the th AIRA weight, Represents the th AIRA's loss function based on the global model parameters ; Represents minimizing the optimization of the parameter ;
[0167] s203, a blockchain network, which is used as a secure and trusted distributed communication and storage platform among AIRAs;
[0168] In PoW, is the hash value obtained by performing a hash calculation on ; is the block header information; is the target hash value, and the node calculates the hash value to satisfy in order to complete the proof of work and obtain the right to record accounts.
[0169] Blockchain consensus difficulty adjustment formula:
[0170]
[0171] Among them, is the target hash value, represents the target hash value of the previous block, represents the actual generation time of the previous block, represents the expected block generation time;
[0172] s204, a policy evaluation and optimization module;
[0173] According to the business service quality and network performance metrics, the immediate reward function definition is still:
[0174] ;
[0175] Among them, represents the immediate reward at the moment, , , respectively represent the first, second, and third weight coefficients, represent the time delay index at the moment, represent the throughput index at the moment, represent the reliability index at the moment;
[0176] Cumulative reward calculation formula:
[0177]
[0178] where represents the cumulative reward starting from time t, is the fifth discount factor, represents the immediate reward at time t + k;
[0179] To solve the multi-objective optimization problem, the Pareto optimal method is adopted to find all non-dominated solution sets , where represents the set of all non-dominated solutions. For any two solutions , there does not exist strictly superior to or strictly superior to cases. Methods such as multi-objective reinforcement learning and game theory are used to evaluate and optimize the intelligent routing strategy of AIRA. Balance multiple performance objectives and improve the effectiveness and global optimality of the strategy.
[0180] s205, the fault diagnosis and recovery module, constructs a fault propagation model based on a graph neural network;
[0181] Learn the causal relationship and propagation law of faults in the network. Achieve fault tracing, impact prediction, and recovery decision-making. Minimize service interruption and performance loss caused by faults.
[0182] Graph neural network layer update formula:
[0183]
[0184] where represents the hidden state of node i in layer, is a non-linear activation function, represents the set of neighbor nodes of node i, is a normalization constant, and are respectively the weight matrix and bias vector of the layer, represents the hidden state of node j in the layer;
[0185] Fault impact range prediction formula:
[0186]
[0187] Among them, represents the impact range of fault f, represents the probability of node i failing under the condition that fault f occurs, represents the importance weight of node i, and N represents the total number of nodes in the network;
[0188] S206, the human-computer interaction interface, provides an intuitive and friendly visualization interface; displays key information such as the global network status, routing strategy, and performance metrics; supports manual intervention and policy adjustment, and realizes the interpretability and controllability of the intelligent routing system.
[0189] Interpretability score calculation formula:
[0190]
[0191] Among them, represents the interpretability score of action a, , , respectively represent the fourth, fifth, and sixth weight coefficients, represents the clarity of the causal relationship of action a, represents the transparency of action a, represents the visualization degree of action a;
[0192] Each module realizes interconnection, interoperability, and data sharing through a standardized interface and works collaboratively; forms an adaptive, intelligent, safe, and efficient integrated satellite interconnection networking routing solution to support the intelligent operation and optimization management of the global logistics network.
[0193] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A satellite internet routing method for integrated space-ground-air communication, characterized in that: include: Acquire network status information, the network status information includes the working status data of the satellite node obtained by the intelligent routing agent from the status monitoring module of the satellite node itself, the link quality related data obtained by the link quality detection device, and the topological structure information of the satellite network; Routing decision, intelligent routing agent is based on deep reinforcement learning model, autonomously learning and optimizing routing strategy according to network status and transportation demand; Multi-agent collaboration: Through the federated learning mechanism, multiple intelligent routing agents exchange routing strategy knowledge to achieve global collaborative optimization; Resource scheduling optimization: the central controller optimizes the scheduling of network resources based on the decisions of the intelligent routing agent and the logistics and transportation requirements; Policy update,The intelligent routing agent updates the routing strategy through the policy gradient algorithm based on the reward feedback obtained by executing actions.
2. The method for routing satellite internet networks integrating space, air and ground according to claim 1, characterized in that: The acquired network status information is cleaned and filtered to remove outliers; then the data is standardized; the graph neural network is used to extract and aggregate features of the standardized data, and the feature representation of nodes and edges is learned through graph convolution operations to generate input data for routing strategy generation.
3. The method for routing satellite internet networks integrating space, air and ground according to claim 1, characterized in that: The deep reinforcement learning model in routing decision-making adopts the ActorCritic framework, which includes a policy network and value network ; The policy network is based on the state Select routing action The probability distribution of the value network estimates the state The expected long-term return.
4. The method for routing satellite internet networks integrating space, air and ground according to claim 1, characterized in that: In multi-agent collaboration, each agent uploads local strategy parameters To the central controller; the central controller aggregates the policy parameters of each agent and updates the global policy , and then distribute it to each agent to update the local strategy.
5. The method for routing satellite internet networks integrating space, air and ground according to claim 1, characterized in that: In resource scheduling optimization, resource scheduling optimization is to minimize the imbalance of network load and maximize the service quality of key businesses.
6. The method for routing satellite internet networks integrating space, air and ground according to claim 1, characterized in that: Also includes: Adaptive network state awareness, adaptive intelligent routing agent AIRA; Intelligent routing strategy adaptation: AIRA is based on a deep reinforcement learning model and can autonomously learn and optimize routing strategies based on the perceived network status. Distributed multi-agent collaboration: AIRA builds a distributed trust mechanism through blockchain technology to achieve secure information sharing and collaborative decision-making; Autonomous fault diagnosis and recovery, AIRA uses anomaly detection and fault diagnosis algorithms to monitor and locate faults and anomalies in the network in real time; Policy evaluation and optimization,AIRA evaluates the effectiveness and optimization of intelligent routing policies,based on business requirements and service quality indicators.
7. The method for routing satellite internet network integrating space, air and ground according to claim 6, characterized in that: The Kalman filter algorithm is used to dynamically estimate and predict the network status to improve the accuracy and real-time performance of perception.
8. The method for routing satellite internet networks integrating space, air and ground according to claim 6, characterized in that: A deep reinforcement learning model is adopted to adapt to the continuous state and action space, and meta-learning and transfer learning mechanisms are introduced to accelerate the strategy learning and adaptation process.
9. The method for routing satellite internet networks integrating space, air and ground according to claim 6, characterized in that: A fault propagation model based on graph neural network is used to learn the propagation law of faults in network topology and predict the scope and extent of fault impact.
10. An air-ground integrated satellite internet network routing system, characterized in that: A method for executing an air-ground integrated satellite internet network routing method as described in any one of claims 1 to 9, comprising: State perception module: responsible for sensing satellite node status and link quality, and then obtaining network status information, and cleaning, filtering and standardizing the obtained network status information to provide accurate and reliable input data for routing strategy generation; Routing strategy generation module: Based on the network status information obtained by the state perception module, the network status information is processed and input into the deep reinforcement learning model, and the routing strategy is autonomously learned and optimized according to the network status and transportation demand; Resource allocation module: dynamically adjust network resource allocation based on the decisions of intelligent routing agents and logistics transportation needs; Policy optimization module: uses intelligent learning algorithms to optimize routing strategies and improve system adaptability; the policy optimization module includes an intelligent learning unit that uses reinforcement learning and deep learning algorithms to learn and optimize routing strategies; the policy update unit updates the optimized routing strategies to the routing strategy generation module to ensure the real-time effectiveness of the routing strategies; Data transmission module: completes data transmission tasks based on optimized routing strategies.
Citation Information
Patent Citations
Seamless credible cross-domain routing system of heterogeneous convergence network and control method of seamless credible cross-domain routing system
CN113660668A
Low earth orbit satellite network flow routing method based on multi-agent reinforcement learning
CN117041129A
Distributed reliable routing design method and system for satellite internet
CN117527688A
Two-hop learning deep reinforcement learning routing strategy method based on dijkstra assistance
CN118175083A
Low earth orbit satellite constellation routing method, system and device and storage medium
CN119966885A
Cited By
Satellite node reliability evaluation method in F6G low earth orbit satellite optical network
CN122052880A