Space-air network layered cooperation route optimization method and system

Through the layered reinforcement learning routing optimization method, combined with dynamic reputation value mechanism and multi-agent algorithm, the routing decision-making problem of drone cluster communication in aerospace information network is solved, and communication efficiency and stability are improved.

CN120264376APending Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM +2
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510431954.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In aerospace information network, traditional routing algorithms are difficult to track and adapt to the rapid changes in communication needs of drone clusters in real time, resulting in unreasonable allocation of satellite resources and insufficient communication efficiency and stability.

Method used

The hierarchical reinforcement learning routing optimization method is adopted to monitor the real-time state of satellites and drone clusters through distributed proxy nodes, introduce a dynamic reputation value mechanism, and combine deep Q networks and multi-agent near-end strategy optimization algorithms to optimize routing decisions across clusters and intra-clusters.

Benefits of technology

It improves the success rate and link utilization rate of satellite and drone collaborative communication, reduces request delay, and enhances the adaptability and robustness of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264376A_ABST
    Figure CN120264376A_ABST
Patent Text Reader

Abstract

The invention discloses an aerospace network hierarchical cooperation routing optimization method and system, and relates to the field of aerospace information network communication, and the method comprises the steps: constructing a system model comprising an unmanned plane, satellite nodes and a communication link, defining link attributes, building a routing optimization objective function, introducing a reputation value-based multi-agent cooperation mechanism, and carrying out the routing optimization of the unmanned plane, the satellite nodes and the communication link. A hierarchical reinforcement learning framework is designed, inter-cluster decision and intra-cluster decision are divided, an upper layer selects an inter-cluster path according to global information, a lower layer plans an unmanned aerial vehicle path and defines a state, an action space and a reward function according to local information, an intelligent agent interacts with an environment to collect empirical data, and optimal routing communication is selected in combination with a real-time reputation value after training. Adjustment can be carried out according to network changes; according to the invention, the cooperative reliability is improved through a reputation value mechanism, the decision complexity is reduced through hierarchical learning, the routing performance is optimized, the network robustness, expandability and adaptability are enhanced, and the cooperative communication efficiency of the unmanned aerial vehicle and the satellite and the system stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aerospace information network communication technology, and in particular to a routing optimization method suitable for a cooperative communication network between unmanned aerial vehicles and satellites. Background Art

[0002] With the continuous evolution of 5G and higher communication technologies, aerospace information networks, as an important part of the modern communication field, are booming at an unprecedented rate. The integrated application of satellites and drones in aerospace information networks is becoming more and more extensive. Satellites provide solid support for long-distance information transmission with their powerful wide-area coverage capabilities and stable communication links; drones have unique advantages in near-Earth space with their high flexibility, maneuverability and rapid deployment characteristics. Under such a complex network architecture, routing optimization has become an extremely challenging task. The high maneuverability of drones makes the network topology structure in continuous and rapid changes. Traditional routing algorithms are difficult to track and adapt to such frequent changes in real time, resulting in delayed or invalid routing decisions. Satellite resources are relatively limited. How to reasonably allocate bandwidth, power and other resources among the communication needs of many drone clusters to ensure the efficiency and stability of communication is a key issue that needs to be solved urgently.

[0003] The multi-agent system can simulate the autonomous decision-making and collaboration process of multiple intelligent individuals and adapt to the distributed characteristics of the network; hierarchical reinforcement learning helps to decompose complex routing decision problems into multiple levels and reduce the complexity of the problem.

[0004] However, in practical applications, the reliability of collaboration between intelligent agents and their adaptability to dynamic environments still need to be improved. To overcome these problems, a reputation value mechanism is introduced, which aims to guide the intelligent agents to make better routing decisions through a comprehensive evaluation of the intelligent agent's behavior, thereby improving the performance of the entire network, effectively solving the above-mentioned routing optimization problem, and ensuring the efficient and stable operation of collaborative communication between satellites and drone clusters. Summary of the invention

[0005] In response to the problems raised in the above background technology, the present invention provides a method and system for optimizing hierarchical collaborative routing in an air-space network to improve the communication success rate and link utilization rate of routing decisions in a collaborative communication network between satellites and unmanned aerial vehicles, while reducing request delays.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] On the one hand, the present invention provides a space-to-air network hierarchical collaborative routing optimization method, comprising:

[0008] Build a collaborative network model that includes low-earth orbit satellite nodes and clustered unmanned aerial vehicle (UAV) swarms. Among them, the satellite nodes form a backbone network for cross-cluster communication, and the UAV swarms are dynamically clustered according to mission requirements. Each cluster contains a cluster head node and several member nodes. Intra-cluster communication is carried out through an ad-hoc network, and data is forwarded between clusters through satellite nodes;

[0009] Monitor the real-time status of satellites and UAV clusters through distributed agent nodes, including: cross-cluster link status, intra-cluster link status, and dynamic reputation values. The link status includes: for each agent in cluster d, there is a set of cross-cluster links E

[0010] ={e d,d}, i ∈ N i,j , j ∈ N d , d ≠ d'}, for each e d ∈ E i,j Define a dynamic reputation value function d,d′

[0011] The composition of the dynamic reputation value function includes: link bandwidth utilization Normalized delay And the number of historical failed nodes fail_count, where the exponential term Ensures that delay-sensitive services preferentially select low-latency links, and the coefficients meet the constraint condition α + β + γ = 1;

[0012] Set the link reputation value update rule formula C i,j (t + 1) = ρ·C i,j (t) + ΔC i,j (t), where ρ is the attenuation coefficient, and its value range is 0 to 1.

[0013] Optionally, for the hierarchical feature extraction and state modeling, extract global cross-cluster link features and construct the upper-layer agent state space, specifically including:

[0014] Observe the remaining bandwidth BW avail Of the cross-cluster link, the historical average link delay delay avg , the satellite node load rate load sat And the candidate cross-cluster link reputation value C i,j (t);

[0015] According to the formulaCalculate the satellite load rate of the next-hop satellite node of the candidate link, where Data processed Represents the amount of data processed by the satellite node within a specific time period, and Capacity max Represents the maximum data processing capacity of the satellite node.

[0016] Optionally, for the hierarchical feature extraction and state modeling, the intra-cluster link features are extracted to construct the state space of the lower-layer agents, which specifically includes:

[0017] The signal-to-noise ratio of the intra-cluster link, link quality , the remaining energy of the node, energy residual , the states of neighbor nodes, neighbor states and the cross-cluster link reputation value provided by the satellite

[0018] Optionally, for implementing the hierarchical multi-agent reinforcement learning routing decision, the upper-layer agent (LEO satellite) adopts the Deep Q-Network (DQN) algorithm. By periodically collecting the global state, a cross-cluster state vector is constructed. The output action of each upper-layer agent corresponds to a set of candidate cross-cluster forwarding node sets, denoted as

[0019] According to the formula r U = ω1e -delay + ω2C selected + ω3(1 - sat_load), the reward function is calculated, and at the same time, the model weights are iteratively optimized using the reward feedback.

[0020] Optionally, for implementing the hierarchical multi-agent reinforcement learning routing decision, the lower-layer agent (UAV) adopts the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm to construct the intra-cluster state vector. The cluster head selects the next-hop node within the cluster, and the action space can be expressed as a a,U =(Path k ), where Path k represents the k-th alternative routing path selected by the lower-layer agent between each pair of nodes within each cluster. The policy gradient updates the Actor-Critic network parameters through the advantage function A(s L , a L ) = Q(s L , a L ) - V(s L ).

[0021] Optionally, the trained hierarchical routing policy is deployed to the actual network. The satellite modifies the inter-satellite link routing table entries through control commands to specify high-reputation value paths. The satellite periodically broadcasts the cross-cluster link reputation value, and the lower-layer cluster head incorporates it into the state space to evaluate the feasibility of cross-cluster requests. Through the satellite reputation feedback, the UAV cluster head dynamically adjusts the forwarding rules of member nodes to respond to local state changes. At the same time, the cluster head feeds back the intra-cluster resource state for the upper layer to adjust the global load balancing policy;

[0022] Continuously collect network topology changes (such as node movement, link interruption), trigger the recalculation of reputation values. If it is detected that the link state change exceeds the threshold, repeat the routing request forwarding judgment, update the routing policy and reissue the flow table;

[0023] Judge whether the task is completed or the network enters a stable state. If continuous optimization is required, return to the above process; otherwise, terminate the routing adjustment process.

[0024] On the other hand, the present invention also provides an airspace network hierarchical collaborative routing optimization system, including:

[0025] Network modeling and initialization module: The satellite constructs a global topology snapshot of the UAV cluster through on-board sensors and inter-satellite links, and provides initial weight configuration suggestions for the clustering algorithm;

[0026] Feature information acquisition module, used to acquire the feature information of each UAV node in the UAV cluster; the feature information includes the geographical position coordinates and speed of the UAV node;

[0027] Status monitoring and acquisition module, which monitors the cross-cluster link status (bandwidth, delay, load) and intra-cluster link status (SNR, energy, queue) in real time through distributed proxy nodes, and synchronizes them to the global status library;

[0028] Dynamic reputation value management module, which collects the current data traffic on the link in real time through the SDN controller, records the time difference between packet sending and receiving through timestamps, calculates the end-to-end delay, counts the number of transmission failures of the link in the recent time window, stores global parameters such as historical reputation values, node load rates, and topological structures in the global status library, dynamically quantifies the link performance, preferentially selects high-reputation paths, and adapts to network topology changes at the same time to ensure communication efficiency and stability;

[0029] Hierarchical feature extraction module, the global feature extraction sub-module uses a deep neural network to extract features such as the reputation distribution and load balance degree of satellite links, and the local feature extraction sub-module aggregates local features such as node energy and link quality within the cluster through GNN message passing;

[0030] Hierarchical routing decision module, the upper-layer decision sub-module (satellite) runs the DDQN algorithm, inputs the global feature vector, and outputs the cross-cluster optimal path selection action; the lower-layer decision sub-module (UAV cluster) runs the MAPPO algorithm, inputs the local feature vector, and outputs the intra-cluster routing action (multi-hop or request satellite forwarding);

[0031] Routing policy deployment module, converts the decision path into a flow table instruction, the satellite issues the inter-satellite link priority and bandwidth reservation policy, and the UAV cluster head adjusts the forwarding rules of member nodes;

[0032] The network status monitoring and re-planning module continuously detects topological changes such as node movement and link interruption, triggers the recalculation of reputation values and the update of routing policies, and starts dynamic re-planning when the threshold is exceeded;

[0033] The resource cooperation feedback module enables the satellite to broadcast the cross-cluster link reputation value to the cluster heads, and the cluster heads feedback the intra-cluster load and energy status to achieve global-local resource cooperation optimization;

[0034] The UAV cluster communication module is used for each UAV node in the UAV cluster to communicate according to the cluster structure.

[0035] In summary, the present invention mainly has the following beneficial effects:

[0036] 1. A cross-cluster communication backbone network is constructed through the satellite to provide network information for the UAV clusters and globally record the link status of the UAV clusters; according to the routing requests, they are divided into cross-cluster and intra-cluster routing requests, and the interaction between the upper-layer intelligent agent (satellite) and the lower-layer intelligent agent (UAV cluster) is realized through state sharing, action triggering, reputation value feedback and dynamic path adjustment; the satellite performs cross-cluster path selection based on the Deep Q-Network (DQN), inputs the global state (cross-cluster link reputation value, bandwidth, delay, load), and selects the next-hop node (satellite or cluster head) by maximizing the Q value, and the reward function encourages low-delay, high-reputation, and low-load paths; the cluster heads elected by the UAV cluster network are responsible for intra-cluster routing decisions. The cluster heads select intra-cluster paths based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, input the local state (SNR, remaining energy, queue, satellite reputation feedback), select the next-hop within the cluster, and the reward function optimizes the bandwidth, energy consumption, and cross-link reputation. The policy gradient is updated through the advantage function, and the cluster heads feedback the intra-cluster link status to the satellite for the upper layer to update the global load balancing parameters.

[0037] 2. By introducing the reputation value mechanism, the present invention dynamically evaluates and reflects the historical performance and reliability of the links between the satellite and the UAVs, can enhance the network's adaptability and robustness, and provides a quantitative standard to guide cross-cluster routing selection by comprehensively considering factors such as bandwidth utilization, transmission delay, and the number of link failures, thereby improving the communication success rate and link utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The flowchart of a method for optimizing hierarchical cooperative routing in an air-space network according to the present invention;

[0039] Figure 2 It is a schematic diagram of the routing decision execution for implementing routing requests according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0041] The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of protection of the present invention. The conditions in the embodiments can be further adjusted according to specific conditions. Any simple improvement to the method of the present invention under the premise of the concept of the present invention falls within the scope of protection required by the present invention.

[0042] Reference Figure 1 and Figure 2 The present invention provides a method and system for optimizing the hierarchical cooperative routing of an airspace network, providing an efficient, reliable, and adaptive routing solution for the airspace network to enhance the stability and anti-interference ability of the network, improve the communication success rate and link utilization rate, and optimize network resource allocation and load balancing.

[0043] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail in conjunction with the accompanying drawings and specific embodiments.

[0044] Refer to Figure 1 An airspace network hierarchical cooperative routing optimization method includes:

[0045] Step 1: Set up a satellite-UAV network. The UAV network is composed of a group of UAV clusters D = {1, 2,..., d}. In each cluster d, there is a group of UAV nodes N d = {1, 2,..., n}. The inter-cluster link is denoted as The intra-cluster link is denoted as

[0046] Step 2: The member node initiates a route n. The set of all links on the path from the source node sour to the destination node des is denoted as p n (sour, des);

[0047] The path can be decomposed into [(sour, node1), (node1, node2),..., (nodeK, des)]. The routing request is divided into cross-cluster routing and intra-cluster routing according to whether the destination node is in the same cluster as the destination node;

[0048] Step 3: For cross-cluster routing requests, the source cluster head sends a routing request to the satellite, carrying the destination cluster ID, monitors the real-time status of the satellite and drone clusters through distributed proxy nodes, and uses a deep Q network (DQN) to select the next hop (satellite or cluster head drone).

[0049] The main purpose of the present invention is to form an upper-level intelligent body through low-orbit satellites. The satellite is responsible for controlling the status of nodes in the global network and the network topology to achieve the establishment and maintenance of information for the entire network. At the same time, the satellite, as an upper-level intelligent body, adopts a DQN network to make cross-cluster routing decisions and select the next-hop routing. In the present invention, the satellite node will establish a channel for interactive information transmission with the drone network plane, be responsible for receiving the node status information sent by the drone network plane, and delegate the sub-goal of the service transmission plan to the drone network. The satellite selects the next hop (satellite or destination cluster head) based on the reputation value and status. If the next hop is a satellite, it is relayed to the target satellite; if it is the target cluster head, the satellite forwards the request to the destination cluster, and the lower-level agent of the destination cluster optimizes the path within the cluster to complete end-to-end transmission;

[0050] Calculate the four sub-states that constitute the state space of the upper-level agent (satellite): the reputation value C of the candidate cross-cluster link i,j (t), link available bandwidth BW avail 、Historical average delay avg And the satellite node load rate load sat ;

[0051] First, the reputation value C of the candidate cross-cluster link i,j (t) can be expressed by the formula Calculated;

[0052] in, represents the link bandwidth utilization, and the cross-cluster link bandwidth between cluster d and cluster d' is defined as BW used Indicates the bandwidth actually being used by the candidate cross-cluster link at the current moment. The current transmission rate can be calculated by the amount of data transmitted on the link and the time, which is the current bandwidth, BW total It represents the total bandwidth of the candidate inter-cluster link, that is, the maximum bandwidth that the link can provide under ideal conditions without interference and loss.

[0053] represents the result of exponential transformation of delay, where delay represents the transmission delay of the current candidate cross-cluster link. The calculation formula can be expressed as It consists of two parts: processing delay and transmission delay τ is a time constant used to adjust the change rate of the exponential function, and its value is usually determined according to the actual network situation and experience;

[0054] fail_count represents the number of consecutive link failures. When a link failure event occurs, the value of fail_count will be automatically incremented by 1. If the link successfully completes data transmission in subsequent data transmissions, the value of fail_count will be reset to 0;

[0055] The parameters α, β, and γ are weight parameters used to balance the impacts of link bandwidth utilization, latency, and the number of consecutive link failures on the cross-cluster link reputation value, and each parameter meets the constraint condition α + β + γ = 1.

[0056] Step 4: Update the cross-cluster link reputation value

[0057] The present invention takes the link reputation value as one of the state space features of the upper-layer agent (satellite) and the lower-layer agent (drone), which affects the routing decisions of the upper and lower-layer agents. The upper-layer agent is responsible for selecting cross-cluster paths (such as between cluster heads or through satellite forwarding), and its core objective is to optimize the global network performance (such as low latency, high reliability). The cross-cluster link reputation value provides link reliability, load balancing, and dynamic adaptability for the upper-layer agent by dynamically quantifying the historical performance of the link (bandwidth utilization, latency, failure times); the lower-layer agent is responsible for in-cluster path optimization, and the action space includes requesting satellite forwarding (i.e., cross-cluster actions). The cross-cluster link reputation value supports cross-cluster request evaluation and resource trade-off. The link reputation value broadcast from the upper-layer agent can help the lower layer weigh cross-cluster and in-cluster paths in local decision-making;

[0058] According to the formula C i,j (t + 1) = ρ·C i,j (t) + ΔC i,j (t) to update the cross-cluster link reputation value, where ρ is the decay coefficient, and its value range is 0 to 1, which is used to prevent the historical value from becoming obsolete.

[0059] Step 5: Collect local state data and perform in-cluster routing and forwarding

[0060] The main objective of the present invention is to use the drone swarm clustering network for task partitioning. Packets sent from the source node can select the optimal path to reach the target node through forwarding by different cluster groups. The cluster head drone, as the lower-layer agent, uses the multi-agent proximal policy optimization (MAPPO) algorithm for routing decision-making. The agents share policy network parameters to improve training efficiency, but at the same time maintain independent value functions to capture local states;

[0061] Calculate four sub-states that make up the state space of the lower-layer agent (drone): the signal-to-noise ratio SNR of the in-cluster link, the remaining energy energy of the noderesidual , the status of neighbor nodes states and the cross-cluster link reputation value provided by the satellite

[0062] Among them, the signal-to-noise ratio SNR can be calculated by the formula to obtain, P s represents the signal power, P n represents the noise power;

[0063] For the intra-cluster routing request, the upper layer provides a sub-goal for the lower layer agent, the cluster head UAV.

[0064] Step 6: Determine whether the intra-cluster multi-hop condition is satisfied

[0065] The intra-cluster multi-hop data forwarding considers multiple factors, such as the connection relationship between nodes, signal strength, energy consumption, etc., and constructs the intra-cluster multi-hop condition from aspects such as path reachability, signal quality, and energy constraint as follows:

[0066]

[0067] Among them, represents node and whether there is a connection between them. When , it means there is a connection between the two nodes; when , it means there is no connection between the two nodes;

[0068] represents the signal-to-noise ratio of the link between node and , SNR th is the set signal-to-noise ratio threshold. Only when the signal-to-noise ratio of the link is greater than or equal to this threshold, is it considered that the link can be used for data transmission. represents the energy consumed when node transmits to node , E max is the maximum energy allowed to be consumed by the entire multi-hop path.

[0069] Step 7: Update the intra-cluster link status

[0070] Update the intra-cluster link signal-to-noise ratio. The signal-to-noise ratio update formula at time t+1 is Obtain ΔP through real-time monitoring and measurement s,i,j and ΔP n,i,j ;

[0071] Update the remaining energy of the node. The formula is expressed as E i (t+1) = E i (t) - ΔEi , ΔE i represents the energy consumed by node i due to operations such as data transmission and signal processing during the time period from t to t + 1. Affected by various factors such as the data transmission volume, transmission power, and processing time, it can be expressed as ΔE i = ΔE trans + ΔE proc , where ΔE trans = P t × T represents the energy consumption for data transmission, with the transmission power being P t , and the transmission time being T; ΔE proc = P proc × T proc represents the energy consumed for signal processing, with the processing power being P proc , and the processing time being T proc ;

[0072] Update the queue length. The update formula for the queue length of node i at time t + 1 is Q i (t + 1)= Q i (t)+ ΔQ arrive - ΔQ transmit , ΔQ arrive is the amount of data newly arriving at node i, and ΔQ transmit is the amount of data successfully transmitted.

[0073] Step 8: Deliver the data to the target node and update the global reputation value library

[0074] The overall method of the UAV swarm communication method based on on-demand weighted clustering and clustering is divided into two parts, including the cross-cluster routing decision of the upper-layer agent (satellite) and the intra-cluster routing decision of the lower-layer agent (UAV). In the cross-cluster routing decision part, the cross-cluster link state characteristics are first obtained, including the bandwidth margin between clusters, link delay, reputation value, and the global congestion index provided by the satellite. A multi-dimensional feature vector and a set of alternative next-hop links are constructed, and the alternative links are dynamically evaluated through a deep Q-network (DQN). The input of the network is the current state feature vector, and the output is the Q-value score of each link. The link with the highest score is selected as the optimal next-hop. The reward function of the DQN integrates the end-to-end delay, link stability, and energy efficiency, and the strategy is optimized through offline training and online fine-tuning. In addition, the satellite synchronizes the global topology change information based on the inter-satellite link, and periodically updates the cross-cluster link feature library to ensure that the routing decision adapts to the dynamic network environment. In the intra-cluster routing decision, the cluster head UAV first constructs a local topology map based on the real-time states (position, speed, remaining energy) of the intra-cluster nodes, performs node feature aggregation and path weight calculation, and selects the intra-cluster transmission path with the minimum comprehensive cost from the pre-computed candidate paths through the multi-agent proximal policy optimization (MAPPO) algorithm combined with the weight distribution. The cost function covers communication delay, energy consumption balance, and link lifetime. The cluster member nodes execute data forwarding according to the routing table issued by the cluster head, and at the same time periodically feedback the local link quality to the cluster head to trigger the update of the dynamic routing table. In response to a sudden link interruption, the cluster head activates the local reinforcement learning module to quickly re-route based on the immediate environmental feedback to ensure the robustness of intra-cluster communication.

[0075] Based on the method provided by the present invention, the present invention also provides an air-space network hierarchical cooperation routing optimization system, including:

[0076] Network modeling and initialization module: The satellite constructs a global topology snapshot of the UAV swarm through on-board sensors and inter-satellite links, and provides initial weight configuration suggestions for the clustering algorithm;

[0077] Feature information acquisition module, used to acquire the feature information of each UAV node in the UAV swarm; the feature information includes the geographical position coordinates and speed of the UAV node.

[0078] State monitoring and acquisition module, which monitors the cross-cluster link state (bandwidth, delay, load) and intra-cluster link state (SNR, energy, queue) in real time through distributed proxy nodes, and synchronizes them to the global state library;

[0079] The dynamic reputation value management module collects the current data traffic on the link in real time through the SDN controller, records the time difference between packet sending and receiving through timestamps, calculates the end-to-end delay, counts the number of transmission failures of the link in the recent time window, stores global parameters such as historical reputation values, node load ratios, and topological structures using the global status library, dynamically quantifies the link performance, preferentially selects high-reputation paths, and adapts to network topology changes to ensure communication efficiency and stability;

[0080] The hierarchical feature extraction module, the global feature extraction sub-module uses a deep neural network to extract features such as the reputation distribution and load balancing degree of the satellite link, and the local feature extraction sub-module aggregates local features such as the energy of nodes within the cluster and the link quality through GNN message passing;

[0081] The hierarchical routing decision module, the upper-layer decision sub-module (satellite) runs the DDQN algorithm, inputs the global feature vector, and outputs the action of selecting the optimal cross-cluster path, and the lower-layer decision sub-module (drone cluster) runs the MAPPO algorithm, inputs the local feature vector, and outputs the in-cluster routing action (multi-hop or request satellite forwarding);

[0082] The routing policy deployment module converts the decision path into flow table instructions, the satellite issues the inter-satellite link priority and bandwidth reservation policy, and the drone cluster head adjusts the forwarding rules of member nodes;

[0083] The network status monitoring and replanning module continuously detects topological changes such as node movement and link interruption, triggers the recalculation of reputation values and the update of routing policies, and starts dynamic replanning when the threshold is exceeded;

[0084] The resource collaboration feedback module, the satellite broadcasts the cross-cluster link reputation value to the cluster head, and the cluster head feedbacks the in-cluster load and energy status to achieve global-local resource collaborative optimization;

[0085] The drone cluster communication module is used for each drone node in the drone cluster to communicate according to the cluster structure.

[0086] Furthermore, the present invention also provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the computer program in the memory, and can perform real-time analysis on a large amount of data from multiple drone clusters and other satellites. The processor can call the computer program in the memory to execute management tasks for the satellite and drone hybrid network, such as cross-cluster routing decision-making, global resource allocation, etc. For example, when receiving a cross-cluster communication request from a certain drone cluster, the processor selects the optimal cross-cluster routing path for it according to the information such as the reputation value and load condition of each link stored in the memory by using the corresponding algorithm.

[0087] In addition, when the computer program in the above-mentioned memory is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a non-transitory computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs. These storage media can conveniently apply the technical solution of the present invention to different devices and systems, promoting the popularization and application of satellite and unmanned aerial vehicle hybrid network technologies. The unmanned aerial vehicle memory stores the flight parameters, energy state, and intra-cluster link information of the unmanned aerial vehicle itself, and the SDN controller is responsible for flexibly controlling and managing the network within the unmanned aerial vehicle cluster and adjusting the intra-cluster routing strategy according to the real-time network state.

[0088] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0089] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that unless otherwise defined, the technical terms or scientific terms used in the present invention should be the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. Each step and algorithm described in the present invention can be appropriately adjusted according to the actual application scenario. For example, different interpolation algorithms, weight calculation formulas, or database storage schemes can be adopted. The streaming computing framework and cross-platform interface specifications can also be customized according to the specific system architecture and user requirements. This embodiment does not limit the protection scope of the present invention, and its technical idea is applicable to a variety of different implementation manners.

[0090] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An optimization method and system for hierarchical collaborative routing in an air-space network, characterized in that, It includes the following steps: Step 1: Establish a collaborative network model including low-earth orbit satellite nodes and clustered unmanned aerial vehicle (UAV) swarms. Among them, the satellite nodes form a cross-cluster communication backbone network, and the UAV swarms are dynamically clustered according to mission requirements. Each cluster contains a cluster head node and several member nodes. Intra-cluster communication is through an ad-hoc network, and inter-cluster data is forwarded through satellite nodes; Step 2: Monitor the real-time states of satellites and UAV clusters through distributed agent nodes, including: cross-cluster link states, intra-cluster link states, and dynamic reputation values; Step 3: Define a dynamic reputation value function for the cross-cluster link between the satellite and the cluster head. The reputation value dynamically reflects the historical performance and reliability of the link, and is used to guide cross-cluster routing selection; Step 4: Perform hierarchical feature extraction and state modeling. Input the cross-cluster link state information into a deep neural network to extract global feature vectors, including reputation value distribution, load balance degree, and path feasibility score; input the intra-cluster link state information into a graph neural network (GNN), and aggregate local topology features through message passing to generate feature vectors for intra-cluster routing decisions; Step 5: Implement hierarchical multi-agent reinforcement learning routing decisions. The upper-layer agent (low-earth orbit satellite) is based on the Deep Q-Network (DQN) algorithm. Input the global state information and output cross-cluster path selection actions, and optimize the Q-value function by minimizing the temporal difference error; Step 6: The lower-layer agent (UAV swarm) adopts the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. Input the local state information and output intra-cluster routing actions, and update the Actor-Critic network parameters through policy gradients; Step 7: Deploy the trained hierarchical routing policy to the actual network. The satellite modifies the inter-satellite link routing table entries through control commands to specify high-reputation value paths, and the UAV cluster heads dynamically adjust the forwarding rules of member nodes to respond to local state changes; Step 8: Continuously collect network topology changes (such as node movement, link interruption), trigger recalculation of reputation values. If it is detected that the link state change exceeds the threshold, repeat Steps 2 to 7, update the routing policy and reissue the flow table; Step 9: Judge whether the task is completed or the network enters a stable state. If continuous optimization is required, return to Steps 5 and 6, otherwise terminate the routing adjustment process.

2. The satellite and UAV cluster cooperative communication routing optimization method based on the reputation value mechanism according to claim 1, wherein, The bandwidth utilization rate, transmission delay, node load rate, and link stability constitute the cross-cluster link state, and the signal-to-noise ratio (SNR), remaining energy of member nodes, and data queue length constitute the intra-cluster link state.

3. The satellite and UAV cluster collaborative communication routing optimization method based on the reputation value mechanism according to claim 1, characterized in that The described dynamic reputation value function comprehensively considers multiple state factors, specifically including: Consider the bandwidth BW occupied by this transmission used , the normalized transmission delay delay, and the number of consecutive link failures fail_count, to form a reputation value function: Among them, the exponential term Ensure that delay-sensitive services prefer low-latency links first, and the coefficients meet the constraint condition α + β + γ = 1; The reputation value update function is: C i,j (t + 1) = ρ·C i,j (t) + ΔC i,j (t) represents the reputation value of link i→j at time t + 1, where ρ is the attenuation coefficient, and its value range is 0 to 1, which is used to prevent historical values from becoming obsolete.

4. The satellite and UAV cluster cooperative communication routing optimization method based on the reputation value mechanism according to claim 1, characterized in that The described hierarchical state modeling includes: The state space of the upper-layer agent (low-earth orbit satellite) can be expressed as: s u = [C i,j (t), BW avail , delay avg , load sat ​ Among which C i,j (t) is the reputation value of the inter-cluster link, BW avail is the remaining available bandwidth, delay avg is the average delay, load sat is the satellite load rate; The state space of the lower-layer agent (UAV) can be expressed as: s L = [link quality , energy residual , neighbor states , C set_link ​ Among them, link quality represents the signal-to-noise ratio of the intra-cluster link, enerhy residual represents the remaining energy of the node, neighbor states represents the status of the neighbor nodes, C set_link represents the cross-cluster link reputation value provided by the satellite.

5. The method for optimizing the collaborative communication routing of satellites and UAV clusters based on the reputation value mechanism according to claim 1, wherein In the described hierarchical routing decision, the interaction between the upper-layer agent (satellite) and the lower-layer agent (UAV cluster) is realized through state sharing, action triggering, reputation value feedback, and dynamic path adjustment.

6. The hierarchical routing decision according to claim 5, wherein The upper-layer routing decision process is that the satellite selects the cross-cluster path based on the Deep Q-Network (DQN), inputs the global state (cross-cluster link reputation value C i,j (t), bandwidth, delay, load), selects the next-hop node (satellite or cluster head) by maximizing the Q value, and the reward function is: r U = ω1e -delay + ω2C selected + ω3(1 - sat_load) Encourage low-latency, high-reputation, and low-load paths. Satellites regularly broadcast the cross-cluster link reputation values to all cluster heads, and the cluster heads incorporate the reputation values into the lower-layer state space to evaluate whether to preferentially select a certain cross-cluster link.

7. The hierarchical routing decision according to claim 5, wherein The lower-layer routing decision process is that the UAV cluster adopts multi-agent proximal policy optimization (MAPPO), inputs the local state (SNR, remaining energy, queue, satellite reputation feedback C set_link ), selects the next hop within the cluster, and the reward function is: Optimize bandwidth, energy consumption, and cross-link reputation. The policy gradient is updated through the advantage function A(s L ,a L ) = Q(s L ,a L ) - V(s L ). The cluster head feeds back the in-cluster link status to the satellite for the upper layer to update the global load balancing parameters.

Citation Information

Cited By

  • Layered deep reinforcement learning routing protocol method based on multi-link ad hoc network

    CN121037284A

  • Hierarchical deep reinforcement learning routing protocol method based on multi-link ad hoc network

    CN121037284B

  • Wireless network adaptive routing method and system based on cluster division and multilayer hypergraph

    CN121218297A

  • Offshore area communication system

    CN121486412A

  • Method and system for dynamically optimizing HPLC (High Performance Liquid Chromatography) route based on link quality

    CN122247908A