Air-ground integrated network spectrum power joint allocation method combining hypergraph and meta reinforcement learning

By combining the hypergraph and meta-reinforcement learning methods, a hypergraph structure is constructed and a graph attention network is used to extract high-order interference features. Combined with the Meta-DQN strategy network, spectrum power is jointly allocated. This solves the problems of dynamic topology changes and link interference in UAV communications, and achieves efficient allocation of spectrum resources and improved communication reliability.

CN120768427APending Publication Date: 2025-10-10NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511032039.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Traditional cellular networks lack the ability to adapt to the dynamic topology changes of airspace nodes and link interference characteristics in drone communications, resulting in interference problems caused by spectrum reuse when air-to-air and air-to-ground links coexist, affecting communication reliability and system throughput performance. In addition, existing deep learning and graph neural network methods have weak generalization capabilities in complex dynamic environments and are unable to meet the scalability and real-time requirements in large-scale node scenarios.

Method used

By combining hypergraph with meta-reinforcement learning, a hypergraph structure and graph attention network are constructed to extract high-order interference relationships, and the Meta-DQN strategy network is combined to perform spectrum power joint allocation, thus achieving efficient allocation of spectrum resources and improving communication reliability in UAV networks.

Benefits of technology

It significantly improves spectrum utilization efficiency and communication reliability, realizes efficient joint allocation of spectrum and power in multi-UAV communication systems, adapts to dynamic topology and task heterogeneity, and provides intelligent and distributed technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768427A_ABST
    Figure CN120768427A_ABST
Patent Text Reader

Abstract

The invention discloses an air-ground integrated network spectrum power joint distribution method combining hypergraph and meta reinforcement learning. The method comprises the following steps: step 1, establishing a multi-unmanned aerial vehicle communication network model; 2, constructing an undirected graph and hypergraph structure; 3, introducing a graph attention network to realize bidirectional aggregation of node features; 4, the aggregated node features are input into a Meta-DQN strategy network, so that the strategy has rapid adaptive capacity when facing environment changes and new tasks; according to the air-ground integrated network spectrum power joint allocation method combining the hypergraph and meta reinforcement learning, node feature representation with global perception capability can be acquired by each unmanned aerial vehicle agent on the premise of only depending on local observation, and rapid adaptation is realized in the face of dynamic environment change, so that efficient migration of a resource allocation strategy is realized, and the resource allocation efficiency is improved. The spectrum utilization rate and the communication success rate of the system are improved, and stable operation of the air-ground converged communication system in a complex scene is effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicles (UAVs), and in particular to a method for jointly allocating spectrum power in an air-ground integrated network by combining hypergraph and meta-reinforcement learning. Background Art

[0002] Amidst rapid technological advancements and profound economic restructuring, the low-altitude economy is rapidly evolving from a conceptual concept to a real industry, becoming a new frontier of international industrial competition. As the digital foundation for its development, the Low-Altitude Intelligent Network (LIN) is a key enabler for the opening up of airspace and the scale-up of its applications. The LIN is increasingly responsible for ensuring communications for drone swarm missions in high-density, complex environments. To meet the data exchange and control requirements of multiple drones performing coordinated missions, achieving efficient and reliable spectrum and power allocation within resource-constrained environments has become a key technological challenge.

[0003] Currently, cellular networks are widely used in drone communications. 5G communications, in particular, are considered a key enabler for building low-altitude intelligent networks due to their high bandwidth and low latency. However, traditional cellular networks are primarily designed for ground users and lack the ability to adapt to dynamic topological changes in airspace nodes and link interference. Especially in the context of scarce uplink resources, interference caused by spectrum reuse when air-to-air and air-to-ground links coexist is a particularly prominent issue, impacting both communication reliability and overall system throughput.

[0004] To address these issues, the academic community has explored the use of deep learning (DL) and deep reinforcement learning (DRL) methods to model and optimize spectrum power resource allocation. While DRL technology possesses the ability to autonomously learn strategies, it still faces challenges such as low training sample efficiency and weak generalization capabilities in complex and dynamic environments (such as high-speed drone movement and frequent link changes). Furthermore, most existing methods utilize centralized learning architectures. While these approaches offer advantages in processing global information, they struggle to meet the scalability and real-time requirements of large-scale node scenarios.

[0005] Graph Neural Networks (GNNs), an effective tool for modeling structured data, have been introduced in recent years to UAV networks to extract high-level features from interference relationships between links. However, traditional GNN architectures primarily target static graphs and are limited in their effectiveness when dealing with the highly dynamic connectivity structures found in UAV networks. Furthermore, conventional graph structures struggle to accurately capture the complex many-to-many interference relationships between U2U links, limiting the model's expressive power.

[0006] In order to improve the representation ability and adaptability, the hypergraph structure is proposed as a higher-order graph modeling tool. It can connect multiple nodes through a hyperedge and effectively characterize the joint interference relationship between multiple links. On this basis, the introduction of the graph attention network (GAT) to perform feature aggregation on the hypergraph can further enhance the model's selective attention to key communication links. Combined with the Meta Reinforcement Learning framework, drones can be equipped with the ability to quickly adapt to new environments and achieve small sample migration in changing scenarios. In view of this, this patent designs a joint spectrum power allocation method for air-ground integrated networks that combines hypergraphs and meta reinforcement learning to overcome the limitations of low-altitude complex environments and limited spectrum resources on data rates, communication reliability, etc. Summary of the Invention

[0007] Purpose of the invention: To provide a spectrum power joint allocation method for air-ground integrated networks that combines hypergraph and meta-reinforcement learning, aiming to overcome the limitations of low-altitude complex environments and limited spectrum resources on data rate, communication reliability, etc.

[0008] Technical solution:

[0009] A method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning, comprising the following steps:

[0010] Step 1: Based on the urban multi-UAV inspection scenario, a multi-UAV communication network scenario model is established, the network channel model and performance indicators are determined, and the communication relationship between adjacent UAVs is obtained;

[0011] Step 2: Construct the drone's communication network into an undirected graph based on the communication relationship. Determine the neighbors of the nodes in the undirected graph and the features contained in each node. Use the maximum clique search algorithm of the undirected graph to obtain all the maximum cliques in the undirected graph. Construct each calculated maximum clique into a hyperedge to generate a hypergraph structure.

[0012] Step 3: Based on the constructed hypergraph structure, a graph attention network is used to process the hypergraph structure so that each node obtains feature embedding covering global information;

[0013] Step 4: The aggregated node features are embedded and input into the Meta-DQN policy network to obtain the optimal channel and power control scheme to maximize the transmission rate of the U2B link from the drone to the base station and the communication success rate of the U2U link from the drone to the drone.

[0014] In a further embodiment, the multi-UAV communication network scenario model established in step 1 specifically includes a scenario area, a base station location, and UAV information, wherein the base station location is located at the center of the scenario area, and the UAV information includes quantity, location, and speed;

[0015] Deploy a group of drones S = {1, 2, 3, ..., s} in the scene area D, where each drone flies in a 3-D space and is added to the scene area in a random distribution. The drones are made to travel at a constant speed at a preset initial velocity and direction at different preset heights.

[0016] The communication between the drones is achieved through U2U links, and the communication between the drones and the base station is achieved through U2B links.

[0017] In a further embodiment, the spectrum resources of the U2B link are directly allocated by the base station, and the spectrum resources of the U2U link are allocated while taking into account the pre-allocation of the spectrum resources of the U2B link. The spectrum resource selection of the UAV is divided into two parts: sub-channel selection and transmit power selection. Each U2U link can only select one sub-channel and one transmit power level;

[0018] Assuming that the number of subchannels is the same as the number of U2B links, denoted as m, and that the U2U link has n transmission power levels, each U2U link has a total of m*n resource selection methods. Therefore, the establishment of the channel model specifically includes the following:

[0019] Assume that all drones are flying at different preset altitudes. At time t, the position of the i-th drone is represented by p i (t) = [x i (t),y i (t),h i ] T ∈R 3 , where t∈[0,T], h i Represents the flight altitude of the i-th UAV;

[0020] UAV i is moving at a constant speed v at time t. i Toward direction d i (t) moves, so the position of drone i at the next time t+1 is p i (t+1)=p i (t)+v i ·d i (t), the distance between UAV i and UAV j is d i,j (t)=||p i (t)-p j (t)||;

[0021] The height of the base station is H bs , so the distance between drone i and the base station is recorded as d i,bs (t);

[0022] The air-to-ground channel is mainly dominated by line-of-sight links (LoS) and non-line-of-sight hybrid links (NLoS). Therefore, the path loss between UAV i and the base station is calculated as follows:

[0023]

[0024] The probability calculation formula for the existence of a line-of-sight link between drone i and the ground base station is as follows:

[0025]

[0026] Where d0 = max(294.05log 10 h i -432.94,18), p1=233.98log 10 h i -0.95, obviously the non-line-of-sight probability is P i NLoS (t)=1-P i LoS (t);

[0027] Therefore, the average path loss between drone i and base station is calculated as follows:

[0028]

[0029] Taking into account small-scale fading, the channel power gain between UAV i and the base station at time t is expressed as follows:

[0030]

[0031] Among them, H i,B (t) represents the fading coefficient between UAV i and the base station;

[0032] The UAV communication channel is mainly dominated by the line-of-sight link (LoS), and the channel state information is determined by the position of the UAV. Therefore, the channel power gain of the U2U link is characterized by the commonly used free-space path loss model as follows:

[0033] g i =β0·(d i,j (t)) -α (5)

[0034] Where β0 represents the channel power gain at the reference distance, and parameter α represents the path loss exponent.

[0035] In a further embodiment, it is assumed that S drones are set up in the scene area D. Each drone has communication needs with the three nearest drones. Then, the drones communicate directly with each other through device-to-device communication. Therefore, the number of U2U links is 3S. It is assumed that there are K pairs of U2U users, expressed as K = {1, 2, 3, ..., k}. Since the U2U link and the U2B link share orthogonally allocated uplink spectrum resources, the signal-to-interference-and-noise ratio (SINR) of the nth U2B link can be expressed as:

[0036]

[0037] in, represents the transmit power of the nth U2B link, h n,B [n] represents the instantaneous channel power gain of the nth U2B link on the nth subchannel; σ 2 is the noise power; ρ i [n] means that if the i-th U2U link uses the spectrum resources of the n-th U2B link, then ρ i [n] is assigned a value of 1, otherwise it is assigned a value of 0;

[0038] It represents the transmit power of the i-th U2U link; It represents the power gain of the interference of the i-th U2U link to the n-th U2B link;

[0039] According to Shannon's formula, the channel capacity of the nth U2B link is expressed as:

[0040]

[0041] Where B represents bandwidth;

[0042] Similarly, the SINR of the i-th U2U link is expressed as:

[0043]

[0044] in, represents the transmit power of the i-th U2U link on the n-th subchannel, g i,i [n] represents the power gain of the i-th U2U link in the n-th subchannel;

[0045] The interference power of the U2B link sharing the same resource block RB to the i-th U2U link is expressed as follows:

[0046]

[0047] in, represents the interference power gain of the nth U2B link to the i-th U2U link;

[0048] The total interference power of all other U2U links sharing the same resource block RB to the i-th U2U link is expressed as follows:

[0049]

[0050] in, represents the interference power gain of the j-th U2U link to the i-th U2U link;

[0051] The channel capacity of the i-th U2U link is:

[0052]

[0053] In a further embodiment, the performance indicator settings specifically include the total transmission rate of the U2B link and the successful transmission of the U2U link;

[0054] Maximize the total transmission rate of the U2B link A successful transmission for each U2U link is defined as:

[0055]

[0056] Where L is the payload size of each U2U link, T and Δ T denote the maximum tolerated delay and the duration of each time slot, t is the time slot index, and t i is the time slot when the i-th U2U link starts transmitting the payload;

[0057] If the duration of transmitting the payload L does not exceed the maximum tolerable delay T, the information transmission of the U2U link is successful, otherwise it fails. In a given transmission period, the successful transmission probability of all U2U links is expressed as follows:

[0058]

[0059] Among them, O i is the number of payloads that the i-th U2U link needs to transmit within the transmission period; It is the indicator of whether the i-th U2U link has successfully transmitted its u-th payload, i.e., when the transmission is successful otherwise

[0060] Therefore, the optimization problem is formulated as follows:

[0061]

[0062] Among them, C1 represents the requirement of the U2U link for successful payload transmission, C2 indicates that each U2U link will only be allocated one resource block, and C3 represents the constraint on the transmission power of the U2U link.

[0063] In a further embodiment, step 2 specifically includes the following steps:

[0064] Step 2-1: According to graph theory, the U2U links are regarded as nodes and the interference relations are regarded as simple edges. The UAV communication network is modeled as an undirected graph G = (V, E), and the node set is represented as V = {u1,u2,...,u k}, the set of edges is determined according to the communication relationship between UAVs;

[0065] Step 2-2: Determine the neighbors of the node and the features of each node. For node i, it contains an initial feature and a list N(i) storing the index of connected hyperedges; the initial features of the node are the channel and interference information observed by the drone itself;

[0066] At time slot t, the channel information of the i-th U2U link is expressed as It contains the instantaneous channel power gain of its own link transmitter in the nth subchannel Instantaneous interference channel power gain from the nth U2B link transmitter in the nth subchannel The instantaneous interference channel power gain from the jth (j≠i) U2U link transmitter in the nth subchannel

[0067] At time slot t, the channel information of the nth U2B link is expressed as It contains the instantaneous channel power gain of its own link transmitter on the nth subchannel Instantaneous interference channel gain from U2U link transmitter to base station The interference signal strength of the i-th U2U link receiver in the n-th subchannel in the previous time slot is

[0068] Therefore, node u i The characteristics are as follows:

[0069]

[0070] Among them, || represents the concatenation of vectors;

[0071] Step 2-3: According to the Bron-Kerbosch algorithm, find all the maximal clusters in the undirected graph and construct each calculated maximal cluster into a hyperedge to obtain the hypergraph structure G = (V, H E );

[0072] Where V represents the vertex set, H E represents a hyperedge set.

[0073] In a further embodiment, based on the constructed hypergraph structure, in the hypergraph structure, the aggregation process of GAT includes a first stage and a second stage;

[0074] The first stage is the information transfer from U2U nodes to hyperedges. Each hyperedge first receives information from the node it connects to. By calculating the attention coefficient between each node and the hyperedge, the features of all nodes are aggregated according to the weights to generate the representation of the hyperedge.

[0075] Calculate the attention coefficient of the aggregation from node to hyperedge, for hyperedge e j ∈H E , calculate each node u it connects one by one i ∈N(e j ) and itself:

[0076] ξ ij =α([Wh i ||Wh j ]) (16)

[0077] Among them, W is a shared parameter obtained through learning, which linearly transforms the node or hyperedge features and projects them into high dimensions; [·||·] is the i , super edge e j The transformed features are spliced; α(·) is a function that calculates the correlation between nodes and hyperedges, and the spliced ​​high-dimensional features are mapped to a real number through a single-layer feedforward neural network; the vertex features contained in the hyperedge are averaged or weighted summed as the initial features of the hyperedge

[0078] The attention coefficient obtained by LeakyReLu activation function and softmax normalization is expressed as follows:

[0079]

[0080] According to the obtained attention coefficient, the hyperedge e j Aggregate the node features contained in the node and calculate its new feature h' j , the specific formula is as follows:

[0081]

[0082] Where σ represents the activation function;

[0083] The second stage is the information transmission from the hyperedge to the U2U node. The node then receives information from the hyperedge to which it belongs, calculates the node's attention weight to its associated hyperedge, and sums the weighted hyperedge features to update the node's representation.

[0084] Calculate the attention coefficient of the aggregation from the hyperedge to the node. The workflow is the same as that of the first stage to obtain the attention coefficient β ij , the new node feature h' obtained by aggregation i , the specific formula is as follows:

[0085]

[0086] Therefore, the first stage operation is performed on each hyperedge in the hypergraph structure, and the second stage operation is performed on each node to obtain the final embedding of each node in the UAV network table view.

[0087] In a further embodiment, the optimization problem of formula (14) is an NP-hard problem, and is therefore modeled as a Markov decision process to define the state space, action space, and reward function of the agent in reinforcement learning;

[0088] State space: Denote the state space as S i , which contains the states of all agents in each time slot; for each time slot t, agent u i Status S i (t) consists of two parts. The first part is the aggregated features extracted by HGNN, and the second part is the local observation of the intelligent agent. Specifically, the local observation of the UAV is also divided into two parts. The first part is the channel and interference information, including the channel information of the U2U link. U2B link channel information Interference signal strength of the previous time slot of the U2U link The second part includes the number of times the neighboring drone selected the nth subchannel in the previous time slot The remaining number of unsent bits and the remaining sending time under the delay constraint and

[0089] Therefore, at time slot t, the environment state of the i-th U2U agent is expressed as follows:

[0090]

[0091] Action space: Denote the action space as A i , based on the locally observed state and strategy, each U2U agent will decide the sub-channel selection one by one and transmit power selection Since the agent needs to select these two actions at the same time, these two actions are combined into a composite action A. i (t), considering three power levels and m number of sub-channels, there are 3*m actions in total, which can be decomposed as follows:

[0092]

[0093] Among them, % represents the modulo operation, / represents division and rounding down;

[0094] Reward function: The immediate reward is formulated by considering three parts: the total rate of the U2B link, the total rate of the U2U link, and the transmission time. Therefore, the immediate reward for time slot t is expressed as follows:

[0095]

[0096] Among them, λ c ,λ p Represents the positive weight of each part, T is the maximum tolerable delay, the expression Indicates the time taken for the transmission.

[0097] In a further embodiment, in step 4, the DQN algorithm includes a Q network and a target Q network, and the two networks have the same DNN structure;

[0098] At the beginning of time slot t, the i-th U2U agent follows the ε-greedy strategy and state S i (t) Select an action A from the action space i (t), which means that the agent randomly selects A with probability ε∈(0,1) i (t), or select an action with probability 1-ε according to the following formula:

[0099]

[0100] Among them, Q(S i (t),A i (t),θ) is a given observation state S i (t) and Action A i The output Q value of the Q network at time (t), θ represents the weight of the Q network;

[0101] The Q value Q* corresponding to the optimal strategy is obtained according to the following update equation:

[0102]

[0103] Use the gradient descent method to minimize the loss function of the i-th U2U agent to update the weight θ of the Q network as follows:

[0104]

[0105] Among them, γ is the discount factor, α is the learning rate, and the weight variable of the target network is Updated periodically with θ.

[0106] In a further embodiment, a meta-learning-based DRL algorithm, namely the HGNN-Meta-DQN algorithm, is used to define a meta-task set T, which includes N T tasks, each task T j (j=1,...,N T ) is considered as a Markov decision process, for each task T j Define a replay buffer for storing experience The support set and query set of each task are defined as and

[0107] Each task sequentially undergoes a meta-training phase, a meta-adaptation phase, and a deployment phase;

[0108] In the meta-training phase, the support set and query set are used to update the weights of a single task network and the global network, respectively. The meta-training phase adopts a two-level training mechanism. The first level training mechanism is individual-level update, and the other level training mechanism is called global-level update. Each task obtains its own network weight by obtaining different experiences from the replay buffer and using the same optimization problem. Therefore, the network weight of each task can be updated by the following optimization problem for task j:

[0109]

[0110] in, represents the weight of the Q network for task j, is the loss function of the Q network of agent i defined in formula (19), is the replay buffer from task j The experience set sampled from the

[0111] The gradient descent method is used to update the network parameters, which is expressed as follows:

[0112]

[0113] in, represents the individual-level update learning rate of the Q network, and the superscript n represents the number of iterations. The first iteration (n=1) of the network parameters of each task has the corresponding global network parameters updated, that is, Then the parameters obtained from the previous iteration are updated;

[0114] After all tasks in the batch complete their respective network parameter updates, a global update is performed. This update process is achieved by aggregating the adaptability of each task's training strategy to the newly sampled experience. These loss functions are added together to form a loss function for optimizing the global network parameters, which is expressed as follows:

[0115]

[0116] Based on the gradient descent method, the parameters of formula (23) are updated by the following formula:

[0117]

[0118] Where α is the learning rate for global level updates;

[0119] When the individual-level update and the global-level update are completed, the algorithm enters the next batch and continues to update the global network parameters;

[0120] In the meta-adaptation phase, each U2U agent is able to quickly adapt to new tasks using small sample data based on well-trained network parameters θ. Similar to the parameter update process in individual-level updates, the network parameters for the new task are updated in the following way:

[0121]

[0122] in, At the beginning of the time step, it is initialized to the global parameter θ obtained by training; the experience playback area of ​​the task is D s ;

[0123] Among them, the HGNN network and the meta-DQN network are updated separately. The HGNN network is used to store the reward information obtained by the U2U agent in each subchannel during the global update phase in a matrix corresponding to the number of subchannels, which is recorded as This matrix is ​​used as the label of the corresponding node, and the label is softened in the following way:

[0124]

[0125] Among them, ω represents the weight value of the HGNN network aggregation result, represents the aggregation result of the lagged network, represents the label of the i-th U2U agent;

[0126] The HGNN network is updated using the mean square error function and expressed as follows:

[0127]

[0128] in, Represents the weight parameter of the HGNN network, y i represents the smoothed label used for network update;

[0129] Until both networks converge;

[0130] During the deployment phase, the HGNN parameters and Meta-DQN policy network parameters obtained during the training phase are solidified and deployed to each drone node. After each deployment, each agent periodically collects local channel state information, interference information, and neighbor node behavior status, and inputs them into the HGNN network to complete the update of node features. The Meta-DQN policy network completes the joint selection of sub-channels and transmission power, and outputs the optimal resource allocation strategy.

[0131] Beneficial effects of the present invention: The method provided by the present invention is a distributed solution. First, a hypergraph modeling is performed on the network's nodes, links and other information. Then, a graph neural network is constructed to perform message clustering and aggregation iteration based on hyperedges, so that a single drone intelligent agent can capture node features containing global information only through local environment observations. After that, a decision is made through the trained Meta-DQN model to achieve a better resource allocation strategy. By constructing a hypergraph structure of the drone communication network, a graph attention mechanism is introduced to extract high-order interference features, and combined with a meta-reinforcement learning framework, rapid generalization and adaptive optimization of resource allocation strategies are achieved, thereby improving spectrum utilization efficiency and communication reliability. This method significantly improves the global perception ability and dynamic environment adaptability of resource allocation, realizes the efficient joint allocation of spectrum and power in multi-drone communication systems, effectively responds to the challenges brought by dynamic topology and task heterogeneity, and provides intelligent and distributed technical support for the integrated development of "artificial intelligence + low-altitude economy". BRIEF DESCRIPTION OF THE DRAWINGS

[0132] Figure 1 This is a schematic diagram of the communication spectrum resource allocation scenario for low-altitude collaborative mission execution based on low-altitude intelligent networked drones of the present invention.

[0133] Figure 2 This is a flow chart of a method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to the present invention.

[0134] Figure 3 Flowchart of the HGNN model of the present invention.

[0135] Figure 4 This is a flowchart of the meta-learning of the present invention.

[0136] Figure 5 This is a schematic diagram of constructing an undirected graph based on the communication relationship of the U2U link of the present invention.

[0137] Figure 6 This is a schematic diagram of the maximal clique hypergraph structure obtained based on an example of the present invention. DETAILED DESCRIPTION

[0138] In the following description, numerous specific details are provided to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, certain technical features well known in the art are not described to avoid confusion with the present invention.

[0139] The present invention will be further described in detail below with reference to the accompanying drawings.

[0140] Reference Figure 1-6 , a method for joint spectrum power allocation of an air-ground integrated network combining hypergraph and meta-reinforcement learning disclosed in the present invention, comprising the following steps:

[0141] Step 1: Based on the urban multi-UAV inspection scenario, a multi-UAV communication network scenario model is established, the network channel model and performance indicators are determined, and the communication relationship between adjacent UAVs is obtained. A UAV has fixed communication requirements with its three nearest neighboring UAVs.

[0142] The multi-UAV communication network scenario model established in step 1 specifically includes the scenario area, base station location, and UAV information, wherein the base station location is located at the center of the scenario area, and the UAV information includes the number, location, and speed;

[0143] A group of drones S = {1, 2, 3, ..., s} is deployed in the scene area D. Each drone flies in 3-D space and is added to the scene area in a random distribution. The drones are made to travel at preset different altitudes with a preset initial velocity and direction. To avoid collisions, each drone is required to travel at a different altitude.

[0144] Communication between drones is achieved via U2U links, and communication between drones and base stations is achieved via U2B links. Assuming each drone is equipped with an antenna, inter-drone communication supports U2U links, transmitting necessary safety information during flight. Furthermore, a group of cellular users M = {1, 2, 3, …, m} exists within the scenario area, requiring high-speed data transmission such as video images. These users communicate with the base station using U2B links.

[0145] Considering the limited spectrum resources and the large demand for uplink in UAV communication networks, in addition, the utilization of uplink is very sparse, so the U2U link and the U2B link share the orthogonally allocated uplink spectrum to improve the utilization of the spectrum. The spectrum resources of the U2B link are directly allocated by the base station, and the spectrum resources of the U2U link are allocated in consideration of the pre-allocation of the spectrum resources of the U2B link. The spectrum resource selection of the UAV is divided into two parts: subchannel selection and transmission power selection, and each U2U link can only select one subchannel and one transmission power level;

[0146] Assuming that the number of subchannels is the same as the number of U2B links, denoted as m, and the U2U link has n transmission power levels, so each U2U link has m*n resource selection methods, therefore, the establishment of the channel model specifically includes the following:

[0147] Assuming that all UAVs fly at a preset specific different height, at time t, the position of the i-th UAV is represented as p i (t)=[x i (t),y i (t),h i ] T ∈R 3 , where t∈[0,T], h i represents the flight height of the i-th UAV;

[0148] The UAV i moves at a constant speed v i in the direction d i (t) at time t, so the position of the UAV i at the next time t+1 is p i (t+1)=p i (t)+v i ·d i (t), the distance between the UAV i and the UAV j is d i,j (t)=||p i (t)-p j (t)||;

[0149] The height of the base station is H bs , so the distance between the UAV i and the base station is denoted as d i,bs (t);

[0150] In this system, under the condition of reasonable flight height and size, the air-to-ground channel is mainly dominated by the line-of-sight link LoS and the non-line-of-sight mixed link NLoS, and the channel model is provided by the 3GPP Release 15 specification, therefore, the calculation formula of the path loss between the UAV i and the base station is as follows:

[0151]

[0152] The probability calculation formula for the existence of a line-of-sight link between drone i and the ground base station is as follows:

[0153]

[0154] Where d0 = max(294.05log 10 h i -432.94,18), p1=233.98log 10 h i -0.95, obviously the non-line-of-sight probability is P i NLoS (t)=1-P i LoS (t);

[0155] Therefore, the average path loss between drone i and base station is calculated as follows:

[0156]

[0157] Taking into account small-scale fading, the channel power gain between UAV i and the base station at time t is expressed as follows:

[0158]

[0159] Among them, H i,B (t) represents the fading coefficient between UAV i and the base station;

[0160] The UAV communication channel is primarily dominated by the line-of-sight (LoS) link, and the channel state information (CSI) is determined by the UAV's position. Therefore, the channel power gain of the U2U link is characterized using the commonly used free-space path loss model as follows:

[0161] g i =β0·(d i,j (t)) -α (5)

[0162] Where β0 represents the channel power gain at the reference distance, and parameter α represents the path loss exponent.

[0163] Assume that there are S drones in scenario area D. Each drone needs to communicate with the three nearest drones. The drones then communicate directly with each other using device-to-device (D2D). Therefore, the number of U2U links is 3S. For convenience, assume that there are K pairs of U2U users, represented by K = {1, 2, 3, …, k}. Since the U2U link and the U2B link share orthogonally allocated uplink spectrum resources, the signal-to-interference-and-noise ratio (SINR) of the nth U2B link can be expressed as:

[0164]

[0165] in, represents the transmit power of the nth U2B link, h n,B [n] represents the instantaneous channel power gain of the nth U2B link on the nth subchannel; σ 2 is the noise power; ρ i [n] means that if the i-th U2U link uses the spectrum resources of the n-th U2B link, then ρ i [n] is assigned a value of 1, otherwise it is assigned a value of 0;

[0166] It represents the transmit power of the i-th U2U link; It represents the power gain of the interference of the i-th U2U link to the n-th U2B link;

[0167] According to Shannon's formula, the channel capacity of the nth U2B link is expressed as:

[0168]

[0169] Where B represents bandwidth;

[0170] Similarly, the SINR of the i-th U2U link is expressed as:

[0171]

[0172] in, represents the transmit power of the i-th U2U link on the n-th subchannel, g i,i [n] represents the power gain of the i-th U2U link in the n-th subchannel;

[0173] The interference power of the U2B link sharing the same resource block (RB) to the i-th U2U link is expressed as follows:

[0174]

[0175] in, represents the interference power gain of the nth U2B link to the i-th U2U link;

[0176] The total interference power of all other U2U links sharing the same resource block RB to the i-th U2U link is expressed as follows:

[0177]

[0178] in, represents the interference power gain of the j-th U2U link to the i-th U2U link;

[0179] The channel capacity of the i-th U2U link is:

[0180]

[0181] The performance indicators include the total transmission rate of the U2B link and the successful transmission of the U2U link.

[0182] Maximize the total transmission rate of the U2B link The U2U link focuses on the reliable transmission of safety-critical information, so the goal is to improve its successful transmission rate. First, the definition of successful transmission of each U2U link is:

[0183]

[0184] Where L is the payload size of each U2U link, T and Δ T denote the maximum tolerated delay and the duration of each time slot, t is the time slot index, and t i is the time slot at which the i-th U2U link starts transmitting the payload. The present invention considers asynchronous communication, and the time at which each U2U link starts transmitting the payload may be different.

[0185] If the duration of transmitting the payload L does not exceed the maximum tolerable delay T, the information transmission of the U2U link is successful, otherwise it fails. In a given transmission period, the successful transmission probability of all U2U links is expressed as follows:

[0186]

[0187] Among them, O i is the number of payloads that the i-th U2U link needs to transmit during the transmission period. Note that the number of payloads that each link needs to transmit during the transmission period may be different. It is the indicator of whether the i-th U2U link has successfully transmitted its u-th payload, i.e., when the transmission is successful otherwise

[0188] Therefore, the optimization problem is formulated as follows:

[0189]

[0190] Here, C1 represents the U2U link's requirement for successful payload transmission, C2 indicates that each U2U link is allocated only one resource block, and C3 represents the U2U link transmit power constraint, preventing excessive transmit power. This resource allocation problem is a multi-objective optimization problem aimed at improving the overall U2B link rate and the U2U link transmission success rate. It is an NP-hard problem.

[0191] Step 2: Construct the drone's communication network into an undirected graph based on the communication relationship. Determine the neighbors of the nodes in the undirected graph and the features contained in each node. Use the maximal clique search algorithm of the undirected graph to obtain all maximal cliques in the undirected graph. Construct each calculated maximal clique into a hyperedge to generate a hypergraph structure to accurately characterize the high-order interference relationship between many-to-many links.

[0192] The step 2 specifically includes the following steps:

[0193] Step 2-1: According to graph theory, the U2U links are regarded as nodes and the interference relations are regarded as simple edges. The UAV communication network is modeled as an undirected graph G = (V, E), and the node set is represented as V = {u1,u2,...,u k}, the set of edges is determined according to the communication relationship between drones, Figure 5 An example of a U2U link is given;

[0194] Step 2-2: Determine the neighbors of the node and the features of each node. For node i, it contains an initial feature and a list N(i) storing the index of connected hyperedges; the initial features of the node are the channel and interference information observed by the drone itself;

[0195] At time slot t, the channel information of the i-th U2U link is expressed as It contains the instantaneous channel power gain of its own link transmitter in the nth subchannel Instantaneous interference channel power gain from the nth U2B link transmitter in the nth subchannel The instantaneous interference channel power gain from the jth (j≠i) U2U link transmitter in the nth subchannel

[0196] At time slot t, the channel information of the nth U2B link is expressed as It contains the instantaneous channel power gain of its own link transmitter on the nth subchannel Instantaneous interference channel gain from U2U link transmitter to base station The interference signal strength of the i-th U2U link receiver in the n-th subchannel in the previous time slot is

[0197] Therefore, node u i The characteristics are as follows:

[0198]

[0199] Among them, || represents the concatenation of vectors;

[0200] The present invention supports distributed deployment, and each drone is loaded with the HGNN model and the Meta-DQN model. Due to the computing resource constraints of drones, if there are a large number of drones in the environment, when all interference relationships are considered, the graph will become extremely complex, and the computational load of aggregating node features will increase dramatically, resulting in extended decision-making time and the introduction of additional delays. To solve this problem, we propose to determine neighbors based on the communication relationship between drones, so that the neighbors are limited to a stable value. Regardless of how the vehicles in the environment change, the node and its neighbors are connected by a simple edge. Specifically, each U2U link is represented as a node in the graph. When the transmitter or receiver of two U2U links is the same device, it is considered that there is an interference relationship between the two, and an undirected edge is established between the corresponding nodes;

[0201] Step 2-3: According to the Bron-Kerbosch algorithm, find all the maximal clusters in the undirected graph and construct each calculated maximal cluster into a hyperedge to obtain the hypergraph structure G = (V, H E );

[0202] Where V represents the vertex set, H E represents a hyperedge set, such as Figure 6 shown.

[0203] Step 3: Based on the constructed hypergraph structure, a graph attention network is used to process the hypergraph structure so that each node obtains a feature embedding that covers global information. The nodes in the undirected graph and the hypergraph are the same, both are U2U links. The graph attention network is used to process the hypergraph model so that each U2U link obtains global information feature embedding that integrates the feature information of neighboring nodes.

[0204] Based on the constructed hypergraph structure, in the hypergraph structure, the aggregation process of GAT includes the first and second stages;

[0205] The first stage is U2U node-to-hyperedge information transfer (Vertex-to-Hyperedge Aggregation): Each hyperedge first receives information from the nodes it connects. Since not all nodes contribute equally to the hyperedge, an attention mechanism is introduced to highlight the nodes that are important to the hyperedge's meaning. The information of these nodes is then aggregated. By calculating the attention coefficient between each node and the hyperedge, the features of all nodes are aggregated according to the weights to generate a representation of the hyperedge.

[0206] Calculate the attention coefficient of the aggregation from node to hyperedge, for hyperedge e j ∈H E , calculate each node u it connects one by one i ∈N(e j) and itself:

[0207] ξ ij =α([Wh i ||Wh j ]) (16)

[0208] Among them, W is a shared parameter obtained through learning, which linearly transforms the node or hyperedge features and projects them into high dimensions; [·||·] is the i , super edge e j The transformed features are concatenated; α(·) is a function that calculates the correlation between nodes and hyperedges. The concatenated high-dimensional features are mapped to a real number and implemented through a single-layer feedforward neural network. To simplify the calculation, we limit the calculation of this correlation to the first-order neighbors, that is, a single-layer attention network. The vertex features contained in the hyperedge are averaged or weighted summed as the initial features of the hyperedge.

[0209] The attention coefficient obtained by LeakyReLu activation function and softmax normalization is expressed as follows:

[0210]

[0211] According to the obtained attention coefficient, the hyperedge e j Aggregate the node features contained in the node and calculate its new feature h' j , the specific formula is as follows:

[0212]

[0213] Among them, σ represents the activation function;

[0214] The second stage is the hyperedge-to-vertex aggregation, where the node receives information from its hyperedges. The reason is the same. The node's attention weight to its associated hyperedge is calculated, and the hyperedge features are weighted and summed to update the node's representation.

[0215] Calculate the attention coefficient of the aggregation from the hyperedge to the node. The workflow is the same as that of the first stage to obtain the attention coefficient β ij , the new node feature h' obtained by aggregation i , the specific formula is as follows:

[0216]

[0217] Therefore, the first stage operation is performed on each hyperedge in the hypergraph structure, and the second stage operation is performed on each node, to obtain the final embedding of each node in the UAV network table view. Through the HyperGAT, not only the high-order relationship between nodes can be modeled, but also the key information of different granularities can be effectively highlighted in the representation learning process, which is very suitable for complex dynamic environments.

[0218] The optimization problem of the formula (14) is an NP difficult problem, so it is modeled as a Markov decision process (MDP), and the state space, action space and reward function of the agent in reinforcement learning are defined;

[0219] The state space is represented as S i , which contains the state of all agents in each time slot; for each time slot t, the state S i of agent u i (t) contains two parts, the first part is the aggregated features extracted by HGNN, and the second part is the local observation of the agent; specifically, the local observation of the UAV is also divided into two parts, the first part is the information of the channel and interference, including the channel information of the U2U link The channel information of the U2B link The interference signal strength of the U2U link in the previous time slot The second part includes the number of times that the neighbor UAV selects the nth subchannel in the previous time slot The number of remaining unsent bits and the remaining transmission time that meet the delay constraint And

[0220] Therefore, at time slot t, the environment state of the i-th U2U agent is represented as follows:

[0221]

[0222] The action space is represented as A i , according to the local observed state and the policy, each U2U agent will decide subchannel selection and transmit power selection Since the agent needs to select both actions at the same time, in order to facilitate calculation, the two actions are combined into a composite action A i (t), the invention considers three power levels, and the number of subchannels is m, so there are 3*m kinds of actions, and the decomposition method is represented as follows:

[0223]

[0224] Where % represents the modulo operation, / represents division and rounding down;

[0225] Reward function: For this invention, good decisions on subchannel allocation and power allocation for each U2U link should maximize the total rate of the U2B link while maximizing the probability of each U2U link successfully transmitting a payload within a specific time. Therefore, this invention considers three components to formulate the immediate reward: the total rate of the U2B link, the total rate of the U2U link, and the transmission time. Therefore, the immediate reward for time slot t is expressed as follows:

[0226]

[0227] Among them, λ c ,λ p Represents the positive weight of each part, T is the maximum tolerable delay, the expression represents the time taken for transmission, which can be considered as a penalty term. As ,increases, the remaining time will decrease, which means that the probability of successfully transmitting the payload within a specific time will decrease. Together with the total rate of the U2U link, this expression reflects the second objective of the optimization problem.

[0228] Step 4: The aggregated node features are embedded and input into the Meta-DQN policy network to make joint decisions on sub-channels and power. During the training phase, a meta-learning mechanism combining individual and global levels is adopted to enable the strategy to quickly adapt to environmental changes and new tasks, obtain the optimal channel and power control scheme, and maximize the transmission rate of the U2B link from drone to base station and the communication success rate of the U2U link from drone to drone, such as Figure 1 When the drone swarm is performing a mission in the air, the U2U links between the drones share the uplink spectrum resources of the U2B link of the drone base station. The HGNN model aggregates neighbor features to obtain feature embedding covering global information. Then, the meta-DQN model is used to find the optimal strategy for the U2U link subchannel and transmission power.

[0229] In step 4, the DQN algorithm includes a Q network and a target Q network. The target Q network is used to improve the stability of the algorithm. The two networks have the same DNN structure.

[0230] At the beginning of time slot t, the i-th U2U agent follows the ε-greedy strategy and state S i (t) Select an action A from the action space i (t), which means that the agent randomly selects A with probability ε∈(0,1) i (t), or select an action with probability 1-ε according to the following formula:

[0231]

[0232] Among them, Q(S i (t),A i (t),θ) is a given observation state S i (t) and Action A i (t) is the output Q value of the Q network, and θ represents the weight of the Q network, which means that the present invention selects the action with the largest Q value;

[0233] The Q value Q* corresponding to the optimal strategy is obtained according to the following update equation:

[0234]

[0235] Use the gradient descent method to minimize the loss function of the i-th U2U agent to update the weight θ of the Q network as follows:

[0236]

[0237] Among them, γ is the discount factor, α is the learning rate, and the weight variable of the target network is Updated periodically with θ.

[0238] The DRL algorithm based on meta-learning, namely HGNN-Meta-DQN algorithm, is used to enhance the rapid adaptability of resource allocation strategy in dynamic environments. A meta-task set T is defined, which contains N T tasks, each task T j (j=1,...,N T ) is considered as a Markov decision process, for each task T j Define a replay buffer for storing experience The support set and query set of each task are defined as and

[0239] Each task sequentially undergoes a meta-training phase, a meta-adaptation phase, and a deployment phase;

[0240] In the meta-training phase, the support set and query set are used to update the weights of a single task network and the global network, respectively. The meta-training phase adopts a two-level training mechanism. The first level training mechanism is individual-level update, and the other level training mechanism is called global-level update. Each task obtains its own network weight by obtaining different experiences from the replay buffer and using the same optimization problem. Therefore, the network weight of each task can be updated by the following optimization problem for task j:

[0241]

[0242] in, represents the weight of the Q network for task j, is the loss function of the Q network of agent i defined in formula (19), is the replay buffer from task j The experience set sampled from the

[0243] Note that the weights of the network for each task of all U2U agents need to be updated. The gradient descent method is used to update the network parameters, which is expressed as follows:

[0244]

[0245] in, represents the individual-level update learning rate of the Q network, and the superscript n represents the number of iterations. The first iteration (n=1) of the network parameters of each task has the corresponding global network parameters updated, that is, Then the parameters obtained from the previous iteration are updated;

[0246] After all tasks in the batch complete their respective network parameter updates, a global update is performed. This update process is achieved by aggregating the adaptability of each task's training strategy to the newly sampled experience. These loss functions are added together to form a loss function for optimizing the global network parameters, which is expressed as follows:

[0247]

[0248] Based on the gradient descent method, the parameters of formula (23) are updated by the following formula:

[0249]

[0250] Where α is the learning rate for global level updates;

[0251] When the individual-level update and the global-level update are completed, the algorithm enters the next batch and continues to update the global network parameters;

[0252] In the meta-adaptation phase, each U2U agent is able to quickly adapt to new tasks using small sample data based on well-trained network parameters θ. Similar to the parameter update process in individual-level updates, the network parameters for the new task are updated in the following way:

[0253]

[0254] in, At the beginning of the time step, it is initialized to the global parameter θ obtained by training; the experience playback area of ​​the task is D s ;

[0255] Among them, the HGNN network and the meta-DQN network are updated separately. The HGNN network is used to store the reward information obtained by the U2U agent in each subchannel during the global update phase in a matrix corresponding to the number of subchannels, which is recorded as This matrix is ​​used as the label of the corresponding node, and the label is softened in the following way:

[0256]

[0257] Among them, ω represents the weight value of the HGNN network aggregation result, represents the aggregation result of the lagged network, represents the label of the i-th U2U agent;

[0258] The HGNN network is updated using the mean square error function and expressed as follows:

[0259]

[0260] in, Represents the weight parameter of the HGNN network, y i represents the smoothed label used for network update;

[0261] Until both networks converge;

[0262] During the deployment phase, the HGNN parameters and Meta-DQN policy network parameters obtained during the training phase are solidified and deployed to each drone node. Each deployed agent periodically collects local channel state information, interference information, and neighboring node behavior status, and inputs this information into the HGNN network to update node features. The Meta-DQN policy network then jointly selects subchannels and transmit power, outputting the optimal resource allocation strategy. Because the present invention supports a distributed decision-making mechanism, each drone agent corresponding to a U2U link is loaded with a pre-trained model and independently makes resource allocation decisions based on its own observations during system operation.

[0263] During actual operation, each agent utilizes local observations of the current time slot (such as channel gain, interference level, remaining data volume, and remaining time), combined with node embedding vectors generated by the HGNN, to rapidly output resource allocation decisions through the Meta-DQN network forward propagation. Because the joint strategy in this invention is modeled in a composite action space, the corresponding subchannel and transmit power pairs can be obtained in one go during the inference process, reducing the computational overhead of traversing the action space and improving decision-making efficiency.

[0264] To ensure the long-term stability and adaptability of the system in the face of emergencies such as dynamic environmental changes and drone topology adjustments, this invention introduces a rapid meta-adaptation mechanism. Specifically, when the agent detects that communication performance indicators (such as successful transmission rate and system throughput) have fallen below a threshold for multiple consecutive time slots, it automatically triggers a fine-tuning module to update the model weights by sampling small batches of samples from the locally cached experience. This enables rapid response to new tasks or new environments, avoiding the delay and overhead of retraining the model.

[0265] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

[0266] The preferred embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.

Claims

1. A method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning, characterized by: The method comprises the following steps: Step 1: Based on the urban multi-UAV inspection scenario, a multi-UAV communication network scenario model is established, the network channel model and performance indicators are determined, and the communication relationship between adjacent UAVs is obtained; Step 2: Construct the drone's communication network into an undirected graph based on the communication relationship. Determine the neighbors of the nodes in the undirected graph and the features contained in each node. Use the maximum clique search algorithm of the undirected graph to obtain all the maximum cliques in the undirected graph. Construct each calculated maximum clique into a hyperedge to generate a hypergraph structure. Step 3: Based on the constructed hypergraph structure, a graph attention network is used to process the hypergraph structure so that each node obtains feature embedding covering global information; Step 4: The aggregated node features are embedded and input into the Meta-DQN policy network to obtain the optimal channel and power control scheme to maximize the transmission rate of the U2B link from the drone to the base station and the communication success rate of the U2U link from the drone to the drone.

2. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 1, characterized in that: The multi-UAV communication network scenario model established in step 1 specifically includes the scenario area, base station location, and UAV information, wherein the base station location is located at the center of the scenario area, and the UAV information includes the number, location, and speed; Deploy a group of drones S = {1, 2, 3, ..., s} in the scene area D, where each drone flies in a 3-D space and is added to the scene area in a random distribution. The drones are made to travel at a constant speed at a preset initial velocity and direction at different preset heights. The communication between the drones is achieved through U2U links, and the communication between the drones and the base station is achieved through U2B links.

3. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 2, characterized in that: The base station directly allocates spectrum resources for the U2B link, and allocates spectrum resources for the U2U link while taking into account the pre-allocation of spectrum resources for the U2B link. The spectrum resource selection for drones is divided into two parts: sub-channel selection and transmit power selection. Each U2U link can only select one sub-channel and one transmit power level. Assuming that the number of subchannels is the same as the number of U2B links, denoted as m, and that the U2U link has n transmission power levels, each U2U link has a total of m*n resource selection methods. Therefore, the establishment of the channel model specifically includes the following: Assume that all drones are flying at different preset altitudes. At time t, the position of the i-th drone is represented by p i (t) = [x i (t),y i (t),h i ] T ∈R 3 , where t∈[0,T], h i Represents the flight altitude of the i-th UAV; UAV i is moving at a constant speed v at time t. i Toward direction d i (t) moves, so the position of drone i at the next time t+1 is p i (t+1)=p i (t)+v i ·d i (t), the distance between UAV i and UAV j is d i,j (t)=||p i (t)-p j (t)||; The height of the base station is H bs , so the distance between drone i and the base station is recorded as d i,bs (t); The air-to-ground channel is mainly dominated by line-of-sight links (LoS) and non-line-of-sight hybrid links (NLoS). Therefore, the path loss between UAV i and the base station is calculated as follows: The probability calculation formula for the existence of a line-of-sight link between drone i and the ground base station is as follows: Where d0 = max(294.05log 10 h i -432.94,18), p1=233.98log 10 h i -0.95, obviously the non-line-of-sight probability is P i NLoS (t)=1-P i LoS (t); Therefore, the average path loss between drone i and base station is calculated as follows: Taking into account small-scale fading, the channel power gain between UAV i and the base station at time t is expressed as follows: Among them, H i,B (t) represents the fading coefficient between UAV i and the base station; The UAV communication channel is mainly dominated by the line-of-sight link (LoS), and the channel state information is determined by the position of the UAV. Therefore, the channel power gain of the U2U link is characterized by the commonly used free-space path loss model as follows: g i =β0·(d i,j (t)) -α (5) Where β0 represents the channel power gain at the reference distance, and parameter α represents the path loss exponent.

4. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 3, characterized in that: Assume that there are S drones in scenario area D. Each drone needs to communicate with the three nearest drones. Drones communicate directly with each other using device-to-device communication. Therefore, the number of U2U links is 3S. Assume that there are K pairs of U2U users, represented by K = {1, 2, 3, …, k}. Since the U2U link and the U2B link share orthogonally allocated uplink spectrum resources, the signal-to-interference-and-noise ratio (SINR) of the nth U2B link can be expressed as: in, represents the transmit power of the nth U2B link, h n,B [n] represents the instantaneous channel power gain of the nth U2B link on the nth subchannel; σ 2 is the noise power; ρ i [n] means that if the i-th U2U link uses the spectrum resources of the n-th U2B link, then ρ i [n] is assigned a value of 1, otherwise it is assigned a value of 0; It represents the transmit power of the i-th U2U link; It represents the power gain of the interference of the i-th U2U link to the n-th U2B link; According to Shannon's formula, the channel capacity of the nth U2B link is expressed as: Where B represents bandwidth; Similarly, the SINR of the i-th U2U link is expressed as: in, represents the transmit power of the i-th U2U link on the n-th subchannel, g i,i [n] represents the power gain of the i-th U2U link in the n-th subchannel; The interference power of the U2B link sharing the same resource block RB to the i-th U2U link is expressed as follows: in, represents the interference power gain of the nth U2B link to the i-th U2U link; The total interference power of all other U2U links sharing the same resource block RB to the i-th U2U link is expressed as follows: in, represents the interference power gain of the j-th U2U link to the i-th U2U link; The channel capacity of the i-th U2U link is:

5. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 4, characterized in that: The performance indicators include the total transmission rate of the U2B link and the successful transmission of the U2U link. Maximize the total transmission rate of the U2B link A successful transmission for each U2U link is defined as: Where L is the payload size of each U2U link, T and Δ T denote the maximum tolerated delay and the duration of each time slot, t is the time slot index, and t i is the time slot when the i-th U2U link starts transmitting the payload; If the duration of transmitting the payload L does not exceed the maximum tolerable delay T, the information transmission of the U2U link is successful, otherwise it fails. In a given transmission period, the successful transmission probability of all U2U links is expressed as follows: Among them, O i is the number of payloads that the i-th U2U link needs to transmit within the transmission period; It is the indicator of whether the i-th U2U link has successfully transmitted its u-th payload, i.e., when the transmission is successful otherwise Therefore, the optimization problem is formulated as follows: Among them, C1 represents the requirement of the U2U link for successful payload transmission, C2 indicates that each U2U link will only be allocated one resource block, and C3 represents the constraint on the transmission power of the U2U link.

6. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2-1: According to graph theory, the U2U links are regarded as nodes and the interference relations are regarded as simple edges. The UAV communication network is modeled as an undirected graph G = (V, E), and the node set is represented as V = {u1,u2,...,u k }, the set of edges is determined according to the communication relationship between UAVs; Step 2-2: Determine the neighbors of the node and the features of each node. For node i, it contains an initial feature and a list N(i) storing the index of connected hyperedges; the initial features of the node are the channel and interference information observed by the drone itself; At time slot t, the channel information of the i-th U2U link is expressed as It contains the instantaneous channel power gain of its own link transmitter in the nth subchannel Instantaneous interference channel power gain from the nth U2B link transmitter in the nth subchannel The instantaneous interference channel power gain from the jth (j≠i) U2U link transmitter in the nth subchannel At time slot t, the channel information of the nth U2B link is expressed as It contains the instantaneous channel power gain of its own link transmitter on the nth subchannel Instantaneous interference channel gain from U2U link transmitter to base station The interference signal strength of the i-th U2U link receiver in the n-th subchannel in the previous time slot is Therefore, node u i The characteristics are as follows: Among them, || represents the concatenation of vectors; Step 2-3: According to the Bron-Kerbosch algorithm, find all the maximal clusters in the undirected graph and construct each calculated maximal cluster into a hyperedge to obtain the hypergraph structure G = (V, H E ); Where V represents the vertex set, H E represents a hyperedge set.

7. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 1, characterized in that: Based on the constructed hypergraph structure, in the hypergraph structure, the aggregation process of GAT includes the first and second stages; The first stage is the information transfer from U2U nodes to hyperedges. Each hyperedge first receives information from the node it connects to. By calculating the attention coefficient between each node and the hyperedge, the features of all nodes are aggregated according to the weights to generate the representation of the hyperedge. Calculate the attention coefficient of the aggregation from node to hyperedge, for hyperedge e j ∈H E , calculate each node u it connects one by one i ∈N(e j ) and itself: ξ ij =α([Wh i ||Wh j ]) (16) Among them, W is a shared parameter obtained through learning, which linearly transforms the node or hyperedge features and projects them into high dimensions; [·||·] is the i , super edge e j The transformed features are spliced; α(·) is a function that calculates the correlation between nodes and hyperedges, and the spliced ​​high-dimensional features are mapped to a real number through a single-layer feedforward neural network; the vertex features contained in the hyperedge are averaged or weighted summed as the initial features of the hyperedge The attention coefficient obtained by LeakyReLu activation function and softmax normalization is expressed as follows: According to the obtained attention coefficient, the hyperedge e j Aggregate the node features contained in the node and calculate its new feature h' j , the specific formula is as follows: Where σ represents the activation function; The second stage is the information transmission from the hyperedge to the U2U node. The node then receives information from the hyperedge to which it belongs, calculates the node's attention weight to its associated hyperedge, and sums the weighted hyperedge features to update the node's representation. Calculate the attention coefficient of the aggregation from the hyperedge to the node. The workflow is the same as that of the first stage to obtain the attention coefficient β ij , the new feature h of the node obtained by aggregation i ', the specific formula is as follows: Therefore, the first stage operation is performed on each hyperedge in the hypergraph structure, and the second stage operation is performed on each node to obtain the final embedding of each node in the UAV network table view.

8. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 5, characterized in that: The optimization problem of formula (14) is an NP-hard problem, so it is modeled as a Markov decision process to define the state space, action space and reward function of the agent in reinforcement learning; State space: Denote the state space as S i , which contains the states of all agents in each time slot; for each time slot t, agent u i Status S i (t) consists of two parts. The first part is the aggregated features extracted by HGNN, and the second part is the local observation of the intelligent agent. Specifically, the local observation of the UAV is also divided into two parts. The first part is the channel and interference information, including the channel information of the U2U link. U2B link channel information Interference signal strength of the previous time slot of the U2U link The second part includes the number of times the neighboring drone selected the nth subchannel in the previous time slot The remaining number of unsent bits and the remaining sending time under the delay constraint and Therefore, at time slot t, the environment state of the i-th U2U agent is expressed as follows: Action space: Denote the action space as A i , based on the locally observed state and strategy, each U2U agent will decide the sub-channel selection one by one and transmit power selection Since the agent needs to select these two actions at the same time, these two actions are combined into a composite action A. i (t), considering three power levels and m number of sub-channels, there are 3*m actions in total, which can be decomposed as follows: Among them, % represents the modulo operation, / represents division and rounding down; Reward function: The immediate reward is formulated by considering three parts: the total rate of the U2B link, the total rate of the U2U link, and the transmission time. Therefore, the immediate reward for time slot t is expressed as follows: Among them, λ c ,λ p Represents the positive weight of each part, T is the maximum tolerable delay, the expression Indicates the time taken for the transmission.

9. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 8, characterized in that: In step 4, the DQN algorithm includes a Q network and a target Q network, and the two networks have the same DNN structure; At the beginning of time slot t, the i-th U2U agent follows the ε-greedy strategy and state S i (t) Select an action A from the action space i (t), which means that the agent randomly selects A with probability ε∈(0,1) i (t), or select an action with probability 1-ε according to the following formula: Among them, Q(S i (t),A i (t),θ) is a given observation state S i (t) and Action A i The output Q value of the Q network at time (t), θ represents the weight of the Q network; The Q value Q* corresponding to the optimal strategy is obtained according to the following update equation: Use the gradient descent method to minimize the loss function of the i-th U2U agent to update the weight θ of the Q network as follows: θ←θ-α▽ θ L i (i) (26) Among them, γ is the discount factor, α is the learning rate, and the weight variable of the target network is Updated periodically with θ.

10. The method for joint spectrum power allocation in an air-ground integrated network combining hypergraph and meta-reinforcement learning according to claim 9, characterized in that: Using the meta-learning-based DRL algorithm, namely the HGNN-Meta-DQN algorithm, we define a meta-task set T, which contains N T tasks, each task T j (j=1,...,N T ) is considered as a Markov decision process, for each task T j Define a replay buffer for storing experience The support set and query set of each task are defined as and Each task sequentially undergoes a meta-training phase, a meta-adaptation phase, and a deployment phase; In the meta-training phase, the support set and query set are used to update the weights of a single task network and the global network, respectively. The meta-training phase adopts a two-level training mechanism. The first level training mechanism is individual-level update, and the other level training mechanism is called global-level update. Each task obtains its own network weight by obtaining different experiences from the replay buffer and using the same optimization problem. Therefore, the network weight of each task can be updated by the following optimization problem for task j: in, represents the weight of the Q network for task j, is the loss function of the Q network of agent i defined in formula (19), is the replay buffer from task j The experience set sampled from the The gradient descent method is used to update the network parameters, which is expressed as follows: in, represents the individual-level update learning rate of the Q network, and the superscript n represents the number of iterations. The first iteration (n=1) of the network parameters of each task has the corresponding global network parameters updated, that is, Then the parameters obtained from the previous iteration are updated; After all tasks in the batch complete their respective network parameter updates, a global update is performed. This update process is achieved by aggregating the adaptability of each task's training strategy to the newly sampled experience. These loss functions are added together to form a loss function for optimizing the global network parameters, which is expressed as follows: Based on the gradient descent method, the parameters of formula (23) are updated by the following formula: Where α is the learning rate for global level updates; When the individual-level update and the global-level update are completed, the algorithm enters the next batch and continues to update the global network parameters; In the meta-adaptation phase, each U2U agent is able to quickly adapt to new tasks using small sample data based on well-trained network parameters θ. Similar to the parameter update process in individual-level updates, the network parameters for the new task are updated in the following way: in, At the beginning of the time step, it is initialized to the global parameter θ obtained by training; the experience playback area of ​​the task is D s ; Among them, the HGNN network and the meta-DQN network are updated separately. The HGNN network is used to store the reward information obtained by the U2U agent in each subchannel during the global update phase in a matrix corresponding to the number of subchannels, which is recorded as This matrix is ​​used as the label of the corresponding node, and the label is softened in the following way: Among them, ω represents the weight value of the HGNN network aggregation result, represents the aggregation result of the lagged network, represents the label of the i-th U2U agent; The HGNN network is updated using the mean square error function and expressed as follows: in, Represents the weight parameter of the HGNN network, y i represents the smoothed label used for network update; Until both networks converge; During the deployment phase, the HGNN parameters and Meta-DQN policy network parameters obtained during the training phase are solidified and deployed to each drone node. After each deployment, each agent periodically collects local channel state information, interference information, and neighbor node behavior status, and inputs them into the HGNN network to complete the update of node features. The Meta-DQN policy network completes the joint selection of sub-channels and transmission power, and outputs the optimal resource allocation strategy.

Citation Information

Cited By

  • Smart home scene-oriented hypergraph influence maximization method

    CN121257593A

  • A method for maximizing the influence of hypergraphs in smart home scenarios

    CN121257593B

  • Unmanned aerial vehicle airspace resource allocation method based on low-altitude economy

    CN121260052A