Dynamic resource allocation method and system for UAV-supported IoT network under time-varying topology

By building a communication model of the IOT network supported by UAV and a hypergraph-based network conflict model, combining Markov decision-making process and hypergraph convolution Q network, dynamic resource allocation of the Internet of Things network supported by drones is realized, solving the problem of low resource allocation efficiency in high-dynamic and high-density network environments, and significantly improving network performance.

CN119052809BActive Publication Date: 2025-05-02ICLOUDSHIELD SECURITY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411165343.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-05-02
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

In a highly dynamic and high-density network environment, how to achieve dynamic resource allocation of IoT networks supported by drones, thereby improving network performance.

Method used

By building a communication model of the IOT network supported by UAV, the network conflict model based on the hypergraph is transformed into a Markov decision-making process model, and a hypergraph convolution Q network is used for dynamic resource allocation to maximize network throughput and spectrum efficiency.

Benefits of technology

It effectively avoids interference from the Internet of Things network supported by drones, improves network throughput and spectrum efficiency, and is better than existing baseline algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052809B_ABST
    Figure CN119052809B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of Internet of Things (IoT), and specifically discloses a method and system for dynamic resource allocation of UAV (unmanned aerial vehicle) supported IOT network under time-varying topology. Firstly, a communication model of the IOT network supported by UAV is given, and then a network conflict model based on a hypergraph is proposed based on the communication model; then, based on the network conflict model, a resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the IoT network supported by UAV; in order to solve this nonlinear and non-convex resource allocation problem, the resource allocation problem is first converted into a Markov decision process (MDP) model; further, a dynamic resource allocation network based on a hypergraph convolutional Q network (HCQN) is used to solve the MDP model and obtain a dynamic resource allocation strategy. Simulation results show that the method and system can avoid interference in the IoT network supported by UAV, and have good performance in terms of network throughput and spectrum utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things (IoT), and in particular to a method and system for dynamically allocating resources of a UAV (unmanned aerial vehicle) supporting an IoT network under a time-varying topology. Background Art

[0002] A large number of terminal devices generate a large amount of data, and the exchange of this data will bring tremendous pressure to the existing Internet of Things, especially in areas with limited connectivity or severe communication congestion. Drones with high three-dimensional maneuverability can be quickly deployed to provide additional network resources. In order to address the limitations of the processing power of a single drone, multi-drone systems are integrated into the Internet of Things to form a drone-supported Internet of Things network by adopting the Flying Self-Organizing Network (FANET). Compared with traditional ground communication networks, drone-supported Internet of Things networks have the advantages of no environmental restrictions and large-scale service coverage. Drones as wireless communication platforms can provide various Internet of Things applications, such as disaster management, border surveillance, and traffic management.

[0003] However, due to the massive deployment of UAVs, scarce spectrum resources and network densification inevitably pose significant challenges to UAV-supported IoT networks. This challenge stems from the fact that there is severe competition for spectrum resources between UAV-to-UAV (U2U) links and creates complex co-channel interference issues, which will degrade the Quality of Service (QoS) of UAVs. On the other hand, the high mobility of multiple UAVs may lead to a time-varying topology of the entire network, which reduces the effectiveness of resource management as well as spectrum efficiency. In a highly dynamic and high-density network environment, how to achieve dynamic resource allocation for UAV-supported IoT networks and thus improve network performance is a key challenge facing UAV IoT.

[0004] However, current research mainly focuses on the resource allocation of static dense networks or non-dense dynamic networks supported by UAVs, without considering the impact of the high mobility and dense deployment of joint UAVs on network throughput. Summary of the invention

[0005] The present invention provides a method and system for dynamic resource allocation of an IoT network supported by UAV under a time-varying topology, and solves the technical problem of how to realize dynamic resource allocation of an IoT network supported by UAV under a highly dynamic and high-density network environment, thereby improving network performance.

[0006] In order to solve the above technical problems, the present invention provides a dynamic resource allocation method for an IOT network supported by a UAV under a time-varying topology, comprising the steps of:

[0007] S1. Build a communication model for the IOT network supported by UAV;

[0008] S2, constructing a hypergraph-based network conflict model based on the communication model;

[0009] S3. Based on the network conflict model, a resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the UAV-supported IoT network;

[0010] S4, transform the resource allocation problem into a Markov decision process model;

[0011] S5. A dynamic resource allocation network based on a hypergraph convolutional Q network is used to solve the MDP model and obtain a dynamic resource allocation strategy.

[0012] Furthermore, in step S1, the communication model includes a ground control station, i.e., GCS, and N UAVs, wherein the N UAVs establish a flying self-organizing network, i.e., FANET, wherein the UAVs can directly interact with each other in multiple hops without going through infrastructure, and data transmitted between UAVs can only be received by adjacent UAVs; the flying self-organizing network performs a large amount of data exchange, aiming to promote information sharing between multiple UAVs and transmit data back to the GCS, and the GCS provides optimal resource allocation information to the UAVs through FANET.

[0013] Further, in the step S2, based on the network conflict type of the communication model, a hypergraph-based network conflict model is constructed;

[0014] The network conflict types of the communication model include direct conflict and hidden conflict; direct conflict means that two U2U links use the same resource block (RB) in a time slot and have the same sending or receiving node, and the U2U link refers to the communication link between two UAVs; hidden conflict means that two U2U links are allocated the same RB in a time slot, and the receiving end of one U2U link is within the coverage of the sending end of another U2U link;

[0015] The network conflict model at time slot t is constructed as

[0016] is the set of vertices at time slot t, It is the set of hyperedges at time slot t, where vertices refer to U2U links and hyperedges refer to the conflicting relationships between U2U links.

[0017] Furthermore, in step S3, the resource allocation problem is constructed as:

[0018]

[0019] Among them, the binary matrix resource matrix M t×K represents the dimension of the binary matrix resource matrix, M t represents the number of U2U links at time slot t, K represents the total number of resource blocks, represents the decision variable for allocating the kth RB to the mth U2U link at time slot t, Indicates that the kth RB is allocated to the mth U2U link in time slot t, otherwise max means maximization, st means what needs to be satisfied, C1, C2, and C3 represent different constraints; f represents the objective function, and f=[f1,f2] represents the objective function composed of sub-objective function f1 and sub-objective function f2. Indicates the solution of C t The maximum value of satisfies the constraints C1, C2, and C3. f1 represents the sub-objective function for calculating the network throughput of the IoT network supported by the drone, and f2 represents the sub-objective function for calculating the spectrum efficiency of the IoT network supported by the drone. represents the network conflict degree of the IoT network supported by the drone at time slot t calculated based on the network conflict model, Indicates that there is no network conflict in the IoT network supported by the drone, otherwise T represents the total number of time slots.

[0020] Furthermore, the network conflict Calculated by the following formula:

[0021]

[0022] Among them, ||Φ t ||1 represents the matrix Φ t 1-norm of ;

[0023] Network conflict matrix Φ t Calculated by the following formula:

[0024]

[0025] Represents the incidence matrix H of the hypergraph-based network conflict model t The transpose of , max(·,1) means comparing each primitive of the matrix with 1 to get the maximum value, log(·) means performing logarithmic operation on each element of the matrix;

[0026] Correlation matrix H t It is expressed as:

[0027]

[0028] means that the e-th hyperedge contains the m-th U2U link, and otherwise Represents the number of hyperedges in the hyperedge set.

[0029] Furthermore, the sub-objective function f1 is expressed as:

[0030]

[0031] in, represents the achievable data rate of the mth U2U link, calculated as:

[0032]

[0033] represents the weight of the kth RB, represents the signal to interference plus noise ratio of the received signal of the mth U2U link at the kth allocated RB;

[0034] The sub-objective function f2 is expressed as:

[0035]

[0036] Further, Calculated by the following formula:

[0037]

[0038] in, and represent the mth U2U link transmitter and the mth The transmit power of the transmitter of the U2U link, is the channel gain between the transmitter and receiver of the mth U2U link on the kth RB allocated at time slot t, is the kth RB allocated at time slot t. The interference channel gain between the transmitter of the mth U2U link and the receiver of the mth U2U link; σ 2 is the noise power at the receiver of the mth U2U link, and it is assumed that the noise power of all receivers in the system is the same;

[0039] Calculated by the following formula:

[0040]

[0041] represents the path loss of the mth U2U link at time slot t, represents the small-scale fading of the mth U2U link on the kth RB at time slot t;

[0042] Calculated by the following formula:

[0043]

[0044] η Los is the additional attenuation factor due to line-of-sight transmission, c is the speed of light, and f c is the carrier frequency; is the Euclidean distance between the transmitter and receiver of the mth U2U link at time slot t.

[0045] Further, in step S5, in the Markov decision process model, the state at time slot t is defined as:

[0046]

[0047] Among them, HyperGCN() represents the hypergraph convolution, that is, the HyperGCN layer, X t represents the feature matrix of the IoT network supported by the drone at time slot t, represents the hypergraph embedding vector of the mth U2U link at time slot t, N Emb represents the size of the hypergraph embedding vector after HyperGCN(·);

[0048] Feature matrix X t It is expressed as:

[0049]

[0050] represents the feature matrix of the mth U2U link. The observation O(s t ,m) is converted into a feature vector D represents the dimension of the feature vector;

[0051] O(s t ,m) is formulated as:

[0052]

[0053] denote the positions of the receiver and transmitter of the mth U2U link in 3D space, represents the transmission power of the transmitter of the mth U2U link, represents the instantaneous channel information of the mth U2U link, represents the transmission rate of the mth U2U link;

[0054] Resource allocation action of the mth U2U link in time slot t is defined as:

[0055]

[0056] Among them, A is the action space, represents the d-th resource allocation action of the m-th U2U link at time slot t;

[0057] For the entire drone-supported IoT network, the action at time slot t is expressed as:

[0058]

[0059] The reward function r after performing an action in time slot t t is defined as:

[0060]

[0061] Among them, λ1 and λ2 represent positive weights used to balance the two objectives;

[0062] The advantage function is defined as:

[0063] A π (S t ,A t )=Q π (S t ,A t )-V π (S t )

[0064] Among them, Q π (S t ,A t ) means that there is action A under strategy π t Status S t The state-action value, V π (S t ) represents the state S under strategy π t The state value function V π (S t ).

[0065] Furthermore, the step S5 specifically includes the steps of:

[0066] The step S5 specifically includes the following steps:

[0067] S51, constructing a dynamic resource allocation network based on the hypergraph convolutional Q network and the Markov decision process model;

[0068] The dynamic resource allocation network includes an IoT environment supported by a UAV, two super GCN layers, two fully connected layers, a behavior Q network, and a target Q network; the IoT environment supported by a UAV is used to provide real-time environmental status information of the IoT network supported by the UAV. t and Ht ; The super GCN layer is used to process the environment state information X t and H t , generate the corresponding feature representation; the fully connected layer is used to integrate the feature representation output by the super GCN layer to generate the Q value; the behavior Q network selects an action A according to the Q value t and execute;

[0069] S52, training the hypergraph convolutional Q network until the training end condition is reached;

[0070] During the training process, the dynamic resource allocation network also has a replay buffer and a loss function; the UAV-supported IOT network performs action A t After that, we observe the new state X t+1 and H t+1 and immediate reward rt; Experience (X t ,H t ,A t ,r t ,X t +1 ,H t+1 ) is stored in the replay buffer; during the training process, a loss function is calculated by randomly sampling a batch of experiences from the replay buffer to update the network parameters of the behavior Q network and the target Q network; the target Q network is used to provide the target Q value when calculating the loss function value;

[0071] S53. Apply the trained hypergraph convolutional Q network to dynamically allocate resources for the UAV-supported IOT network under time-varying topology.

[0072] The present invention also provides a dynamic resource allocation system for an IOT network supported by a UAV under a time-varying topology, the key of which is that a dynamic resource allocation module is provided to execute steps S1 to S5 in the dynamic resource allocation method.

[0073] The present invention provides a method and system for dynamic resource allocation of an IOT network supported by UAV under time-varying topology. First, a communication model of the IOT network supported by UAV is given. Secondly, a network conflict model based on a hypergraph is proposed based on the communication model to analyze the conflict relationship between U2U links. Then, based on the network conflict model, the influence of interference and time-varying topology is considered at the same time, and the resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the Internet of Things network supported by UAV. In order to solve this nonlinear and non-convex resource allocation problem, the resource allocation problem is first converted into a Markov decision process (MDP) model. A dynamic resource allocation network based on a hypergraph convolutional Q network (HCQN) is further used to solve the MDP model and obtain a dynamic resource allocation strategy. Simulation results show that the method and system can avoid interference in the Internet of Things network supported by UAV, and have good performance in terms of network throughput and spectrum utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a communication model diagram of an IOT network supported by a UAV under a time-varying topology provided by an embodiment of the present invention;

[0075] Figure 2 is an example diagram of a time-varying topology of an Internet of Things network supported by a drone provided by an embodiment of the present invention;

[0076] Figure 3 is an example diagram of the relationship between co-channel interference and network conflict provided by an embodiment of the present invention;

[0077] Figure 4 It is a schematic diagram of the HCQN provided by an embodiment of the present invention for dynamic resource allocation in an Internet of Things network supported by a UAV;

[0078] Figure 5 is a graph showing the relationship between rewards and the number of iterations of a dynamic resource allocation network with an action mask and without an action mask provided by an embodiment of the present invention;

[0079] Figure 6 is a diagram of a dynamic resource allocation process provided by an embodiment of the present invention;

[0080] Figure 7 is a graph showing the relationship between network throughput and time slots of three algorithms provided in an embodiment of the present invention;

[0081] Figure 8 is a histogram of average network throughput of three algorithms provided in an embodiment of the present invention;

[0082] Fig. 9 is a relationship diagram between frequency efficiency and time slot of three algorithms provided in an embodiment of the present invention;

[0083] Fig.10It is a bar chart of average frequency efficiency of three algorithms provided by the embodiments of the present invention. DETAILED DESCRIPTION

[0084] The following specifically illustrates the implementation mode of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0085] The method for dynamic resource allocation of an IOT network supported by a UAV under a time-varying topology provided by an embodiment of the present invention comprises the following steps:

[0086] S1. Build a communication model for the IOT network supported by UAV;

[0087] S2, constructing a hypergraph-based network conflict model based on the communication model;

[0088] S3. Based on the network conflict model, a resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the UAV-supported IoT network;

[0089] S4, transform the resource allocation problem into a Markov decision process model;

[0090] S5. A dynamic resource allocation network based on a hypergraph convolutional Q network is used to solve the MDP model and obtain a dynamic resource allocation strategy.

[0091] (1) Step S1: Constructing the communication model of the IoT network supported by UAV

[0092] The communication model of the Internet of Things network supported by the drone constructed in this embodiment is as follows: Figure 1 As shown, it includes a ground control station (GCS) and N UAVs. The set of N UAVs is represented as The nth UAV is denoted as U n , n=1,2,…,N, or n∈N. N drones are randomly distributed in a bounded 3D free space, and the boundary points are (x max ,0,0),(0,y max ,0),(0,0,z max )(x max Indicates the x-axis boundary point, y max Indicates the y-axis boundary point, z max represents the z-axis boundary point), where the coordinates of the nth UAV, i.e., UAV n, after being projected into the 3D Cartesian coordinates are p n =(x n ,y n ,z n). N drones establish a flying self-organizing network (FANET).

[0093] With all drones establishing FANET, drones are able to interact with each other directly in multiple hops without going through the infrastructure. Due to the channel fading effect and limited transmit power, data transmitted between drones can only be received by adjacent drones. This example assumes that Orthogonal Frequency Division Multiple Access (OFDMA) is used for wireless communications between large-scale drones, which reuse the same spectrum resources in the drone-supported IoT network. The total spectrum resources are equally divided into K available resource blocks (RBs), whose set is denoted as The kth resource is denoted as RB k , k = 1, 2, …, K, or k∈K. In general, FANET performs a large amount of data exchange, aiming to facilitate information sharing among multiple UAVs and data transmission back to the ground control station (GCS). GCS provides optimal resource allocation information to UAVs through FANET. Assume that the IoT network supported by UAVs works in time slot mode, with a total of T time slots, represented by the set T slot ={st1,st2,…,st T}, where the tth time slot is denoted as st t , where for simplicity, the time interval [t, t+1) is called time slot t, t=1,2,…,T, or t∈T.

[0094] Due to the mobility of UAVs and the resulting time-varying topology, the number of U2U (UAV-to-UAV) links will vary over time, e.g. Figure 2 Therefore, in this example, the U2U link set of the IoT network supported by the drone at time slot t is defined as If there are M in the time slot t t U2U links, where the mth U2U link is denoted as UU m ,m=1,2,…,M t , or expressed as m∈M t In addition, let represents the decision variable for allocating the kth RB to the mth U2U link at time slot t, Indicates that the kth RB is allocated to the mth U2U link in time slot t, otherwise Since all drones fly in 3D free space, there is no occlusion between drones, and drones can transmit data according to the rural macrocellular aerial vehicle path loss model. Therefore, the path loss of the mth U2U link at time slot t is It can be expressed as:

[0095]

[0096] Among them, η Los is the additional attenuation factor due to line-of-sight transmission, c is the speed of light, and f c is the carrier frequency. is the Euclidean distance between the transmitter and receiver of the mth U2U link at time slot t. In addition, the Rice distribution is used to define the small-scale fading of the mth U2U link at time slot t on the kth RB The channel gain between the transmitter and receiver of the mth U2U link on the kth allocated RB at time slot t is given by:

[0097]

[0098] Therefore, for the mth U2U link at the allocated kth RB, the signal to interference plus noise ratio (SINR) of the received signal can be expressed as:

[0099]

[0100] in, and represent the mth U2U link transmitter and the mth The transmit power of the transmitter of the U2U link, is the channel gain between the transmitter and receiver of the mth U2U link on the kth RB allocated at time slot t, is the kth RB allocated at time slot t. The interference channel gain between the transmitter of the mth U2U link and the receiver of the mth U2U link; σ 2 is the noise power at the receiver of the mth U2U link, and it is assumed that the noise power is the same for all receivers in the system.

[0101] The achievable data rate of the mth U2U link is calculated as follows:

[0102]

[0103] in, Represents the weight of the kth RB.

[0104] (2) Step S2: Constructing a network conflict model based on a hypergraph

[0105] In the dynamic resource allocation process of the UAV-supported IoT network, the purpose of interference avoidance can be considered as the absence of network conflict in the data transmission at time slot t, which can be expressed as the network conflict degree of the UAV-supported IoT network at time slot t Figure 3 The figure shows the relationship between co-channel interference and network conflict. Figure 3As shown, network conflicts are divided into two types, direct conflicts and hidden conflicts.

[0106] 1) Direct conflict: Two U2U links use the same RB in a time slot and have the same sending or receiving node. Figure 3 For example, in the t=2 time slot, the third RB is allocated to the mth U2U link and the U2U link, the mth U2U link and the The U2U links have the same receiving node, and the receivers of the U2U links will be subject to co-channel interference.

[0107] 2) Hidden conflict: Two U2U links are allocated the same RB in a time slot, and the receiving end of one U2U link is within the coverage of the transmitting end of the other U2U link, e.g. Figure 3 In the time slot t=3, the receiving end of the mth U2U link is located at Within the coverage of the transmitter of the mth U2U link, the receiver of the mth U2U link will suffer from the interference from the mth U2U link due to using the same allocated RB 3 in time slot 3. Co-channel interference from transmitters of a U2U link.

[0108] exist Figure 3 In this paper, each U2U link in the UAV-supported IoT network should be assigned a different color from all its neighbors within two hops, which can also be regarded as graph star coloring, which is more complicated than traditional graph coloring. Therefore, this example establishes a new hypergraph-based network conflict model to deal with the complexity of graph star coloring, and analyzes the network conflict problem in the sequel. For the convenience of analysis, this example abstracts U2U links as vertices and abstracts their mutual conflict relationships as hyperedges.

[0109] Definition 1: Denoted as G H A hypergraph consisting of a finite set of vertices V H and a finite hyperedge set E H composition.

[0110] In particular, at time slot t, this example considers a hypergraph of a dynamic network such as Figure 3 As shown, according to the definition, it is expressed as here

[0111] is the set of vertices at time slot t,

[0112] is the set of hyperedges at time slot t.

[0113] Definition 2: A vertex is a set constructed from all U2U links that need to allocate resources in a UAV-supported IoT network.

[0114] Here, the size of the vertex set is equal to the number of U2U links in the dynamic network with time slot t, i.e. by Figure 3 As an example, according to Definition 2, the vertex set can be expressed as in

[0115] Definition 3: Due to the conflicting relations of U2U links, a hyperedge consists of collective relations within multiple interconnected vertices.

[0116] by Figure 3 As an example, according to Definition 3, Figure 3 As shown, the hyperedge set can be expressed as Where e1 represents the mth U2U link, U2U link and The conflict relationship between the U2U links.

[0117] For hypergraphs, hypergraphs The correlation matrix is ​​expressed as in means that the e-th hyperedge contains the m-th U2U link, and otherwise represents the number of hyperedges in the hyperedge set. In order to t To quantify the degree of conflict, this example uses a binary matrix resource matrix To represent the relationship between U2U links and resources. Therefore, based on the hypergraph association matrix, this example further derives the network conflict matrix Φ t for:

[0118]

[0119] in, Represents the hypergraph incidence matrix H t The transpose of . log(·) means to perform logarithmic operation on each element of the matrix, for example:

[0120]

[0121] max(·,1) means comparing each element of the matrix with 1 to get the maximum value, for example:

[0122]

[0123] In the Network Conflict Matrix In which x e,k > 0 means that the U2U link belonging to the e-th hyperedge is assigned the k-th RB, resulting in a conflict. Otherwise, x e,k ≤0. Therefore, the overall network conflict degree of the entire network It can be calculated as follows:

[0124]

[0125] Among them, ||Φ t ||1 represents the matrix Φ t The 1-norm of Indicates that there is no network conflict in the IoT network supported by the drone, otherwise

[0126] (3) Step S3: Constructing the resource allocation problem

[0127] In the present invention, the goal of this example is to achieve dynamic resource allocation to maximize network throughput and spectrum efficiency while ensuring interference avoidance and time-varying topology requirements in the IoT network supported by drones. Then, the problem of maximizing the network throughput f1 of the entire network in each time slot can be mathematically described as:

[0128]

[0129] Among them, the binary matrix resource matrix C t is the decision vector matrix of the optimization objective.

[0130] Furthermore, the spectral efficiency f2 of the entire UAV-supported IoT network is defined as:

[0131]

[0132] Based on the above analysis, the multi-objective optimization problem can be mathematically expressed as:

[0133]

[0134] Wherein, f=[f1,f2] represents the objective function composed of sub-objective function f1 and sub-objective function f2 (linear combination of sub-objective function f1 and sub-objective function f2), Indicates the solution of C t The maximum value of satisfies the constraints C1, C2, and C3. Constraint C1 means that the resources allocated to each U2U link are conflict-free for the IoT network supporting UAVs. Constraint C2 means that at most one RB is allocated to each U2U link in a time slot. Constraint C3 limits the resource allocation variables. range.

[0135] (4) Step S4: Convert the resource allocation problem into a Markov decision process (MDP) model. The above resource optimization problem is difficult to handle due to the nonlinearity and non-convexity of the objective function and the combinatorial nature of the decision variables. In addition, the problem is further exacerbated in scenarios involving time-varying topology and large-scale dense interference. On this basis, this example first converts the problem into an MDP model, and then proposes a dynamic resource allocation method based on HRL.

[0136] In this example, the MDP model is used to describe the dynamic resource allocation process of the UAV-supported IoT network. For the highly dynamic UAV-supported IoT network scenario, the GCS is defined as an agent that learns a strategy to allocate resources to the system utility, i.e., the system design goal is defined in Eq. (11). For the dynamic resource allocation problem, the four key elements of the formulated MDP model, i.e., the state space, action space, reward function, and value function, are defined as follows:

[0137] 1) State space:

[0138] At time slot t, the nodes and topology information of the conflict hypergraph are expressed as where D is the number of features for each U2U link and is the feature matrix, which conveys the properties of all U2U links. Specifically, the true environment state s t Various information of each U2U link can be included, and GCS collects information that helps the IoT network supporting UAVs at each time slot t. For each U2U link, the following information can be observed.

[0139] ① The location of the receiver of the mth U2U link in 3D space can be defined as:

[0140]

[0141] ② The position of the transmitter of the mth U2U link in 3D space can be defined as:

[0142] ③ The transmit power of the transmitter of the mth U2U link.

[0143] ④ The instantaneous channel information of the mth U2U link can be accurately estimated by the receiver of the mth U2U link at the beginning of each time slot t, and this example assumes that it is also available instantaneously at the transmitter through undelayed feedback.

[0144] ⑤ The transmission rate of the mth U2U link.

[0145] Therefore, the observation function of the mth U2U link at time slot t is formulated as:

[0146]

[0147] in, Then, in order to extract features from high-dimensional observations, this example uses a feature extractor to transform the observation O(s t ,m) is converted into a feature vector In this example, the state space is represented as S, and the state of the mth U2U link is Therefore, the feature matrix X of the entire network t It can be expressed as:

[0148]

[0149] Hypergraph convolution, denoted as HyperGCN, can be used to transform H t and X t As input parameters, the embedding vector of the hypergraph model is then learned by aggregating the features of all U2U links based on hyper-edges. The processing of a HyperGCN layer can be expressed mathematically as:

[0150]

[0151] Where l represents the neural network layer sequence number. Here H and X are H t and X t , D v is the degree matrix of the vertices. W is the diagonal hyperedge weight matrix. D ε is the degree matrix of the hyperedge. is the weight matrix separating the neural network layers at the lth and l+1th layers, and F(l) and F(l+1) represent the activation functions of the lth and l+1th layers, respectively, which are used to calculate the output value of the neurons in the current layer. In addition, F(·) is a nonlinear activation function. Therefore, the embedded state of the IoT network supported by the UAV can be expressed as:

[0152]

[0153] in, The hypergraph embedding vector representing the U2U link, N Emb represents the size of the hypergraph embedding vector after HyperGCN(·).

[0154] 2) Action Space: For the IoT network supported by drones, GCS as an agent can decide the resource allocation action for each time slot, that is, which RBs to allocate to the transmitters of all U2U links. Therefore, the resource allocation action of the mth U2U link in time slot t is (A is the action space) is defined as:

[0155]

[0156] in, It represents the dth resource allocation action of the mth U2U link at time slot t.

[0157] For the entire network, the action of the IoT network supported by the drone at time slot t can be expressed as:

[0158]

[0159] in Resource matrix C that can be directly used as a UAV-supported IoT network t .

[0160] 3) Reward Function

[0161] The reward function of the MDP should be designed based on the objectives of the originally formulated dynamic resource allocation problem (11), which includes maximizing the network throughput of the entire network and maximizing the spectrum efficiency of the entire UAV-supported IoT network. In addition, the network conflict degree constraint of the UAV-supported IoT network should be expressed as a penalty. Therefore, in this paper, the reward function r after executing an action at time slot t is t is defined as:

[0162]

[0163] Among them, λ1 and λ2 represent positive weights used to balance the two objectives.

[0164] 4) Value function: denoted as V π (S t ) represents the state value function in state S t The expected cumulative reward obtained by the agent adopting strategy π when . In the MDP model, the state value function V is used π (S t ) to evaluate each state, defined as:

[0165]

[0166] in, represents the expected value under strategy π, γ τ represents the discount rate at time τ, r t+τ+1 represents the reward at time t+τ+1, τ represents a certain time. Under strategy π, there is action A t Status S t The state-action value function Q π (S t ,A t) is defined as follows:

[0167]

[0168] By following the Bellman criterion, the optimal Q function Q * (S t ,A t ) can be estimated as follows:

[0169] Q*(S t ,A t )= s' [r t +γmax A' Q*(S t+1 ,A')|S t ,A t ] (twenty one)

[0170] s′ represents the state at the next moment (the subscript s′ is used as an identifier), A′ represents the action at the next moment, γ represents the discount rate, and Q * (S t+1 ,A′) represents the optimal state-action value under the state and action at the next moment.

[0171] Therefore, the advantage function can be defined as:

[0172] A π (S t ,A t )=Q π (S t ,A t )-V π (S t ) (twenty two)

[0173] (5) Step S5: Solving the MDP model

[0174] In brief, step S5 specifically includes the following steps:

[0175] S51, constructing a dynamic resource allocation network based on the hypergraph convolutional Q network and the Markov decision process model;

[0176] The dynamic resource allocation network includes an IoT environment supported by a UAV, two super GCN layers, two fully connected layers, a behavior Q network, and a target Q network; the IoT environment supported by a UAV is used to provide real-time environmental status information of the IoT network supported by the UAV. t and H t ; The super GCN layer is used to process the environment state information X t and H t, generate the corresponding feature representation; the fully connected layer is used to integrate the feature representation output by the super GCN layer to generate the Q value; the behavior Q network selects an action A according to the Q value t and execute;

[0177] S52, training the hypergraph convolutional Q network until the training end condition is reached;

[0178] During the training process, the dynamic resource allocation network also has a replay buffer and a loss function; the UAV-supported IOT network performs action A t After that, we observe the new state X t+1 and H t+1 and immediate reward rt; Experience (X t ,H t ,A t ,r t ,X t +1 ,H t+1 ) is stored in the replay buffer; during the training process, a loss function is calculated by randomly sampling a batch of experiences from the replay buffer to update the network parameters of the behavior Q network and the target Q network; the target Q network is used to provide the target Q value when calculating the loss function value;

[0179] S53. Apply the trained hypergraph convolutional Q network to dynamically allocate resources for the UAV-supported IOT network under time-varying topology.

[0180] Based on the formulated MDP model, the present invention proposes a hypergraph convolutional Q network (HCQN / hyperGCN) algorithm to handle the dynamic resource allocation problem (11) of the UAV-supported IoT network under time-varying topology. The entire dynamic resource allocation network is as follows: Figure 4 As shown, it includes an IoT environment (network) supported by drones, two super GCN layers, two fully connected layers, a behavioral Q network, a target Q network, and a replay buffer and loss function during training.

[0181] Figure 4 Each module serves the following purpose:

[0182] UAV-supported IoT environment: This is the input source of the entire model, providing real-time environmental status information, including the interaction relationship between nodes and resource requirements.

[0183] HyperGCN layer: This is a Hypergraph Convolutional Neural Network (HyperGCN) layer that handles information propagation on a hypergraph. It has two layers, each of which contains embedding operations to transform the original node and edge features into a representation suitable for subsequent calculations. Hypergraph Convolutional Network (HyperGCN) is a special type of convolutional neural network that can process hypergraph data, that is, there are higher-level structures (called hyperedges or hypernodes) in addition to nodes and edges.

[0184] Fully connected layers: There are two fully connected layers after the Super GCN layer, which further process the output of the Super GCN layer to facilitate the calculation of the Q function. The fully connected layer integrates the information of nodes and edges to form a global representation.

[0185] Behavioral Q-Network: The behavioral Q-Network is a network used to select actions. It selects actions based on the current state and historical information (H t ) predicts the Q value, that is, the value of each possible action. The Q function estimates the expected cumulative reward after performing an action in a given state.

[0186] Advantage layer and value layer: The behavioral Q network contains an advantage layer and a value layer, both of which are different methods for calculating Q values. The advantage layer focuses on the relative value of an action relative to other actions, while the value layer focuses on the intrinsic value of the state itself.

[0187] Replay buffer: stores past experiences and is used to train the Q network. It saves information such as historical states, actions, next states, and rewards so that it can be updated by randomly sampling during training.

[0188] Small batch samples: A small batch of samples is randomly drawn from the playback buffer for training the Q network.

[0189] Target Q-Network: The target Q-Network is used to calculate the target value. It is a copy of the behavioral Q-Network whose parameters are kept fixed for a period of time and is used to calculate the target value to minimize the loss function.

[0190] The loss function L(θ t ): The loss function measures the gap between the behavior Q network and the target Q network and is used to guide the parameter update of the Q network.

[0191] RB allocation operation: refers to the resource allocation operation, which determines how to allocate resources to each node based on the output of the behavioral Q network.

[0192] The workflow of the entire network is as follows:

[0193] At each time step t, the UAV-supported IoT environment generates a new state X t and H tThe super GCN layer processes this state information and converts it into a suitable representation. The fully connected layer further processes these representations to generate Q values. The behavior Q network selects an action A based on the Q value. t After executing the action, the new state X is observed. t+1 and H t+1 , and the immediate reward r t The new information is added to the replay buffer. The samples in the replay buffer are used to train the behavior Q network, and the parameters θ of the behavior Q network are updated by comparing the Q values ​​of the behavior Q network and the target Q network. t The parameters of the target Q network are The parameters θ of the self-behavior Q network are constantly copied t , but remains fixed for a period of time to stabilize the training process. Through repeated iterations, the behavioral Q network gradually learns the best resource allocation strategy to maximize the long-term cumulative reward.

[0194] In HCQN, there is an agent that performs actions and observes the next state of the system. The agent can then randomly explore the best action by applying a greedy strategy. In particular, HyperGCN adapts to any network topology input by converting a hypergraph model into an embedding vector corresponding to the number of U2U links. The agent then selects actions for each embedding vector using dueling dual deep Q networks (DQNs), ultimately achieving resource allocation across the network. In addition, this example further develops a dynamic resource allocation network based on action masks for hypergraphs.

[0195] 1) Hypergraph-based action mask:

[0196] like Figure 4 As shown, the HCQN architecture is introduced to approximate the action-value function Q π (S t ,A t ). By using the state value and state-related action advantages estimated by the HyperGCN layer, the efficiency of learning the state value function is improved. The coefficient of HCQN is denoted as θ t , which can be trained to π (S t ,A t θ t ) output and Q π (S t ,A t ) is minimized to a sufficiently small level. Therefore, the action value function Q based on HCQN is π (S t ,A t θ t )for:

[0197]

[0198] in It represents the embedding vector d of the mth U2U link state when the action is in time slot t m The value of is estimated. represents the dynamic action value matrix between all U2U links and all allocated resources, that is, the resource allocation scheme can be based on the input network state information X t and H t Dynamically adapt to the highly dynamic changes in the network. However, before the model is trained to have conflict-free resource allocation capabilities, A t Directly determined by the following formula:

[0199]

[0200] This will easily lead to resource conflicts because the resource allocation is obtained directly for each node through the argmax operation, ignoring the node affinity.

[0201] In order to satisfy constraint C1, this example proposes a hypergraph-based action mask to ensure that each resource allocation is conflict-free and further exploit the network conflict relations available in the hypergraph. Therefore, this example adopts a method called invalid action masking to avoid resource conflicts and then reduce interference. This method is generally used to prevent in extended discrete action spaces (i.e., multiple discrete action spaces). According to the association matrix of the hypergraph containing the network conflict relations of the UAV-supported IoT network, the hypergraph-based action mask can be shown in Algorithm 1 shown in Table 1.

[0202] Table 1

[0203]

[0204] Table 1 describes the hypergraph-based action masking algorithm, named HyperMASK (Q π (X H ,H H ,·;θ),H H ), whose input parameters are: Q π Function (Q π (X H ,H H ,·;θ),H H ) and the hypergraph H H , output an estimated Q π value Algorithm 1 mainly includes the following steps:

[0205] 1. From the hypergraph H H Get the adjacency matrix A from H ;

[0206] 2. Traverse all resource types m=1,2,3...M. For each resource type m, traverse all possible allocation schemes m*;

[0207] 3. If resource type m and allocation scheme m * If the correlation between them is 1, then perform operation a m = argmax a (Q π (d m ,a;θ))(find the solution that maximizes Q π Function allocation scheme a m , where d m is the resource demand, a is all the allocation options, and θ is Q π Function parameters);

[0208] 4. Q π The function value is set to a small positive number Much smaller than Q π Minimum value of a function To avoid repeatedly selecting the same allocation scheme;

[0209] 5. Complete all possible allocation schemes m * The loop completes the loop for all resource types m and returns the estimated Q π The value is the result of the algorithm.

[0210] 2) Parameter update

[0211] HCQN combines the advantages of duel DQN and double DQN, such as Figure 4 As shown, it can more accurately estimate the action value function and select actions. HCQN overcomes the instability of the activation function of the neural network due to its own characteristics by introducing the target Q network. According to the idea of ​​time difference, the target value can be expressed as formula (25). The agent receives a reward r from the environment t , and then use the experience (X t ,H t ,A t ,r t ,X t+1 ,H t+1 ) is stored in the replay buffer. To minimize the loss, the loss function is calculated by randomly sampling a small batch of experiences from the replay buffer and then used to update the network parameters. t ) is calculated by the mean square error defined as (26). In order to minimize L(θ t ) parameter θ t, using the gradient-based method, which can be defined as Equation (27). Therefore, the parameter update with action mask can be defined as Equation (28).

[0212]

[0213] in, represents the target Q value, which is the immediate reward r t Add the discount factor γ and multiply it by the maximum estimated Q value at the next moment. t ) represents the loss function, which is defined as the behavior Q network (with parameters θ t ) and the target Q value. This loss function is used to train the behavioral Q network to make it closer to the target Q value. is the gradient used in the gradient descent method, which represents the loss function with respect to the parameter θ t The gradient of . α is the learning rate when the behavior Q network parameters are updated. || ||1 means 1 norm.

[0214] In summary, the above parameter updating process is summarized in Algorithm 2 as shown in Table 2.

[0215] Table 2

[0216]

[0217] The specific process of the HCQN-based resource allocation algorithm can be described as follows:

[0218] The algorithm first initializes some parameters, such as the replay buffer, learning rate, discount factor, and behavioral Q network. It then runs in a series of iterations, each of which resets the state from the environment and starts a new round. At each time step t, the algorithm selects an action based on the behavioral Q network and some randomness, performs this action in the environment, and then records the reward and the new state. These experiences are stored in the replay buffer. The algorithm then randomly extracts a batch of experiences from the replay buffer and updates the parameters of the behavioral Q network using gradient descent. After a certain number of time steps, the parameters of the target Q network are updated. Finally, the algorithm returns the trained behavioral Q network parameters.

[0219] In order to implement the above method, an embodiment of the present invention further provides a dynamic resource allocation system for an IOT network supported by a UAV under a time-varying topology, which is provided with a dynamic resource allocation module for executing steps S1 to S5 in the above dynamic resource allocation method.

[0220] The following is a simulation.

[0221] 1) Parameter settings

[0222] All simulation experiments were performed on Dell servers using Python 3.9.13, Pytorch 2.0.1, NetworkX 3.2, and PyTorchGeometric 2.3.1. To demonstrate the performance of the proposed algorithm (labeled as HCQN), this example compares the proposed dynamic resource allocation network with baseline algorithms (labeled as A2C and Random). For easy reference, the main system parameters used in the simulation are summarized in Table 3.

[0223] Table 3 Simulation parameters

[0224]

[0225] 2) Convergence performance of training

[0226] This example first evaluates the convergence performance of the proposed dynamic resource allocation network in a drone-supported IoT network. The training of the HGCN model includes iterative updates, and the number of iterations is set to 1000, such as Figure 5 As the number of iterations increases, both the update methods with and without action masks show Figure 5 The results show that the updating without action mask leads to the gradual convergence of the reward curve in the above figure. Updating the HGCN parameters without action mask may cause the parameters to fail to converge to the optimal result. Moreover, after convergence with a sufficient number of iterations, the updating with action mask produces a slightly higher cumulative reward than the updating without action mask. This result suggests that during the dynamic resource allocation process of UAV-supported IoT networks, the updating without action mask leads to ineffective learning, resulting in lower cumulative reward values ​​compared to the updating with action mask. Therefore, compared with the updating without action mask, the updating with action mask has the potential advantage of avoiding the local optimal space through diversified exploration.

[0227] 3) Performance comparison of different algorithms

[0228] In order to demonstrate the performance of the proposed algorithm, especially its performance gain over some existing algorithms, the proposed algorithm is compared with existing A2C algorithms and random algorithms. Figure 6 The simulation results of the proposed dynamic resource allocation network, A2C algorithm and random algorithm in the allocation process of 10ms, 20ms, 30ms, 40ms and 50ms are given in the paper. Figure 6 (a), (b) and (c) correspond to the proposed dynamic resource allocation network, A2C algorithm and random algorithm respectively. Figure 6It can be seen that the random algorithm will cause interference in the reused RBs due to its uncertainty, while the proposed dynamic resource allocation network and A2C algorithm are machine learning algorithms that can evenly use all available RBs to avoid network performance degradation caused by overuse of a single RB. In addition, the proposed dynamic resource allocation network effectively avoids interference through action masks, which means that it uses RBs more efficiently than the A2C algorithm, which will allow the network to achieve higher network throughput and more efficient spectrum efficiency.

[0229] The proposed dynamic resource allocation network is further compared with other dynamic resource allocation schemes in terms of network throughput. Figure 7 As shown. Figure 7 It can be seen that in the dynamic resource allocation process, the throughput of the proposed scheme is much higher than other schemes. In the time slot of 20ms to 30ms, the topology of the IoT network supported by drones is complex and the structure is highly dynamic, and the overall network throughput shows a downward trend. In addition, the proposed dynamic resource allocation network can effectively handle the negative impact of high dynamic changes through the HGCN model and action mask, improve the anti-interference ability of the drone IoT network, and achieve higher network throughput than the A2C algorithm and the random algorithm.

[0230] Figure 8 The average network throughput of different dynamic resource allocation schemes in the UAV-supported IoT network is shown. Compared with the random algorithm, the proposed dynamic resource allocation network and the A2C algorithm have learning capabilities. Through model training, the ability to avoid interference can be effectively improved, thereby improving the network throughput. Figure 8 As shown in Figure 2, the proposed dynamic resource allocation network performs better than other solutions. The average network throughput of the proposed scheme is 8.17Mbps higher than that of the A2C algorithm and 10.81Mbps higher than that of the random algorithm. In addition, since the dynamic resource allocation network has HGCN and action mask, it can more effectively improve the network performance and obtain higher network throughput than the A2C algorithm.

[0231] The performance of the proposed dynamic resource allocation network is compared with other dynamic resource allocation networks and Fig. 9 Through continuous training of the neural network model, the proposed dynamic resource allocation network and A2C algorithm improve the ability of resource reuse and reduce the number of overall resource usage, thereby ensuring Fig. 9 The spectrum efficiency of the network is improved under the condition of high network performance. However, due to its uncertainty, the spectrum efficiency obtained by the random algorithm fluctuates more than that of the intelligent algorithm, and it is impossible to guarantee high spectrum efficiency under the premise of high network throughput.

[0232] Fig.10The average spectrum efficiency under different resource allocation schemes is shown. Fig.10 , the proposed HCQN algorithm performs better than other solutions. The average network efficiency of the scheme is 0.35 bit / s / Hz higher than that of the A2C algorithm and 0.44 bit / s / Hz higher than that of the random algorithm. In addition, because the HCQN algorithm has HGCN and hypergraph mask, it can better improve the resource allocation performance and obtain higher spectrum efficiency than the A2C algorithm.

[0233] In summary, the dynamic resource allocation method and system of the UAV-supported IOT network under time-varying topology provided by the present invention firstly gives the communication model of the UAV-supported IOT network, and then proposes a network conflict model based on a hypergraph based on the communication model to analyze the conflict relationship between U2U links; then based on the network conflict model, the influence of interference and time-varying topology is considered at the same time, and the resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the UAV-supported Internet of Things network; in order to solve this nonlinear and non-convex resource allocation problem, the resource allocation problem is first converted into a Markov decision process (MDP) model; further, a dynamic resource allocation network based on a hypergraph convolutional Q network (HCQN) is used to solve the MDP model and obtain a dynamic resource allocation strategy. Simulation results show that the proposed method and system can avoid interference in the UAV-supported Internet of Things network, and are superior to the existing baseline algorithms in terms of network throughput and spectrum efficiency.

[0234] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A dynamic resource allocation method for an IOT network supported by UAV under a time-varying topology, characterized in that: Includes steps: S1. Build a communication model for the IOT network supported by UAV; S2, constructing a hypergraph-based network conflict model based on the communication model; S3. Based on the network conflict model, a resource allocation problem is constructed with the goal of maximizing the network throughput and spectrum efficiency of the UAV-supported IoT network; S4, transform the resource allocation problem into a Markov decision process model; S5. A dynamic resource allocation network based on a hypergraph convolutional Q network is used to solve the MDP model and obtain a dynamic resource allocation strategy.

2. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 1, characterized in that: In step S1, the communication model includes a ground control station, i.e., GCS, and N UAVs. The N UAVs establish a flying self-organizing network, i.e., FANET. The UAVs can directly interact with each other in multiple hops without going through infrastructure. Data transmitted between UAVs can only be received by adjacent UAVs. The flying self-organizing network performs a large amount of data exchange, aiming to promote information sharing between multiple UAVs and transmit data back to the GCS. The GCS provides the best resource allocation information to the UAVs through FANET.

3. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 2, characterized in that: In the step S2, based on the network conflict type of the communication model, a hypergraph-based network conflict model is constructed; The network conflict types of the communication model include direct conflict and hidden conflict; direct conflict means that two U2U links use the same resource block (RB) in a time slot and have the same sending or receiving node, and the U2U link refers to the communication link between two UAVs; hidden conflict means that two U2U links are allocated the same RB in a time slot, and the receiving end of one U2U link is within the coverage of the sending end of another U2U link; The network conflict model at time slot t is constructed as is the set of vertices at time slot t, It is the set of hyperedges at time slot t, where vertices refer to U2U links and hyperedges refer to the conflicting relationships between U2U links.

4. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 3 is characterized in that: In step S3, the resource allocation problem is constructed as: Among them, the binary matrix resource matrix M t ×K represents the dimension of the binary matrix resource matrix, M t represents the number of U2U links at time slot t, K represents the total number of resource blocks, represents the decision variable for allocating the kth RB to the mth U2U link at time slot t, Indicates that the kth RB is allocated to the mth U2U link in time slot t, otherwise max means maximization, st means what needs to be satisfied, C1, C2, and C3 represent different constraints; f represents the objective function, and f=[f1,f2] represents the objective function composed of sub-objective function f1 and sub-objective function f2. Indicates the solution of C t The maximum value of satisfies the constraints C1, C2, and C3. f1 represents the sub-objective function for calculating the network throughput of the IoT network supported by the drone, and f2 represents the sub-objective function for calculating the spectrum efficiency of the IoT network supported by the drone. represents the network conflict degree of the IoT network supported by the drone at time slot t calculated based on the network conflict model, Indicates that there is no network conflict in the IoT network supported by the drone, otherwise T represents the total number of time slots.

5. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 4 is characterized in that: Network Conflict Calculated by the following formula: Among them, ||Φ t ||1 represents the matrix Φ t 1-norm of ; Network conflict matrix Φ t Calculated by the following formula: Represents the incidence matrix H of the hypergraph-based network conflict model t The transpose of , max(·,1) means comparing each primitive of the matrix with 1 to get the maximum value, log(·) means performing logarithmic operation on each element of the matrix; Correlation matrix H t It is expressed as: means that the e-th hyperedge contains the m-th U2U link, and otherwise Represents the number of hyperedges in the hyperedge set.

6. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 5, characterized in that: The sub-objective function f1 is expressed as: in, represents the achievable data rate of the mth U2U link, calculated as: represents the weight of the kth RB, represents the signal to interference plus noise ratio of the received signal of the mth U2U link at the kth allocated RB; The sub-objective function f2 is expressed as:

7. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 6 is characterized in that: Calculated by the following formula: in, and represent the mth U2U link transmitter and the mth The transmit power of the transmitter of the U2U link, is the channel gain between the transmitter and receiver of the mth U2U link on the kth RB allocated at time slot t, is the kth RB allocated at time slot t. The interference channel gain between the transmitter of the mth U2U link and the receiver of the mth U2U link; σ 2 is the noise power at the receiver of the mth U2U link, and it is assumed that the noise power of all receivers in the system is the same; Calculated by the following formula: represents the path loss of the mth U2U link at time slot t, represents the small-scale fading of the mth U2U link on the kth RB at time slot t; Calculated by the following formula: η Los is the additional attenuation factor due to line-of-sight transmission, c is the speed of light, and f c is the carrier frequency; is the Euclidean distance between the transmitter and receiver of the mth U2U link at time slot t.

8. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 7, characterized in that: In step S5, in the Markov decision process model, the state at time slot t is defined as: Among them, HyperGCN() represents the hypergraph convolution, that is, the HyperGCN layer, X t represents the feature matrix of the IoT network supported by the drone at time slot t, represents the hypergraph embedding vector of the mth U2U link at time slot t, N Emb represents the size of the hypergraph embedding vector after HyperGCN(·); Feature matrix X t It is expressed as: represents the feature matrix of the mth U2U link. The observation O(s t ,m) is converted into a feature vector D represents the dimension of the feature vector; O(s t ,m) is formulated as: denote the positions of the receiver and transmitter of the mth U2U link in 3D space, represents the transmission power of the transmitter of the mth U2U link, represents the instantaneous channel information of the mth U2U link, represents the transmission rate of the mth U2U link; Resource allocation action of the mth U2U link in time slot t is defined as: Among them, A is the action space, represents the d-th resource allocation action of the m-th U2U link at time slot t; For the entire drone-supported IoT network, the action at time slot t is expressed as: The reward function r after performing an action in time slot t t is defined as: Among them, λ1 and λ2 represent positive weights used to balance the two objectives; The advantage function is defined as: A π (S t ,A t )=Q π (S t ,A t )-V π (S t ) Among them, Q π (S t ,A t ) means that there is action A under strategy π t Status S t The state-action value, V π (S t ) represents the state S under strategy π t The state value function V π (S t ).

9. The method for dynamic resource allocation of an IOT network supported by UAV under a time-varying topology according to claim 8, characterized in that: The step S5 specifically comprises the following steps: S51, constructing a dynamic resource allocation network based on the hypergraph convolutional Q network and the Markov decision process model; The dynamic resource allocation network includes an IoT environment supported by a UAV, two super GCN layers, two fully connected layers, a behavior Q network, and a target Q network; the IoT environment supported by a UAV is used to provide real-time environmental status information X of the IoT network supported by the UAV t and H t ; The super GCN layer is used to process the environment state information X t and H t , generate the corresponding feature representation; the fully connected layer is used to integrate the feature representation output by the super GCN layer to generate the Q value; the behavior Q network selects an action A according to the Q value t and execute; S52, training the hypergraph convolutional Q network until the training end condition is reached; During the training process, the dynamic resource allocation network also has a replay buffer and a loss function; the UAV-supported IOT network performs action A t After that, we observe the new state X t+1 and H t+1 and immediate reward rt; experience (X t ,H t ,A t ,r t ,X t+1 ,H t+1 ) is stored in the replay buffer; during the training process, a loss function is calculated by randomly sampling a batch of experiences from the replay buffer to update the network parameters of the behavior Q network and the target Q network; the target Q network is used to provide the target Q value when calculating the loss function value; S53. Apply the trained hypergraph convolutional Q network to dynamically allocate resources for the UAV-supported IOT network under time-varying topology.

10. A dynamic resource allocation system for UAV-supported IOT networks under time-varying topology, characterized by: A dynamic resource allocation module is provided, which is used to execute steps S1 to S5 in the dynamic resource allocation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • B5G mass Internet of Things resource allocation method and system based on hypergraph

    CN117376355A

  • Internet of Things resource conflict-free allocation method and system based on graph reinforcement learning

    CN117440442A