Unmanned aerial vehicle multi-hop network intelligent resource allocation method based on deep reinforcement learning

By proposing an intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning, the signal conflict and QoS issues in wireless UAV ad hoc networks are solved, achieving efficient resource allocation and QoS guarantee, and improving network performance.

CN121334883APending Publication Date: 2026-01-13ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511410317.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Traditional VNE technology cannot effectively handle signal conflicts and interference in wireless unmanned aerial vehicle (UAV) ad hoc networks, lacks QoS awareness capabilities, and is difficult to achieve efficient resource scheduling and policy decision-making in complex and dynamic environments.

Method used

We adopt a deep reinforcement learning-based intelligent resource allocation method for UAV multi-hop networks. By constructing a PPO model and combining GCN and MPNN, we optimize the resource allocation strategy and comprehensively consider the embedding success rate, resource cost-effectiveness ratio and QoS performance to achieve self-learning and optimization of the strategy.

Benefits of technology

It improves the task embedding efficiency and resource utilization of UAV multi-hop networks, ensures the QoS performance of the network in dynamic environments, and realizes intelligent scheduling and collaborative optimization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334883A_ABST
    Figure CN121334883A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle multi-hop network intelligent resource allocation method based on deep reinforcement learning, and the method comprises the steps: abstracting task input into a virtual network, abstracting an unmanned aerial vehicle cluster into a physical network, and regarding a resource allocation process as embedding from the virtual network to the physical network. And processing a single node in the virtual network every time by taking the network state information as input, and updating the physical network resource information according to an output result. And a return value is obtained by counting the embedding success rate, evaluating the resource cost-effectiveness ratio and the network QoS (Quality of Service) performance. And calculating a dominant function expectation corresponding to the action by using the intelligent resource allocation network, and updating the strategy under constraint based on a strategy gradient optimization method until the network converges. Finally, task input is converted into a virtual network request, the virtual network request is input into the trained resource allocation network, an optimal embedding strategy is obtained, and efficient allocation of resources is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of UAV self-organizing networks, and particularly to an intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning. Background Technology

[0002] In recent years, the application of unmanned aerial vehicle (UAV) swarm technology in both civilian and military fields has attracted widespread attention, demonstrating its enormous potential in scenarios such as disaster relief, environmental monitoring, and logistics transportation. Today, as UAV swarms grow in size, the communication overhead and management complexity within the swarms are also increasing. Driven by the increasing complexity and diverse needs of missions, traditional dedicated self-organizing network solutions based on single scenarios can no longer meet the requirements for flexibility and scalability.

[0003] Wireless ad hoc networks, a key technology for UAV swarm communication, are characterized by the ability to spontaneously form networks to complete tasks without the need for fixed infrastructure. However, this network architecture faces numerous challenges in terms of high dynamism, limited resources, and diverse Quality of Service (QoS) requirements. How to dynamically allocate physical nodes to meet different QoS requirements based on unique task needs has become a core issue in UAV swarm network design.

[0004] Virtual Network Embedding (VNE) technology offers a flexible and effective solution. By mapping Virtual Network Requests (VNRs) to the physical network layer, VNE can dynamically schedule tasks for different cluster topologies and achieve reasonable resource allocation, meeting task requirements while optimizing resource utilization.

[0005] However, current VNE technology is mainly applied in wired network environments, and it still has many shortcomings when facing wireless UAV ad hoc networks with unstable links and limited resources. On the one hand, traditional VNE methods are generally designed for wired network scenarios and cannot handle signal conflicts and interference that may occur during the concurrent execution of multiple tasks, leading to performance degradation or failure of the embedding strategy. On the other hand, existing embedding models often aim to maximize the mapping success rate or resource utilization, lacking a comprehensive consideration of network QoS indicators, and making it difficult to guarantee the actual needs of tasks in terms of latency, bandwidth, and stability. In addition, traditional VNE algorithms generally adopt heuristic or approximate algorithms, lacking the ability to optimize long-term benefits, and are difficult to achieve efficient resource scheduling and policy decision-making in complex and uncertain dynamic environments.

[0006] Therefore, there is an urgent need for a virtual network embedding method that can adapt to wireless network environments, has QoS awareness capabilities, and can achieve policy self-learning and optimization, so as to improve the intelligent level of resource allocation and service assurance capabilities of UAV multi-hop networks in complex mission scenarios. Summary of the Invention

[0007] The present invention aims to at least partially solve one of the technical problems in the related art.

[0008] Therefore, the purpose of this invention is to propose an intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning, which can improve the embedding success rate, resource utilization and QoS performance of dispatched tasks.

[0009] To achieve the above objectives, a first aspect of the present invention proposes an intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning, comprising:

[0010] The task input is abstracted into a virtual network with resource requirements and network topology, the drone swarm is abstracted into a physical network with resource capacity and network topology, and the resource allocation process is abstracted into an embedding process from the virtual network to the physical network.

[0011] The network state information is used as the input to the intelligent resource allocation network. The network state information includes the topology of the physical network, available resource capacity, resource allocation status, and resource requirements of the virtual network.

[0012] Each time a single node in the virtual network is considered, the resource allocation result of the virtual network, i.e. the embedding result, is output according to the intelligent resource allocation network, and the resource information of the physical network is updated.

[0013] The embedding success rate is obtained based on the statistical data of virtual network embedding. The resource cost-effectiveness ratio is evaluated based on the embedding benefits and resource consumption of the virtual network. The network QoS performance is evaluated based on the QoS prediction module to obtain three indicators. The reward value is obtained based on the three indicators.

[0014] The advantage function corresponding to different actions is calculated by the intelligent resource allocation network, and then the expectation of the advantage function is obtained. The policy is updated under the preset magnitude constraint based on the policy gradient optimization method, and the resource allocation network is trained until convergence.

[0015] A series of task inputs are converted into virtual network requests, which are then input into a trained resource allocation network to obtain a suitable embedding strategy.

[0016] In addition, the intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning according to the above embodiments of the present invention may also have the following additional technical features:

[0017] Furthermore, in one embodiment of the present invention, the step of using network state information as input to the intelligent resource allocation network in a multi-hop network includes:

[0018] The topology information of the current physical network is collected as input to the PPO model of the intelligent resource allocation network. This topology includes the number of nodes in the physical network, the connection relationships between physical nodes, and the connection weights. Resource capacity information and resource allocation status of the physical network are also collected, expressed as s. p_net ={c free ,c max ,b free ,b max ,emb}, where c free and c max b represents the node's remaining computing resources and total computing resources, respectively. free and b max These represent the node's remaining bandwidth resources and total bandwidth resources, respectively; emb indicates whether the physical node is occupied; resource demand information and resource allocation status of virtual nodes are collected and expressed as... Where c req and b req These represent the node's computing resource requirements and bandwidth resource requirements, respectively. pand It represents the proportion of embedded nodes in a virtual network.

[0019] Furthermore, in one embodiment of the present invention, the step of outputting the virtual network embedding result according to the intelligent resource allocation network and updating the resource information of the physical network includes:

[0020] The state s = {s} is obtained by interacting with the environment through the intelligent resource allocation network. p_net ,s v_node}, then generate the action space distribution A = {(a i ,p i )|i=0,1,…,n-1}},a i This represents the mapping action that maps the current virtual node to the physical node numbered i. i Choose action a under the current strategy i The probability is given by n, where n represents the total number of physical nodes in the physical network. After the embedding operation is completed, the remaining resource information of the relevant physical nodes and physical links is updated, and the embedding result is used as environmental feedback for policy training and optimization of the reinforcement learning model.

[0021] Furthermore, in one embodiment of the present invention, the reward function of the intelligent resource allocation network is defined as:

[0022]

[0023] Among them, G vRepresents the virtual network to be embedded; |N v | Represents the total number of virtual nodes in the virtual network; R2C(G v ) represents the resource cost-effectiveness ratio after the virtual network is embedded, that is, the ratio of resource benefits to resource costs; W represents the weight value of the QoS performance improvement of the entire physical network after the virtual node and its associated links are successfully embedded, which is generated by the QoS prediction module.

[0024] Furthermore, in one embodiment of the present invention, the objective function of the policy gradient optimization process is defined as:

[0025]

[0026] in, This represents the ratio of the probability of taking action a in state s to the probability of taking action a under the current policy versus the probability of taking action a under the old policy. The advantage function representing the current state-action pair is approximated using the generalized advantage estimation method. This represents the reward value resulting from taking action a in state s. The value function represents state s; ∈ is the clipping factor, used to limit the magnitude of policy updates; clip() is the clipping function, used to clip the policy ratio to the range [1-∈, 1+∈] to avoid over-updating the policy.

[0027] To achieve the above objectives, a second aspect of the present invention provides a deep reinforcement learning-based intelligent resource allocation device for unmanned aerial vehicle (UAV) multi-hop networks, comprising the following modules:

[0028] The input module is used to input network status information as input to the intelligent resource allocation network, wherein the network status information includes the topology of the physical network, available resource capacity, resource allocation status, and resource requirements of the virtual network.

[0029] The update module is used to output the resource allocation results of the virtual network according to the intelligent resource allocation network, and update the resource information of the physical network.

[0030] The evaluation module is used to obtain the embedding success rate based on the statistical results of virtual network embedding, evaluate the resource cost-effectiveness ratio based on the embedding benefits and resource consumption of the virtual network, and evaluate the network QoS performance based on the QoS prediction module, so as to obtain three indicators and obtain the reward value based on the three indicators.

[0031] The training module is used to calculate the advantage function corresponding to different actions through the intelligent resource allocation network, thereby obtaining the expectation of the advantage function, updating the policy based on the policy gradient optimization method under the preset magnitude constraint, and training the resource allocation network until convergence.

[0032] The generation module is used to convert a series of task inputs into virtual network requests, input the virtual network requests into a trained resource allocation network, and obtain a suitable embedding strategy.

[0033] To achieve the above objectives, a third aspect of the present invention provides a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning as described above.

[0034] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the above-described intelligent resource allocation method for UAV multi-hop networks based on deep reinforcement learning.

[0035] This invention proposes a resource allocation method for unmanned aerial vehicle (UAV) networks in multi-hop transmission network environments. This method introduces deep reinforcement learning to abstract the resource allocation process into a Virtual Object (VNE) problem. By constructing a PPO-based VNE model and taking the resource status of the UAV physical network and the resource requirements of task requests as input, it outputs a resource mapping strategy to guide the embedding process of different tasks in the multi-hop network. This method effectively improves the embedding efficiency of task requests and the utilization rate of physical resources, while ensuring the QoS performance of the network in dynamic environments, achieving intelligent scheduling and collaborative optimization of resources in UAV multi-hop networks. Attached Figure Description

[0036] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0037] Figure 1 This is a schematic diagram of a virtual network embedding scenario for an unmanned aerial vehicle (UAV) self-organizing network system according to an embodiment of this application.

[0038] Figure 2 This is a schematic diagram of the virtual network embedding process according to an embodiment of this application;

[0039] Figure 3 This is a flowchart of an intelligent resource allocation method for unmanned aerial vehicle (UAV) ad hoc networks based on deep reinforcement learning, according to an embodiment of this application.

[0040] Figure 4 This is a schematic diagram of the structure of an intelligent resource allocation model based on deep reinforcement learning according to an embodiment of this application;

[0041] Figure 5 These are schematic diagrams of three physical network topologies used in the experimental simulation;

[0042] Figure 6 These are simulation graphs showing the embedding performance of resource allocation methods and other methods based on deep reinforcement learning under different network structures.

[0043] Figure 7 This is a schematic diagram of an intelligent resource allocation device for a multi-hop UAV network based on deep reinforcement learning, provided in an embodiment of this application. Detailed Implementation

[0044] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0045] The following describes an embodiment of the present invention, with reference to the accompanying drawings, a method for intelligent resource allocation in UAV multi-hop networks based on deep reinforcement learning.

[0046] Figure 1 This is a schematic diagram of a virtual network embedding scenario for an unmanned aerial vehicle (UAV) self-organizing network system according to an embodiment of this application. Figure 1 As shown, the system adopts a two-layer architecture, including an upper virtual network request layer and a lower physical resource layer. The virtual network request layer contains multiple virtual network units, each forming a specific network topology, including multiple virtual nodes and virtual links, each corresponding to different computing and communication resource requirements to reflect diverse service request characteristics. The physical resource layer consists of multiple UAV nodes, forming a dynamically changing network cluster. Each node establishes a physical link connection based on its communication range. Both nodes and links have quantifiable resource attributes, such as computing power and bandwidth capacity. Cross-layer connections represent the mapping relationship between virtual nodes and virtual links to physical nodes and physical links, characterizing the embedding process of virtual network requests in the physical network.

[0047] Figure 2This is a schematic diagram of the virtual network embedding process according to an embodiment of the present invention. The process includes two stages: virtual node embedding and virtual link embedding. In the virtual node embedding stage, it is necessary to ensure that the computing resource capacity of the selected physical node meets the resource requirements of the corresponding virtual node. This embodiment employs a reinforcement learning-based policy generation module to generate a mapping policy distribution based on the state information of the node to be embedded in the current virtual network, and thereby determine the target physical node of the virtual node. In the virtual link embedding stage, it is necessary to ensure that the bandwidth resource capacity of the physical path meets the transmission requirements of the virtual link, and on this basis, improve QoS performance. This embodiment uses a QoS prediction module to evaluate the performance of candidate physical paths that meet the bandwidth requirements, and selects the path with the best QoS performance to complete the mapping of the virtual link.

[0048] Figure 3 This is a flowchart of an intelligent resource allocation method for unmanned aerial vehicle (UAV) ad hoc networks based on deep reinforcement learning, according to an embodiment of this application. Specifically, it includes the following steps:

[0049] Step S310: The task input is abstracted into a virtual network with resource requirements and network topology, the drone swarm is abstracted into a physical network with resource capacity and network topology, and the resource allocation process is abstracted into an embedding process from the virtual network to the physical network.

[0050] In the embodiments of this application, the task input is abstracted as a virtual network, represented as an undirected weighted graph G. v =(N v ,L v ), where N v and L v These represent the sets of virtual nodes and virtual links, respectively. For each node n... v N v Considering CPU requirements For each link l v L v Considering bandwidth requirements The drone swarm is abstracted as a physical network, represented as an undirected weighted graph G. s =(N s ,L s ), where N s and L s Let n represent the drone nodes and the set of communication links, respectively. For each node n... s N s Considering CPU capacity For each link l s L s Considering bandwidth capacity The resource allocation process is abstracted as a mapping process, represented as a mapping relationship from virtual network to physical network.

[0051] Step S320: The network status information is used as input to the intelligent resource allocation network. The network status information includes the topology of the physical network, available resource capacity, resource allocation status, and resource requirements of the virtual network.

[0052] In this embodiment, the state information of the entire physical network and the state information of a single virtual node to be embedded are considered. The state information of the entire physical network can be represented as s. p_net ={c free ,c max ,b free ,b max ,emb}, where c free and c max b represents the node's remaining computing resources and total computing resources, respectively. free and b max These represent the node's remaining bandwidth resources and total bandwidth resources, respectively. `emb` indicates whether the physical node is occupied. For a single virtual node, its state information can be represented as... Where c req and b req These represent the node's computing resource requirements and bandwidth resource requirements, respectively. pand It represents the proportion of embedded nodes in a virtual network.

[0053] Step S330: Each time, consider a single node in the virtual network, output the resource allocation result of the virtual network according to the intelligent resource allocation network, i.e., the embedding result, and update the resource information of the physical network.

[0054] In this embodiment, the virtual node embedding stage employs a Graph Convolutional Network (GCN) to extract the topology and resource features of the physical network, while simultaneously using a Multi-Layer Perceptron (MLP) to extract the resource requirement features of the virtual nodes. Based on these features, a PPO-based reinforcement learning algorithm is used to generate a mapping strategy for the virtual nodes. The output is a set of probability distributions, representing the likelihood of the virtual node embedding into each physical node. Guided by this probability distribution, the embedding scheme for the current virtual node can be determined. The virtual link embedding stage first obtains K candidate physical paths that satisfy bandwidth resource constraints based on the shortest path algorithm. Then, a QoS prediction module is used to evaluate the service quality of each candidate path, selecting the physical path with the best performance to generate the virtual link embedding scheme.

[0055] It is worth noting that, in this embodiment, the QoS prediction module is built on a Message Passing Neural Network (MPNN) and trained offline using network operation data obtained through simulation on the OPNET platform. This model takes the feature information of the physical nodes and links involved in all working paths in the physical network as input and outputs a service quality prediction index for the corresponding path, which is used to assist in the optimal path selection during the virtual link mapping process.

[0056] Step S340: Obtain the embedding success rate based on the statistical results of virtual network embedding, evaluate the resource cost-effectiveness ratio based on the embedding benefits and resource consumption of the virtual network, evaluate the network QoS performance based on the QoS prediction module, and obtain three indicators. The reward value is obtained based on the three indicators.

[0057] In this embodiment, the reward function is designed by comprehensively considering factors such as VNR embedding success rate, resource cost-effectiveness, and QoS performance. By jointly modeling the above multiple indicators, the constructed reward function not only ensures the smooth embedding of task requests but also effectively guides the policy network to improve the allocation efficiency of physical resources, achieve higher resource utilization, and maintain good network communication performance.

[0058] In one possible implementation, the reward function r of the intelligent resource allocation network t Defined as:

[0059]

[0060] Among them, G v Represents the virtual network to be embedded; |N v | Represents the total number of virtual nodes in the virtual network; R2C(G v ) represents the resource cost-effectiveness ratio after the virtual network is embedded, that is, the ratio of resource benefits to resource costs; W represents the weight value of the QoS performance improvement of the entire physical network after the virtual node and its associated links are successfully embedded, which is generated by the QoS prediction module.

[0061] Step S350: Calculate the advantage function corresponding to different actions through the intelligent resource allocation network, and then obtain the expectation of the advantage function. Update the policy based on the policy gradient optimization method under a certain amplitude constraint, and train the resource allocation network until convergence.

[0062] In this embodiment, a PPO-based intelligent resource allocation model is used to generate and update the strategy. A schematic diagram of the model's structure is shown below. Figure 4 .like Figure 4As shown, the PPO-based intelligent resource allocation model can include an environment, an actor network, a critic network, a memory experience pool, and a PPO objective function. The main functional descriptions of each module are as follows:

[0063] The environment module provides the current network status. t Action a is selected based on the policy distribution output by the Actor network. t ~π θ (a|s t It generates a feedback reward r based on the set reward function. t The Actor network receives network status s transmitted from the environment. t And output strategy π θ (a|s t The Critic network evaluation is applied to state s. t Perform an evaluation and output the state value function. The Memory experience pool is used to store the state-action-reward sequence (s) generated during the interaction between the agent and the environment. t ,a t ,r t ,s t+1 This is used for subsequent estimation of return value and advantage function.

[0064] The objective of the policy gradient optimization based on PPO is to maximize the expected value of the advantage function, which can be expressed as:

[0065]

[0066] in, This represents the ratio of the probability of taking action a in state s to the probability of taking action a under the current policy versus the probability of taking action a under the old policy. The advantage function representing the current state-action pair is approximated using the generalized advantage estimation method. V represents the reward value resulting from taking action a in state s. πθ (s) represents the value function of state s; ∈ is the clipping factor, used to limit the magnitude of policy updates; clip() is the clipping function, used to clip the policy ratio to the range [1-∈, 1+∈] to avoid over-updating the policy.

[0067] Through the above mechanism, the resource allocation strategy can converge stably during the training process, effectively improving strategy performance and ensuring resource utilization and service quality during virtual network embedding.

[0068] Step S360: Convert a series of task inputs into virtual network requests, input the virtual network requests into the trained resource allocation network, and obtain a suitable embedding strategy.

[0069] Figure 5 These are schematic diagrams of three physical network topologies used in the experimental simulation. Specifically, from left to right: Physical Topology 1 contains 30 physical nodes, and the node spatial distribution is generated using the Rejection Sampling (RS) method; Physical Topology 2 also contains 30 physical nodes, and the node spatial distribution is generated using the Poisson Point Process (PPP) method; Physical Topology 3 contains 100 physical nodes, and the node distribution is also generated based on the rejection sampling method.

[0070] Figure 6 This is a simulation graph showing the embedding performance of resource allocation methods based on deep reinforcement learning and other methods under different network architectures. To verify the effectiveness of the DRL-VNE algorithm proposed in this application, it is compared and evaluated with existing VNE algorithms, such as NRM-RANK and A3C-GCN. The experiment is based on... Figure 5 The figure illustrates three different physical network topologies, comparing the embedding performance of each algorithm under different VNR arrival rates. The horizontal axis represents the VNR arrival rate, i.e., the number of virtual network requests submitted per unit time, reflecting the frequency of task issuance; the vertical axis represents multiple embedding performance metrics, including embedding success rate, average resource cost-effectiveness, average network latency, and packet loss rate. Through multi-dimensional metric comparison, the comprehensive optimization capabilities of the proposed algorithm in terms of resource utilization and communication performance can be fully evaluated.

[0071] Because the NRM-RANK algorithm is a heuristic method, its adaptability to multiple optimization objective scenarios is limited. This results in significantly worse embedding performance metrics compared to the A3C-GCN algorithm and the DRL-VNE algorithm proposed in this application across the three physical network topologies. In contrast, the A3C-GCN algorithm performs better in terms of network embedding rate and average cost-effectiveness ratio, especially in complex physical network environments with 100 nodes, where it maintains an embedding success rate exceeding 0.7 even when the VNR arrival rate is between 0.5 and 1.0. Furthermore, its average cost-effectiveness ratio remains consistently above 0.55 across all three topologies and different VNR arrival rates, demonstrating good resource utilization efficiency. However, because this algorithm does not fully consider link QoS performance in the virtual link embedding stage and reward function design, its performance fluctuates significantly in metrics such as average network latency and packet loss rate, resulting in poor performance. The DRL-VNE algorithm proposed in this application comprehensively considers multiple metrics, including embedding success rate, resource cost-effectiveness ratio, and QoS performance, during policy generation and path selection. While maintaining an embedding success rate similar to A3C-GCN, the algorithm improves network cost-effectiveness and achieves better performance in terms of average latency and packet loss rate, resulting in more stable and superior QoS guarantees. These results demonstrate that the algorithm proposed in this application can effectively guide the resource allocation process of task inputs in UAV networks, achieving the dual optimization goals of resource utilization efficiency and communication quality.

[0072] Figure 7 This is a schematic diagram of an intelligent resource allocation device for a multi-hop UAV network based on deep reinforcement learning, provided in an embodiment of this application. Figure 7 As shown, the device includes the following modules:

[0073] The input module 710 is used to input network status information as input to the intelligent resource allocation network, wherein the network status information includes the topology of the physical network, available resource capacity, resource allocation status and resource requirements of the virtual network.

[0074] The update module 720 is used to output the resource allocation results of the virtual network according to the intelligent resource allocation network, and update the resource information of the physical network.

[0075] The evaluation module 730 is used to obtain the embedding success rate based on the statistical results of virtual network embedding, evaluate the resource cost-effectiveness ratio based on the embedding benefits and resource consumption of the virtual network, evaluate the network QoS performance based on the QoS prediction module, and obtain three indicators to obtain the reward value based on the three indicators.

[0076] Training module 740 is used to calculate the advantage function corresponding to different actions through the intelligent resource allocation network, thereby obtaining the expectation of the advantage function, updating the policy based on the policy gradient optimization method under the preset magnitude constraint, and training the resource allocation network until convergence.

[0077] The generation module 750 is used to convert a series of task inputs into virtual network requests, input the virtual network requests into a trained resource allocation network, and obtain a suitable embedding strategy.

[0078] This invention presents an intelligent resource allocation device for UAV multi-hop networks based on deep reinforcement learning. By introducing deep reinforcement learning, the resource allocation process is abstracted into a Virtual Object (VNE) problem. By constructing a PPO-based VNE model and taking the resource status of the UAV physical network and the resource requirements of task requests as input, a resource mapping strategy can be output to guide the embedding process of different tasks in the multi-hop network. This method can effectively improve the embedding efficiency of task requests and the utilization rate of physical resources, while ensuring the QoS performance of the network in dynamic environments, achieving intelligent scheduling and collaborative optimization of resources in UAV multi-hop networks.

[0079] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0080] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0081] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for intelligent resource allocation of a multi-hop network of unmanned aerial vehicles based on deep reinforcement learning, characterized in that, The method comprises the following steps: abstracting a task input as a virtual network with resource requirements and network topology, abstracting a UAV cluster as a physical network with resource capacity and network topology, and abstracting a resource allocation process as an embedding process of the virtual network into the physical network; inputting state information of the network into the intelligent resource allocation network, wherein the state information of the network comprises a topology structure, available resource capacity, resource allocation, and resource requirements of the virtual network of the physical network; outputting a resource allocation result of the virtual network, i.e., an embedding result, from the intelligent resource allocation network and updating resource information of the physical network each time a single node in the virtual network is considered; obtaining an embedding success rate according to the embedding of the virtual network, evaluating a resource cost-benefit ratio according to embedding benefits and resource losses of the virtual network, and evaluating a QoS performance of the network according to a QoS prediction module to obtain three indexes, and obtaining a return value according to the three indexes; calculating an advantage function corresponding to different actions by using the intelligent resource allocation network, obtaining an expected value of the advantage function, and updating a policy under a preset amplitude constraint based on a policy gradient optimization method to train the resource allocation network until convergence; converting a series of task inputs into virtual network requests, inputting the virtual network requests into the trained resource allocation network, and obtaining a suitable embedding strategy.

2. The method of claim 1, wherein, The resource allocation process is abstracted as a virtual network embedding process, which comprises: A task input is abstracted into a virtual network, which is composed of a plurality of virtual nodes with resource requirements and their connection relationship, expressed as an undirected weighted graph G v = (N v , L v ), wherein N v and L v represent the virtual node and virtual link set respectively; The UAV cluster is abstracted as a physical network consisting of a plurality of physical nodes with computing resources and bandwidth resources and wireless communication links therebetween, expressed as an undirected weighted graph G s = (N s , L s ), where N s and L s represent the set of physical nodes and physical links, respectively; selecting physical nodes that meet constraint conditions in the physical network according to resource requirements of virtual nodes in the virtual network to form a node mapping relationship, finding a multi-hop path that meets bandwidth constraints and QoS limits in the physical network according to a logical connection relationship between the virtual nodes to embed a virtual link, and forming a link mapping relationship, and constructing a complete embedding mapping of the virtual network into the physical network.

3. The method of claim 1, wherein, The network state information is inputted into the intelligent resource allocation network in the multi-hop network, which comprises: Topology information of the current physical network is collected, which includes the number of nodes of the physical network, the connection relationship between the physical nodes and the connection weight; resource capacity information and resource allocation state of the physical network are collected, expressed as s p_net ={c free ,c max ,b free ,b max ,emb} where c free and c max represent the remaining computing resources and the total computing resources of the node respectively, b free and b max represent the remaining bandwidth resources and the total bandwidth resources of the node respectively, and emb represents whether the physical node is occupied; resource requirement information and resource allocation state of the virtual node are collected, expressed as s vnode ={c req ,b req ,r pand} where c req and b req represent the computing resource requirement and the bandwidth resource requirement of the node respectively, and r pand represents the proportion of embedded nodes in the virtual network. encoding the network state information into a state vector as an input of a Proximal Policy Optimization (PPO) model of the intelligent resource allocation network to guide generation of a resource allocation policy.

4. The method of claim 1, wherein, The embedding result of the virtual network outputted from the intelligent resource allocation network is used to update the resource information of the physical network, which comprises: by the intelligent resource allocation network to obtain the state s = {s p_net , s v_node}, and then generate an action space distribution A = {(a i , p i )|i = 0, 1, …, n-1}}, a i represents a mapping action of mapping the current virtual node to the physical node numbered i, p i represents the probability of selecting the action a i under the current policy, and n represents the total number of physical nodes in the physical network; The mapping action comprises two stages: a first stage of embedding a current virtual node into a candidate physical node, i.e., mapping of the virtual node, and a second stage of selecting a multi-hop physical path that meets bandwidth requirements between corresponding physical nodes according to a link relationship between the virtual nodes, i.e., mapping of the virtual link; in the path selection process, a QoS prediction module is also used to evaluate a service quality index of the path to preferentially select a path that meets a preset QoS performance requirement to ensure overall performance of the network; After the embedding operation is completed, remaining resource information of the related physical nodes and physical links is updated, and the embedding result is used as an environment feedback for policy training and optimization of the reinforcement learning model.

5. The method of claim 1, wherein, The reward function of the intelligent resource allocation network is defined as follows: For resource allocation decisions in a certain state, the system will calculate the reward value based on multiple standards according to the embedding result of the current virtual network, and realize feedback of different degrees: wherein G v represents the current virtual network to be embedded; |N v represents the total number of virtual nodes in the virtual network; R2C(G v ) represents the resource efficiency ratio of the virtual network after embedding, i.e. the ratio of resource benefit to resource cost; and W represents the weight value of the improvement of the QoS performance of the entire physical network after successful embedding of the virtual nodes and their associated links, generated by the QoS prediction module.

6. The method of claim 1, wherein, The objective function of the policy gradient optimization process is defined as: where, represents the probability ratio of the current policy to the old policy taking action a in state s; represents the advantage function of the current state-action pair, which is approximated using the generalized advantage estimation method, represents the return value of taking action a in state s, represents the value function of state s; ∈ is a clipping factor used to limit the magnitude of policy updates; clip() is a clipping function used to clip the policy ratio to the range of [1-∈, 1+∈] to avoid excessive policy updates.

7. A deep reinforcement learning-based intelligent resource allocation device for multi-hop UAV networks, characterized in that, It comprises the following modules: An input module for inputting network state information as an input of the intelligent resource allocation network, wherein the network state information comprises a topology structure, available resource capacity, resource allocation, and resource requirements of a virtual network of a physical network; An update module for outputting a resource allocation result of the virtual network by the intelligent resource allocation network and updating resource information of the physical network; An evaluation module for obtaining an embedding success rate according to an embedding situation of the virtual network, evaluating a resource cost-benefit ratio according to embedding benefits and resource losses of the virtual network, and evaluating network QoS performance according to a QoS prediction module to obtain three indexes, and obtaining a reward value according to the three indexes; A training module for calculating an advantage function corresponding to different actions by the intelligent resource allocation network, obtaining an advantage function expectation, updating a policy under a preset amplitude constraint based on a policy gradient optimization method, and training the resource allocation network until convergence; A generation module for converting a series of task inputs into a virtual network request, inputting the virtual network request into the trained resource allocation network, and obtaining a suitable embedding strategy.

8. A computer device, comprising: The computer program is executed by the processor to implement the deep reinforcement learning-based intelligent resource allocation method for a multi-hop network of unmanned aerial vehicles according to any one of claims 1-6.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the deep reinforcement learning-based intelligent resource allocation method for a multi-hop network of unmanned aerial vehicles according to any one of claims 1-6.