Graph attention and double-Q network driven agent dynamic resource allocation method

Through the dynamic resource allocation method of the agent driven by graph attention and dual Q network, the problem of traditional scheduling methods in low-altitude logistics networks is solved that it is difficult for traditional scheduling methods to cope with dynamic task requirements and coordinate the complementarity of drones and ground unmanned vehicles, and efficient coordination, dynamic adaptation and resource optimization are achieved, and task completion rate and efficiency are improved.

CN120218573AActive Publication Date: 2025-06-27湖南工商大学

Patent Information

Application Number
CN202510696264.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Traditional scheduling methods in low-altitude logistics networks are difficult to cope with dynamic task requirements, coordinate the complementarity between drones and ground unmanned vehicles, effectively model dynamic environmental interference, difficult to balance global optimization and local efficiency, and significantly increase in computing complexity when the task scale is expanded.

Method used

The dynamic resource allocation method of the agent driven by graph attention and dual Q network is adopted, and the graph structure of state information and task information is generated through the self-attention mechanism. The graph convolution network and graph attention network calculate the matching attention weight between the task and the agent, and intelligent decisions are made based on the comprehensive Q value and individual Q value to realize the dynamic resource allocation of the agent in the edge server.

Benefits of technology

It realizes efficient coordination between drones and ground unmanned vehicles in dynamic mission scenarios, improves task completion rate, reduces task delay, energy consumption and cost, dynamically adapts to complex environments, and balances global optimality with local optimality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218573A_ABST
    Figure CN120218573A_ABST
Patent Text Reader

Abstract

The invention relates to a graph attention and double-Q network driven agent dynamic resource allocation method, which is applied to an edge server and comprises the following steps: receiving task information issued by a central control platform and state information uploaded by an agent in real time; generating a graph structure of state information and task information through a self-attention mechanism, and enabling the graph structure to sequentially pass through a graph convolutional network and a graph attention network to obtain attention weights of edges; calculating a matching attention weight between the task and the intelligent agent based on the attention weight of the edge; calculating a comprehensive Q value based on the matching attention weight and the individual Q value of the agent; the current network in the double-depth Q network selects the corresponding action of the intelligent agent in the corresponding state based on the comprehensive Q value, and the target network in the double-depth Q network calculates the evaluation value of the action; the state comprises state information and task information, and the action is that the agent executes a task; and allocating a corresponding task to the intelligent agent in real time based on the action with the highest evaluation value in the corresponding state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of resource allocation, and particularly to an intelligent agent dynamic resource allocation method driven by graph attention and double Q network. Background Art

[0002] The low-altitude logistics network is a new type of logistics transportation system constructed based on unmanned equipment such as unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs). Utilizing low-altitude airspace resources, it realizes efficient delivery of goods, information, and services in areas such as towns, rural areas, and industrial parks. There are many deficiencies in traditional scheduling methods in the low-altitude logistics network, such as: 1. The personnel requirements in the low-altitude logistics network are highly dynamic (such as emergency deliveries, sudden tasks, etc.), and traditional static scheduling methods are difficult to respond to.

[0003] 2. There are significant differences in resource capabilities (computing, load capacity, endurance) between UAVs (lightweight tasks, fast response) and UGVs (computation-intensive, heavy transportation), and traditional scheduling methods are difficult to coordinate their complementarity.

[0004] 3. The low-altitude logistics network faces dynamic environmental interferences (such as weather changes, obstacles, no-fly zones, traffic congestion), and traditional scheduling methods cannot effectively model these uncertainties.

[0005] 4. Traditional scheduling methods often have difficulty balancing global optimization (such as total system energy consumption) and local efficiency.

[0006] 5. When the task scale expands, the computational complexity of traditional optimization algorithms (such as linear programming) increases significantly, making it difficult to meet the requirements of real-time scheduling. Summary of the Invention

[0007] Based on this, it is necessary to provide an intelligent agent dynamic resource allocation method driven by graph attention and double Q network. This method is applied to an edge server and includes: S1: Receiving task information sent by the central control platform and status information uploaded by the intelligent agent in real time; the intelligent agent includes a UAV cluster and / or a UGV; S2: Generating a graph structure of the status information and task information through a self-attention mechanism, and passing the graph structure through a graph convolutional network and a graph attention network in sequence to obtain the attention weights of the edges; calculating the matching attention weights between the task and the intelligent agent based on the attention weights of the edges; S3: Calculating a comprehensive Q value based on the matching attention weights and the individual Q values of the intelligent agent; S4: The current network in the double deep Q network selects the corresponding action of the intelligent agent in the corresponding state based on the comprehensive Q value, and the target network in the double deep Q network calculates the evaluation value of the action; the state includes status information and task information, and the action is for the intelligent agent to execute the task; S5: Based on the action with the highest evaluation value in the corresponding state, allocate the corresponding task to the agent in real time.

[0008] Preferably, the task information includes: the size of the task data, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.

[0009] Preferably, the state information includes the agent state information and the environmental state information; The agent state information includes: the current 3D position coordinates of the agent, the current speed of the agent, the current remaining power of the agent, the current available computing resources of the agent, the progress of the task currently executed by the agent; The environmental state information includes: obstacles, weather, no-fly zones or the road traffic conditions for the unmanned vehicle to travel.

[0010] Preferably, the graph structure for generating the state information and the task information through the self-attention mechanism includes: Perform feature embedding on the state information and the task information respectively to obtain the agent state feature representation and the task feature representation; Take the agent state feature representation and the task feature representation as the nodes of the graph structure, and take the connection relationship between the agent state feature representation and the task feature representation as the edges of the graph structure; The calculation of the connection relationship includes: Multiply the task feature representation by the first parameter matrix, and multiply the agent state feature representation by the second parameter matrix; Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio; Pass the first ratio through the softmax function to obtain the weight of the edge; Traverse all the task feature representations, calculate the weights of the edges between all the task feature representations and each agent state feature representation, and obtain the affinity matrix of the state information and the task information.

[0011] Preferably, in S2, the process of obtaining the matching attention weight includes: Use the k-means clustering algorithm to dynamically adjust the edges between the nodes in the graph structure; Normalize the affinity matrix to obtain the adjacency matrix of the graph; Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure by the feature representation of any node in the graph structure and the third parameter matrix, and pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node; Traverse all the nodes in the graph structure, update the feature representations of all the nodes, and obtain the embedding representation of the graph; Input the embedded representation of the graph into the graph attention network. In the graph attention network, multiply the embedded representation of the graph by the third parameter matrix, and pass the obtained second product through the LeakyReLU activation function to obtain the attention weights of the edges in the graph. Calculate the matching attention weights between the task and the agent based on the attention weights of the edges.

[0012] Preferably, the specific calculation of the matching attention weights is as follows: Pass the attention weights of the edges and negative infinity through the indicator function, and pass the obtained value through the softmax function to obtain the matching attention weights between the task and the agent; The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.

[0013] Preferably, the calculation formula for the comprehensive Q value is: Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the agents ; Where, represents the comprehensive Q value corresponding to the i th agent; represents the state at the t th moment; represents the action selected at the t th moment; M represents the number of agents; represents the matching attention weights between the task and the m th agent; represents the individual Q value of the m th agent; represents the network parameters of the m th agent.

[0014] Preferably, the training process of the double deep Q network includes: Define the state, action, and immediate reward; Create two deep Q networks with the same structure, respectively as the current network and the target network; The current network selects an action based on the state information and task information, and adopts the ε-greedy strategy; the ε-greedy strategy randomly selects an action with probability ε or selects the corresponding action in the state where the agent has the maximum comprehensive Q value; Initialize the experience replay pool; The agent executes the selected action, updates the state, and calculates the immediate reward; Take the state before the update, the selected action, the immediate reward, and the state after the update as experience, and store it in the experience replay pool; When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experience from the experience replay pool, select the optimal action in the updated state using the current network, and calculate the evaluation value of the optimal action using the target network; Pass the optimal action in the updated state selected by the current network through the current network to obtain the Q value of the current network; Based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network, calculate the mean squared error loss function; minimize the mean squared error loss function and backpropagate to update the Q value of the current network; Every fixed number of steps, perform a soft update on the network parameters of the target network based on the network parameters in the Q value of the current network.

[0015] Preferably, the calculation formula for the immediate reward is: ; ; ; ; where represents the immediate reward at the t th moment; represents the task completion reward; represents the path optimization reward; represents the resource utilization reward; represents the first weight coefficient; represents the second weight coefficient; represents the third weight coefficient; K represents the total number of tasks; represents whether the k th task is completed, taking the value of 1 if completed and 0 if not completed; represents the delay tolerance of the k th task; M represents the number of agents; represents the path length of the i th agent; represents the energy consumption of the i th agent in path planning; represents the resources actually used by the i th agent; represents the total available resources of the i th agent; represents the idle time of the i th agent.

[0016] Preferably, the update formula for the Q value of the current network is: ; Among them, represents the Q value after the current network update; represents the state at the t moment; represents the action at the t moment; represents the network parameters after the current network update; represents the Q value before the current network update; represents the network parameters before the current network update; represents the learning rate; represents the t instantaneous reward at the moment; represents the discount factor; represents the Q value of the target network; represents the network parameters of the target network; represents the t state at the +1 moment; represents the Q value of the optimal action of the current network in the state .

[0017] Beneficial effects: The method can achieve dynamic optimization through intelligent decision-making. The goal is to achieve efficient cooperation between drones and ground unmanned vehicles in dynamic task scenarios and improve the task completion rate; the method reduces task latency, energy consumption, and costs. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is the flowchart of the intelligent agent dynamic resource allocation method driven by graph attention and double Q network in the embodiments of the present application. Detailed Embodiments

[0020] In order to make the above objects, features, and advantages of the present application more obvious and understandable, the following will make a detailed description of the specific embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description in order to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0021] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0022] As Figure 1 shown, this embodiment provides an intelligent agent dynamic resource allocation method driven by graph attention and double Q-network. This method is applied to an edge server and includes: S1: Receive the task information sent by the central control platform and the status information uploaded by the intelligent agent in real time; the intelligent agent includes a drone cluster and / or a driverless vehicle.

[0023] In this embodiment, the task information includes: the task data size, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.

[0024] The status information includes intelligent agent status information and environmental status information; The intelligent agent status information includes: the current 3D position coordinates of the intelligent agent, the current speed of the intelligent agent, the current remaining power of the intelligent agent, the current available computing resources of the intelligent agent, and the progress of the task currently executed by the intelligent agent; The environmental status information includes: obstacles, weather, no-fly zones, or the road traffic conditions for the driverless vehicle to travel.

[0025] S2: Generate a graph structure of the status information and the task information through the self-attention mechanism, and sequentially pass the graph structure through a graph convolutional network and a graph attention network to obtain the attention weights of the edges; calculate the matching attention weights between the task and the intelligent agent based on the attention weights of the edges.

[0026] Specifically, the generation of the graph structure of the status information and the task information through the self-attention mechanism includes: Perform feature embedding on the status information and the task information respectively to obtain the intelligent agent status feature representation and the task feature representation; Use the intelligent agent status feature representation and the task feature representation as the nodes of the graph structure, and use the connection relationship between the intelligent agent status feature representation and the task feature representation as the edges of the graph structure; The connection relationship calculation includes: Multiply the task feature representation by the first parameter matrix, and multiply the intelligent agent status feature representation by the second parameter matrix; Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio; Pass the first ratio through the softmax function to obtain the weight of the edge; Traverse all task feature representations, calculate the weights of the edges between all task feature representations and each agent state feature representation, and obtain the affinity matrix of state information and task information.

[0027] Generate the affinity matrix of the graph through the self-attention mechanism, and use the graph structure to represent the relationship between tasks and agents. These graph structures provide a basis for the subsequent calculation of matching attention weights.

[0028] Furthermore, the process of obtaining the matching attention weights includes: Use the k-means clustering algorithm to dynamically adjust the edges between nodes in the graph structure to enhance the connection relationship between nodes; Normalize the affinity matrix to obtain the adjacency matrix of the graph; Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure, the feature representation of any node in the graph structure, and the third parameter matrix. Pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node; Traverse all nodes in the graph structure, update the feature representations of all nodes, and obtain the embedding representation of the graph; Input the embedding representation of the graph into the graph attention network. In the graph attention network, multiply the embedding representation of the graph by the third parameter matrix. This is a linear transformation used to map the embedding representation to a new space to better capture the relationship between nodes. Pass the obtained second product through the LeakyReLU activation function. The LeakyReLU activation function can introduce non-linearity to help the model capture the complex relationship between nodes. Compared with the traditional ReLU activation function, the LeakyReLU activation function also has an output in the negative value region, avoiding the problem of gradient disappearance; Calculate the attention weights of each edge in the graph through the above linear transformation and LeakyReLU activation function. The attention weights of the edges reflect the association strength between nodes. The higher the value, the more important the relationship between nodes; Calculate the matching attention weights between tasks and agents based on the attention weights of the edges.

[0029] Even further, the specific calculation of the matching attention weights is as follows: Pass the attention weights of the edges and negative infinity through the indicator function, and pass the obtained value through the softmax function to obtain the matching attention weights between tasks and agents; The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.

[0030] Through the fusion of information, the most important features in the graph embedding can be focused on, the matching attention weights between the evaluation task and the agent can be evaluated, which provides a basis for the subsequent calculation of the comprehensive Q value and further improves the effect of task allocation.

[0031] In task scheduling, the goal is to assign tasks to the most suitable agents (UAVs or UGVs). Each task has specific computing requirements, latency requirements, and data requirements, and the selection of agents needs to consider the matching of these requirements with resources. In this embodiment, the matching problem can be modeled as: ; ; ; where, represents the matching degree between the task and the resources (such as the matching degree between computing resources and battery power, etc.), , 1 represents a perfect match, and 0 represents a complete mismatch; represents the matching degree between the task and the environment (such as the matching degree between obstacle avoidance and weather impact, etc.), , 1 represents a perfect match, and 0 represents a complete mismatch; represents the k-th task; represents the resource information; represents the environmental information; represents the weight coefficient between the task and the resources; represents the weight coefficient between the task and the environment; , , , are the coefficients of each resource factor respectively; Computation, Storage, Energy, Communication represent computing power, storage capacity, required energy consumption, and communication capacity; , , , are the coefficients of each environmental factor respectively; Weather, Obstacle, Traffic, No-fly_zones represent weather conditions, obstacles and flight / travel space, traffic conditions, and no-fly zones.

[0032] S3: Calculate the comprehensive Q value based on the matching attention weight and the individual Q value of the agent.

[0033] Specifically, the calculation formula for the comprehensive Q value is: Calculate the comprehensive Q value based on the matching attention weight and the individual Q value of the agent ; where, Represents the comprehensive Q value corresponding to the i th agent; Represents the state at the t th moment; Represents the action selected at the t th moment; M Represents the number of agents; Represents the matching attention weight between the task and the m th agent; Represents the individual Q value of the m th agent; Represents the network parameters of the m th agent.

[0034] S4: The current network in the double deep Q network selects the corresponding action of the agent in the corresponding state based on the comprehensive Q value, and the target network in the double deep Q network calculates the evaluation value of the action; The state includes state information and task information, and the action is for the agent to execute the task (including the assigned task and the resource configuration of the agent).

[0035] In this embodiment, the training process of the double deep Q network includes: Define the state, action, and immediate reward; Create two deep Q networks with the same structure, serving as the current network and the target network respectively; The current network selects an action based on the state information and task information, and adopts the ε-greedy strategy; The ε-greedy strategy randomly selects an action with probability ε or selects the corresponding action of the agent in the state when the comprehensive Q value is the largest; Initialize the experience replay pool; The agent executes the selected action, updates the state, and calculates the immediate reward; The calculation formula of the immediate reward is: ; ; ; ; Among them, Represents the immediate reward at the t th moment; Represents the task completion reward; Represents the path optimization reward; Represents the resource utilization reward; Represents the first weight coefficient; Represents the second weight coefficient; Represents the third weight coefficient; K Represents the total number of tasks; Indicates whether the k th task is completed. If completed, the value is 1; if not completed, the value is 0; Indicates the delay tolerance of the k th task; M Indicates the number of agents; Indicates the path length of the i th agent; Indicates the energy consumption of the i th agent in path planning; Indicates the resources actually used by the i th agent; Indicates the total available resources of the i th agent; Indicates the idle time of the i th agent.

[0036] Take the state before update, the selected action, the immediate reward, and the state after update as experience and store it in the experience replay pool.

[0037] When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experience from the experience replay pool, use the current network to select the optimal action in the state after update, and use the target network to calculate the evaluation value of the optimal action.

[0038] Pass the optimal action selected by the current network in the state after update through the current network to obtain the Q value of the current network.

[0039] Based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network, calculate the mean squared error loss function; minimize the mean squared error loss function and backpropagate to update the Q value of the current network.

[0040] The expression of the mean squared error loss function is: ; ; Among them, represents the mean squared error loss function; represents the network parameters of the current network before update; B represents the batch size; Indicates the evaluation value of the optimal action calculated by the target network in the b th batch; represents the Q value of the current network before update; Indicates the state in the b th batch; Indicates the action in the b th batch; Indicates the immediate reward in the b th batch; Represents the discount factor, which is the weight of future rewards; Represents the Q-value of the target network; Represents the b’ state in the Represents the b’ optimal action in the state of the

[0041] In complex scenarios (such as sudden weather changes, UGV path blockages), the double deep Q-network reduces policy oscillations caused by single network errors by separating action decision-making and value evaluation, ensuring the continuity of task allocation. The generated matching attention weights provide structured prior knowledge (such as task-agent matching degree) for the action selection of the current network. The target network generates stable Q-value estimates based on global state information (such as no-fly zone distribution, UGV load capacity), avoiding local noise interference. The update formula for the Q-value of the current network is: ; Among them, Represents the updated Q-value of the current network; Represents the state at the t th moment; Represents the action at the t th moment; Represents the network parameters of the current network after update; Represents the Q-value of the current network before update; Represents the network parameters of the current network before update; Represents the learning rate; Represents the immediate reward at the t th moment; Represents the discount factor; Represents the Q-value of the target network; Represents the network parameters of the target network; Represents the t +1 state at the moment; Represents the Q-value of the optimal action of the current network in the state ; ;

[0042] At fixed intervals of steps, the network parameters of the target network are softly updated based on the network parameters in the Q-value of the current network. The update formula is: ; Among them, Represents the network parameters of the target network after update; Represents the adjustment parameter, ; Represents the network parameters of the current network before update; Represents the network parameters of the target network.

[0043] S5: Based on the action with the highest evaluation value in the corresponding state, allocate the corresponding task to the agent in real time.

[0044] This embodiment also provides a resource allocation system, which includes: a drone cluster, an unmanned vehicle, a central control platform, and an edge server.

[0045] The working process of the system is as follows: 1. Task reception and allocation: After receiving a task request, the central control platform performs global task allocation according to factors such as task type, urgency, and device status (UAV and UGV). Relatively simple tasks with high real-time requirements (such as data collection and video stream analysis) are allocated to the drone cluster; while computationally intensive tasks that require strong processing capabilities (such as path planning and environmental perception) are allocated to the ground unmanned vehicle (UGV) or the edge server.

[0046] 2. Data processing and computing: The drone cluster is responsible for performing preliminary data processing, such as data screening, image preprocessing, and preliminary object detection. The edge server provides computing support for the UGV at the ground node to process more complex tasks, such as real-time path planning and environmental analysis.

[0047] 3. Task execution and coordination: UAVs and UGVs execute tasks according to the schedule, dynamically adjust tasks and paths based on environmental feedback, and update their status in real time. UAVs execute tasks such as data collection and light cargo delivery, and UGVs are responsible for computationally intensive tasks and heavy cargo transportation tasks. The drones and UGVs execute specific tasks according to the instructions of the central control platform and the edge server, and at the same time, they feedback the task status in real time. The central control platform monitors the task execution situation according to the feedback information, adjusts the task allocation strategy in a timely manner, and optimizes resource utilization.

[0048] 4. Real-time monitoring and adjustment: The central control platform continuously monitors the operating status of the system from a global perspective, identifies potential problems (such as equipment failures and path obstacles), and makes adjustments in a timely manner. In case of emergencies or failures, the platform responds quickly to perform task rescheduling or activate backup solutions.

[0049] The intelligent agent dynamic resource allocation method driven by graph attention and double Q network provided in this embodiment has the following beneficial effects: 1. Support the collaborative work of heterogeneous agents: There are significant differences in resources and task capabilities between drones (UAVs) and ground unmanned vehicles (UGVs). Through the comprehensive Q value and individual Q value, this model can efficiently coordinate the task execution of heterogeneous agents, enabling them to give full play to their respective capabilities.

[0050] 2. Dynamic Adaptation to Complex Environments: Based on the learning ability of the attention mechanism and deep reinforcement learning, the agent can autonomously adjust task allocation and path planning strategies in dynamic task requirements (such as peak logistics tasks) and changing environments (such as bad weather and communication obstacles). By dynamically evaluating the matching degree between tasks and the environment through the attention mechanism, the most suitable resource nodes are preferentially allocated, thereby reducing latency and saving energy consumption.

[0051] 3. Balance between Global Optimum and Local Optimum: The comprehensive Q-value provides a global perspective. By optimizing the overall system performance (such as task completion rate and energy efficiency), global-optimal task allocation and resource scheduling are achieved. The individual Q-value focuses on the local performance of a single agent (such as path planning and computing resource utilization) to ensure the efficient operation of each agent. The two work together to take into account the actual needs and constraints of individual agents while maintaining global optimality.

[0052] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0053] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An intelligent agent dynamic resource allocation method driven by graph attention and double Q network, which is applied to an edge server, and is characterized in that Including: S1: Receive the task information sent by the central control platform and the status information uploaded by the agents in real time; the agents include a drone cluster and / or unmanned vehicles; S2: Generate a graph structure of the status information and the task information through the self-attention mechanism, and sequentially pass the graph structure through a graph convolutional network and a graph attention network to obtain the attention weights of the edges; calculate the matching attention weights between the task and the agents based on the attention weights of the edges; S3: Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the agents; S4: The current network in the double deep Q network selects the corresponding actions of the agents in the corresponding state based on the comprehensive Q value, and the target network in the double deep Q network calculates the evaluation value of the actions; the state includes the status information and the task information, and the action is for the agents to execute the task; S5: Allocate the corresponding tasks to the agents in real time based on the actions with the highest evaluation value in the corresponding state.

2. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, wherein The task information includes: the task data size, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.

3. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that The status information includes the agent status information and the environmental status information; The agent status information includes: the current 3D position coordinates of the agent, the current speed of the agent, the current remaining power of the agent, the current available computing resources of the agent, the progress of the task currently executed by the agent; The environmental status information includes: obstacles, weather, no-fly zones or the road traffic conditions for the unmanned vehicle to travel.

4. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, wherein, The generation of the graph structure of the status information and the task information through the self-attention mechanism includes: Perform feature embedding on the status information and the task information respectively to obtain the agent status feature representation and the task feature representation; Use the agent status feature representation and the task feature representation as the nodes of the graph structure, and use the connection relationship between the agent status feature representation and the task feature representation as the edges of the graph structure; The calculation of the connection relationship includes: Multiply the task feature representation by the first parameter matrix, and multiply the agent status feature representation by the second parameter matrix; Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio; Pass the first ratio through the softmax function to obtain the weights of the edges; Traverse all the task feature representations, calculate the weights of the edges between all the task feature representations and each agent status feature representation, and obtain the affinity matrix of the status information and the task information.

5. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 4, wherein In S2, the process of obtaining the matching attention weights includes: Use the k-means clustering algorithm to dynamically adjust the edges between the nodes in the graph structure; Perform normalization processing on the affinity matrix to obtain the adjacency matrix of the graph; Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure by the feature representation of any node in the graph structure and the third parameter matrix, and pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node; Traverse all the nodes in the graph structure and update the feature representations of all the nodes to obtain the embedding representation of the graph. Input the embedded representation of the graph into the graph attention network. In the graph attention network, multiply the embedded representation of the graph by the third parameter matrix, and pass the obtained second product through the LeakyReLU activation function to obtain the attention weights of the edges in the graph; Calculate the matching attention weights between the task and the agent based on the attention weights of the edges.

6. The method for dynamic resource allocation of an agent driven by graph attention and double Q network according to claim 5, wherein The specific calculation of the matching attention weights is as follows: Pass the attention weights of the edges and negative infinity through the indicator function, and pass the obtained value through the softmax function to obtain the matching attention weights between the task and the agent; The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.

7. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that The calculation formula for the comprehensive Q value is: Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the agents. ; Among them, represents the comprehensive Q-value corresponding to the i th agent; represents the state at the t th moment; represents the action selected at the t th moment; M represents the number of agents; represents the matching attention weight between the task and the m th agent; represents the individual Q-value of the m th agent; represents the network parameters of the m th agent.

8. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, wherein The training process of the double deep Q network includes: Define the state, action, and immediate reward; Create two deep Q networks with the same structure, which are used as the current network and the target network respectively; The current network selects an action based on the state information and task information, and adopts the ε-greedy strategy; the ε-greedy strategy randomly selects an action with probability ε or selects the corresponding action in the state where the agent has the maximum comprehensive Q value; Initialize the experience replay pool; The agent executes the selected action, updates the state, and calculates the immediate reward; Take the state before the update, the selected action, the immediate reward, and the state after the update as the experience, and store it in the experience replay pool; When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experiences from the experience replay pool, use the current network to select the optimal action in the updated state, and use the target network to calculate the evaluation value of the optimal action; Pass the optimal action selected by the current network in the updated state through the current network to obtain the Q value of the current network; Calculate the mean squared error loss function based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network; minimize the mean squared error loss function and backpropagate to update the Q value of the current network; Every fixed number of steps, perform a soft update on the network parameters of the target network based on the network parameters in the Q value of the current network.

9. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 8, characterized in that, The calculation formula for the immediate reward is: ; ; ; ; Among them, represents the immediate reward at the t moment; represents the task completion reward; represents the path optimization reward; represents the resource utilization reward; represents the first weight coefficient; represents the second weight coefficient; represents the third weight coefficient; K represents the total number of tasks; represents whether the k th task is completed. If completed, the value is 1; if not completed, the value is 0; represents the delay tolerance of the k th task; M represents the number of agents; represents the path length of the i th agent; represents the energy consumption of the i th agent in path planning; represents the resources actually used by the i th agent; represents the total available resources of the i th agent; represents the idle time of the i th agent.

10. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 8, characterized in that, The update formula for the Q value of the current network is: ; Among them, represents the Q value after the current network update; represents the t state at time represents the t action at time represents the network parameters after the current network update; represents the Q value before the current network update; represents the network parameters before the current network update; represents the learning rate; represents the t immediate reward at time represents the discount factor; represents the Q value of the target network; represents the network parameters of the target network; represents the t state at time represents the optimal action of the current network in the state

Citation Information

Patent Citations

  • Multi-agent migration reinforcement learning method based on graph attention network

    CN115936058A

  • Internet of vehicles multi-agent edge computing content cache decision-making method based on graph attention

    CN116634396A

  • Marine wireless network resource allocation method based on attention mechanism

    CN119136249A

  • Demonstration-conditioned reinforcement learning for few-shot imitation

    US20220395975A1

Cited By

  • Computer resource allocation management method and system based on deep learning

    CN121349707A