An Agent Dynamic Resource Allocation Method Driven by Graph Attention and Double Q-Network
Through the dynamic resource allocation method of the agent driven by graph attention and dual Q network, the problem of coordinated dispatch of drones and unmanned vehicles in low-altitude logistics networks is solved, efficient and dynamic resource allocation and task completion are achieved, delay and energy consumption are reduced, and the overall performance of the system is improved.
Patent Information
- Application Number
- CN202510696264.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-28
AI Technical Summary
There are high dynamics in low-altitude logistics networks, significant resource capabilities differences, large dynamic environment interference, traditional scheduling methods are difficult to achieve efficient coordination and real-time response, and the computational complexity is high, making it difficult to meet task requirements.
The dynamic resource allocation method of the agent driven by graph attention and dual Q network is adopted, and the graph structure of the task and the agent is generated through the self-attention mechanism, the graph convolution and graph attention network are used to calculate the attention weight, and the agent action selection and resource allocation are combined with the dual-depth Q network to achieve matching the task and the agent.
It realizes efficient coordination between drones and ground unmanned vehicles, reduces task delays and energy consumption, improves task completion rates, dynamically adapts to complex environments, and balances global and local optimizations.
Smart Images

Figure CN120218573B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of resource allocation, and particularly to an intelligent agent dynamic resource allocation method driven by graph attention and double Q network. Background Art
[0002] The low-altitude logistics network is a new type of logistics transportation system constructed based on unmanned equipment such as unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs). Utilizing low-altitude airspace resources, it realizes efficient delivery of goods, information, and services in areas such as towns, rural areas, and industrial parks. There are many deficiencies in traditional scheduling methods in the low-altitude logistics network, such as:
[0003] 1. The personnel requirements in the low-altitude logistics network are highly dynamic (such as emergency deliveries, sudden tasks, etc.), and it is difficult for traditional static scheduling methods to respond.
[0004] 2. There are significant differences in resource capabilities (computing, load capacity, endurance) between UAVs (lightweight tasks, fast response) and UGVs (computing-intensive, heavy transportation), and it is difficult for traditional scheduling methods to coordinate their complementarity.
[0005] 3. The low-altitude logistics network faces dynamic environmental interferences (such as weather changes, obstacles, no-fly zones, traffic congestion), and traditional scheduling methods cannot effectively model these uncertainties.
[0006] 4. Traditional scheduling methods often struggle to balance global optimization (such as total system energy consumption) and local efficiency.
[0007] 5. When the task scale expands, the computational complexity of traditional optimization algorithms (such as linear programming) increases significantly, making it difficult to meet the requirements of real-time scheduling. Summary of the Invention
[0008] Based on this, it is necessary to provide an intelligent agent dynamic resource allocation method driven by graph attention and double Q network. This method is applied to an edge server and includes:
[0009] S1: Receive the task information sent by the central control platform and the status information uploaded by the intelligent agent in real time; the intelligent agent includes a UAV cluster and / or a UGV;
[0010] S2: Generate a graph structure of the status information and the task information through a self-attention mechanism, and sequentially pass the graph structure through a graph convolutional network and a graph attention network to obtain the attention weights of the edges; calculate the matching attention weights between the task and the intelligent agent based on the attention weights of the edges;
[0011] S3: Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the intelligent agent;
[0012] S4: The current network in the double deep Q-network selects the corresponding action of the agent in the corresponding state based on the comprehensive Q-value, and the target network in the double deep Q-network calculates the evaluation value of the action; the state includes state information and task information, and the action is the agent executing the task.
[0013] S5: Based on the action with the highest evaluation value in the corresponding state, allocate the corresponding task to the agent in real time.
[0014] Preferably, the task information includes: the size of the task data, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.
[0015] Preferably, the state information includes agent state information and environmental state information;
[0016] The agent state information includes: the current 3D position coordinates of the agent, the current speed of the agent, the current remaining battery power of the agent, the current available computing resources of the agent, the progress of the task currently executed by the agent;
[0017] The environmental state information includes: obstacles, weather, no-fly zones or the road traffic conditions for the unmanned vehicle to travel.
[0018] Preferably, the graph structure for generating state information and task information through the self-attention mechanism includes:
[0019] Perform feature embedding on the state information and task information respectively to obtain the agent state feature representation and the task feature representation;
[0020] Use the agent state feature representation and the task feature representation as the nodes of the graph structure, and use the connection relationship between the agent state feature representation and the task feature representation as the edges of the graph structure;
[0021] The connection relationship calculation includes:
[0022] Multiply the task feature representation by the first parameter matrix, and multiply the agent state feature representation by the second parameter matrix;
[0023] Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio;
[0024] Pass the first ratio through the softmax function to obtain the weight of the edge;
[0025] Traverse all the task feature representations, calculate the weights of the edges between all the task feature representations and each agent state feature representation, and obtain the affinity matrix of the state information and the task information.
[0026] Preferably, in S2, the process of obtaining the matching attention weight includes:
[0027] Dynamically adjust the edges between nodes in the graph structure using the k-means clustering algorithm;
[0028] Normalize the affinity matrix to obtain the adjacency matrix of the graph;
[0029] Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure with the feature representation of any node in the graph structure and the third parameter matrix, and pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node;
[0030] Traverse all nodes in the graph structure to update the feature representations of all nodes and obtain the embedding representation of the graph;
[0031] Input the embedding representation of the graph into the graph attention network. In the graph attention network, multiply the embedding representation of the graph with the third parameter matrix, and pass the obtained second product through the LeakyReLU activation function to obtain the attention weights of the edges in the graph;
[0032] Calculate the matching attention weights between tasks and agents based on the attention weights of the edges.
[0033] Preferably, the specific calculation of the matching attention weights is as follows:
[0034] Pass the attention weights of the edges and negative infinity through the indicator function, and pass the obtained value through the softmax function to obtain the matching attention weights between tasks and agents;
[0035] The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.
[0036] Preferably, the calculation formula for the comprehensive Q value is:
[0037] Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the agents
[0038] ;
[0039] where, represents the comprehensive Q value corresponding to the i -th agent; represents the state at the t -th moment; represents the action selected at the t -th moment; M represents the number of agents; represents the matching attention weight between the task and the m -th agent; represents the m -th agent's individual Q value; represent the network parameters of the m th agent.
[0040] Preferably, the training process of the double deep Q network includes:
[0041] Define the state, action, and immediate reward;
[0042] Create two deep Q networks with the same structure, serving as the current network and the target network respectively;
[0043] The current network selects an action based on the state information and task information, and adopts the ε-greedy strategy; the ε-greedy strategy randomly selects an action with probability ε or selects the corresponding action in the state where the agent has the maximum comprehensive Q value;
[0044] Initialize the experience replay pool;
[0045] The agent executes the selected action, updates the state, and calculates the immediate reward;
[0046] Take the state before update, the selected action, the immediate reward, and the state after update as experience, and store it in the experience replay pool;
[0047] When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experience from the experience replay pool, use the current network to select the optimal action in the updated state, and use the target network to calculate the evaluation value of the optimal action;
[0048] Pass the optimal action selected by the current network in the updated state through the current network to obtain the Q value of the current network;
[0049] Based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network, calculate the mean squared error loss function; minimize the mean squared error loss function and backpropagate to update the Q value of the current network;
[0050] Every fixed number of steps, perform a soft update on the network parameters of the target network based on the network parameters in the Q value of the current network.
[0051] Preferably, the calculation formula for the immediate reward is:
[0052] ;
[0053] ;
[0054] ;
[0055] ;
[0056] where represent thet Immediate reward at a moment Indicates the task completion reward Indicates the path optimization reward Indicates the resource utilization reward Indicates the first weight coefficient Indicates the second weight coefficient Indicates the third weight coefficient K Indicates the total number of tasks Indicates the k completion status of the Indicates the k delay tolerance of the M Indicates the number of agents Indicates the i path length of the Indicates the i energy consumption of the Indicates the i actual resources used by the Indicates the i total available resources of the Indicates the i idle time of the
[0057] Preferably, the update formula for the Q value of the current network is:
[0058] ;
[0059] Wherein, Indicates the Q value after the update of the current network Indicates the t status at a moment Indicates the t action at a moment Indicates the network parameters after the update of the current network Indicates the Q value before the update of the current network Indicates the network parameters before the update of the current network Indicates the learning rate Indicates the t immediate reward at a moment Indicates the discount factor Indicates the Q value of the target network Indicates the network parameters of the target network Indicates the t status at the Indicates the Q value of the optimal action of the current network in the state .
[0060] Beneficial effects: The method can achieve dynamic optimization through intelligent decision-making, aiming to achieve efficient cooperation between drones and ground unmanned vehicles in a dynamic task scenario and improve the task completion rate; the method reduces task latency, energy consumption, and costs. Description of the Drawings
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0062] Figure 1 It is a flowchart of the intelligent agent dynamic resource allocation method driven by graph attention and double Q network in the embodiment of the present application. Detailed Embodiments
[0063] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe the detailed embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0064] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0065] As Figure 1 shown, this embodiment provides an intelligent agent dynamic resource allocation method driven by graph attention and double Q network. This method is applied to an edge server and includes:
[0066] S1: Receive the task information sent by the central control platform and the status information uploaded by the intelligent agent in real time; the intelligent agent includes a drone cluster and / or an unmanned vehicle.
[0067] In this embodiment, the task information includes: the size of the task data, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.
[0068] The status information includes intelligent agent status information and environmental status information;
[0069] The agent state information includes: the current 3D position coordinates of the agent, the current speed of the agent, the current remaining power of the agent, the current available computing resources of the agent, and the progress of the task currently being executed by the agent;
[0070] The environmental state information includes: obstacles, weather, no-fly zones, or the road traffic conditions for the autonomous vehicle to travel.
[0071] S2: Generate a graph structure of the state information and task information through the self-attention mechanism, and sequentially pass the graph structure through a graph convolutional network and a graph attention network to obtain the attention weights of the edges; calculate the matching attention weights between the task and the agent based on the attention weights of the edges.
[0072] Specifically, the generation of the graph structure of the state information and task information through the self-attention mechanism includes:
[0073] Perform feature embedding on the state information and task information respectively to obtain the agent state feature representation and the task feature representation;
[0074] Use the agent state feature representation and the task feature representation as the nodes of the graph structure, and use the connection relationship between the agent state feature representation and the task feature representation as the edges of the graph structure;
[0075] The connection relationship calculation includes:
[0076] Multiply the task feature representation by the first parameter matrix, and multiply the agent state feature representation by the second parameter matrix;
[0077] Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio;
[0078] Pass the first ratio through the softmax function to obtain the weights of the edges;
[0079] Traverse all the task feature representations, calculate the weights of the edges between all the task feature representations and each agent state feature representation, and obtain the affinity matrix of the state information and task information.
[0080] Generate the affinity matrix of the graph through the self-attention mechanism, and use the graph structure to represent the relationship between the task and the agent. These graph structures provide a basis for the subsequent calculation of the matching attention weights.
[0081] Furthermore, the process of obtaining the matching attention weights includes:
[0082] Use the k-means clustering algorithm to dynamically adjust the edges between the nodes in the graph structure to enhance the connection relationship between the nodes;
[0083] Perform normalization processing on the affinity matrix to obtain the adjacency matrix of the graph;
[0084] Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure with the feature representation of any node in the graph structure and the third parameter matrix, and pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node.
[0085] Traverse all nodes in the graph structure to update the feature representations of all nodes and obtain the embedding representation of the graph.
[0086] Input the embedding representation of the graph into the graph attention network. In the graph attention network, multiply the embedding representation of the graph with the third parameter matrix. This is a linear transformation used to map the embedding representation to a new space to better capture the relationships between nodes. Pass the obtained second product through the LeakyReLU activation function. The LeakyReLU activation function can introduce non-linearity to help the model capture the complex relationships between nodes. Compared with the traditional ReLU activation function, the LeakyReLU activation function also has an output in the negative value region, avoiding the problem of gradient disappearance. Calculate the attention weights of each edge in the graph through the above linear transformation and LeakyReLU activation function. The attention weights of the edges reflect the association strength between nodes. The higher the value, the more important the relationship between nodes.
[0087] Calculate the matching attention weights between tasks and agents based on the attention weights of the edges.
[0088] Furthermore, the specific calculation of the matching attention weights is as follows:
[0089] Pass the attention weights of the edges and negative infinity through the indicator function, and then pass the obtained value through the softmax function to obtain the matching attention weights between tasks and agents.
[0090] The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.
[0091] Through the fusion of information, the most important features in the graph embedding can be focused on, the matching attention weights between tasks and agents can be evaluated, which provides a basis for the subsequent calculation of the comprehensive Q value and further improves the effect of task allocation.
[0092] In task scheduling, the goal is to assign tasks to the most suitable agents (UAVs or UGVs). Each task has specific computing requirements, latency requirements, and data requirements. The selection of agents needs to consider the matching of these requirements and resources. In this embodiment, the matching problem can be modeled as:
[0093] ;
[0094] ;
[0095] ;
[0096] Among them, represents the matching degree between the task and the resource (such as the matching degree between computing resources and battery power, etc.), , 1 represents a perfect match, and 0 represents a complete mismatch; represents the matching degree between the task and the environment (such as the matching degree between obstacle avoidance and weather impact, etc.), , 1 represents a perfect match, and 0 represents a complete mismatch; represents the k-th task; represents the resource information; represents the environmental information; represents the weight coefficient between the task and the resource; represents the weight coefficient between the task and the environment; , , , are the coefficients of each resource factor respectively; Computation, Storage, Energy, Communication represent computing power, storage capacity, required energy consumption, and communication ability respectively; , , , are the coefficients of each environmental factor respectively; Weather, Obstacle, Traffic, No-fly_zones represent weather conditions, obstacles and flight / travel space, traffic conditions, and no-fly zones respectively.
[0097] S3: Calculate the comprehensive Q value based on the matching attention weight and the individual Q value of the agent.
[0098] Specifically, the calculation formula for the comprehensive Q value is:
[0099] Calculate the comprehensive Q value based on the matching attention weight and the individual Q value of the agent
[0100] ;
[0101] Among them, represents the comprehensive Q value corresponding to the i -th agent; represents the state at the t -th moment; represents the action selected at the t -th moment; M represents the number of agents; represents the task and the mThe matching attention weights among agents; Denote the m individual Q value of the th agent; m Denote the network parameters of the
[0102] S4: The current network in the double deep Q network selects the corresponding action of the agent in the corresponding state based on the comprehensive Q value, and the target network in the double deep Q network calculates the evaluation value of the action; The state includes state information and task information, and the action is for the agent to execute tasks (including the assigned tasks and the resource allocation of the agent).
[0103] In this embodiment, the training process of the double deep Q network includes:
[0104] Define the state, action, and immediate reward;
[0105] Create two deep Q networks with the same structure, serving as the current network and the target network respectively;
[0106] The current network selects an action based on the state information and task information, and adopts the ε-greedy strategy; The ε-greedy strategy randomly selects an action with probability ε or selects the corresponding action of the agent in the state when the comprehensive Q value is the largest;
[0107] Initialize the experience replay pool;
[0108] The agent executes the selected action, updates the state, and calculates the immediate reward;
[0109] The calculation formula of the immediate reward is:
[0110] ;
[0111] ;
[0112] ;
[0113] ;
[0114] Among them, Denote the t immediate reward at the th moment; Denote the task completion reward; Denote the path optimization reward; Denote the first weight coefficient; Denote the second weight coefficient; Denote the third weight coefficient; K Denote the total number of tasks; Denote the kWhether a task is completed, if it is completed, the value is 1, if it is not completed, the value is 0; Indicates the k delay tolerance of the M th task; Indicates the number of agents; i Indicates the path length of the th agent; i Indicates the energy consumption of the th agent in path planning; i Indicates the resources actually used by the th agent; i Indicates the total resources available to the th agent; i Indicates the idle time of the
[0115] Take the state before update, the selected action, the immediate reward, and the state after update as experience, and store it in the experience replay pool.
[0116] When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experience from the experience replay pool, use the current network to select the optimal action in the state after update, and use the target network to calculate the evaluation value of the optimal action.
[0117] Pass the optimal action selected by the current network in the state after update through the current network to obtain the Q value of the current network.
[0118] Based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network, calculate the mean squared error loss function; minimize the mean squared error loss function and backpropagate to update the Q value of the current network.
[0119] The expression of the mean squared error loss function is:
[0120] ;
[0121] ;
[0122] Among them, Indicates the mean squared error loss function; Indicates the network parameters of the current network before update; B represents the batch size; Indicates the b th evaluation value of the optimal action calculated by the target network in the batch; Indicates the Q value of the current network before update; Indicates the b th state in the batch; Indicates the b th action in the batch; Indicates the bImmediate rewards in a batch; Denotes the discount factor, representing the weight of future rewards; Denotes the Q-value of the target network; Denotes the b’ State in the batch; Denotes the b’ Optimal action in the state of the batch.
[0123] In complex scenarios (such as sudden weather changes, UGV path blockages), the double deep Q-network reduces policy oscillation caused by single network errors by separating action decision-making and value evaluation, ensuring the continuity of task allocation. The generated matching attention weights provide structured prior knowledge (such as task-agent matching degree) for the action selection of the current network. The target network generates stable Q-value estimates based on global state information (such as no-fly zone distribution, UGV load capacity), avoiding local noise interference. The update formula for the Q-value of the current network is:
[0124] ;
[0125] Where, Denotes the updated Q-value of the current network; Denotes the t State at time Denotes the t Action at time Denotes the updated network parameters of the current network; Denotes the Q-value of the current network before update; Denotes the network parameters of the current network before update; Denotes the learning rate; Denotes the t Immediate reward at time Denotes the discount factor; Denotes the Q-value of the target network; Denotes the network parameters of the target network; Denotes the t State at time +1; Denotes the Q-value of the optimal action of the current network in the state .
[0126] At fixed intervals of steps, the network parameters of the target network are softly updated based on the network parameters in the Q-value of the current network. The update formula is:
[0127] ;
[0128] Where, Denotes the updated network parameters of the target network; Denotes the adjustment parameter, ; represents the network parameters before the current network update; represents the network parameters of the target network.
[0129] S5: Based on the action with the highest evaluation value in the corresponding state, allocate the corresponding task to the agent in real time.
[0130] This embodiment also provides a resource allocation system, which includes: a drone cluster, an unmanned vehicle, a central control platform, and an edge server.
[0131] The working process of the system is as follows:
[0132] 1. Task reception and allocation: After receiving a task request, the central control platform performs global task allocation according to factors such as task type, urgency, and device status (UAV and UGV). Relatively simple tasks with high real-time requirements (such as data collection and video stream analysis) are allocated to the drone cluster; while computationally intensive tasks that require strong processing capabilities (such as path planning and environmental perception) are allocated to the ground unmanned vehicle (UGV) or the edge server.
[0133] 2. Data processing and computing: The drone cluster is responsible for performing preliminary data processing, such as data screening, image preprocessing, and preliminary target detection. The edge server provides computing support for the UGV at the ground node and processes more complex tasks, such as real-time path planning and environmental analysis.
[0134] 3. Task execution and coordination: The UAV and UGV execute tasks according to the schedule, dynamically adjust tasks and paths according to environmental feedback, and update the status in real time. The UAV executes tasks such as data collection and light cargo delivery, and the UGV is responsible for computationally intensive tasks and heavy cargo transportation tasks. The drones and UGV execute specific tasks according to the instructions of the central control platform and the edge server, and at the same time, feedback the task status in real time. The central control platform monitors the task execution situation according to the feedback information, adjusts the task allocation strategy in a timely manner, and optimizes resource utilization.
[0135] 4. Real-time monitoring and adjustment: The central control platform continuously monitors the operating status of the system from a global perspective, identifies potential problems (such as equipment failures and path obstacles), and makes adjustments in a timely manner. When an emergency event or failure occurs, the platform quickly responds, performs task rescheduling, or activates a backup plan.
[0136] The intelligent agent dynamic resource allocation method driven by graph attention and double Q network provided in this embodiment has the following beneficial effects:
[0137] 1. Support collaborative work of heterogeneous agents: There are significant differences in resources and task capabilities between unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs). This model can efficiently coordinate the task execution of heterogeneous agents by integrating the comprehensive Q-value and the individual Q-value, enabling them to fully utilize their respective capabilities.
[0138] 2. Dynamically adapt to complex environments: Based on the learning ability of the attention mechanism and deep reinforcement learning, agents can autonomously adjust task allocation and path planning strategies in dynamic task requirements (such as peak logistics tasks) and changing environments (such as bad weather and communication obstacles). By dynamically evaluating the task-environment matching degree through the attention mechanism, the most suitable resource nodes are preferentially allocated, thereby reducing latency and saving energy consumption.
[0139] 3. Balance between global optimality and local optimality: The comprehensive Q-value provides a global perspective and realizes the globally optimal task allocation and resource scheduling by optimizing the overall system performance (such as task completion rate and energy efficiency). The individual Q-value focuses on the local performance of a single agent (such as path planning and computing resource utilization) to ensure the efficient operation of each agent. The two work together to take into account the actual needs and constraints of individual agents while maintaining global optimality.
[0140] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0141] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An intelligent agent dynamic resource allocation method driven by graph attention and double Q network, which is applied to an edge server, and is characterized in that Including: S1: Receive the task information sent by the central control platform and the status information uploaded by the agents in real time; the agents include a drone swarm and / or unmanned vehicles; S2: Generate a graph structure of the status information and the task information through the self-attention mechanism, and sequentially pass the graph structure through a graph convolutional network and a graph attention network to obtain the attention weights of the edges; Calculate the matching attention weights between the tasks and the agents based on the attention weights of the edges; The generating of the graph structure of the status information and the task information through the self-attention mechanism includes: Perform feature embedding on the status information and the task information respectively to obtain the agent status feature representation and the task feature representation; Take the agent status feature representation and the task feature representation as the nodes of the graph structure, and take the connection relationship between the agent status feature representation and the task feature representation as the edges of the graph structure; The connection relationship calculation includes: Multiply the task feature representation by the first parameter matrix, and multiply the agent status feature representation by the second parameter matrix; Multiply the two products and divide by the square root of the feature dimension to obtain the first ratio; Pass the first ratio through the softmax function to obtain the weights of the edges; Traverse all the task feature representations, calculate the weights of the edges between all the task feature representations and each agent status feature representation, and obtain the affinity matrix of the status information and the task information; The process of obtaining the matching attention weights includes: Use the k-means clustering algorithm to dynamically adjust the edges between the nodes in the graph structure; Perform normalization processing on the affinity matrix to obtain the adjacency matrix of the graph; Input the adjusted graph structure into the graph convolutional network. In the graph convolutional network, multiply the adjacency matrix of the graph structure by the feature representation of any node in the graph structure and the third parameter matrix, and pass the obtained first product through the ReLU activation function to update the feature representation of the corresponding node; Traverse all the nodes in the graph structure, update the feature representations of all the nodes, and obtain the embedding representation of the graph; Input the embedding representation of the graph into the graph attention network. In the graph attention network, multiply the embedding representation of the graph by the third parameter matrix, and pass the obtained second product through the LeakyReLU activation function to obtain the attention weights of the edges in the graph; Calculate the matching attention weights between the tasks and the agents based on the attention weights of the edges; S3: Calculate the comprehensive Q value based on the matching attention weights and the individual Q values of the agents; S4: The current network in the double deep Q network selects the corresponding actions of the agents in the corresponding states based on the comprehensive Q value, and the target network in the double deep Q network calculates the evaluation value of the actions; the states include the status information and the task information, and the actions are for the agents to execute tasks; S5: Based on the actions with the highest evaluation value in the corresponding states, allocate the corresponding tasks to the agents in real time.
2. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, wherein The task information includes: the task data size, the computing requirements of the task, the latency requirements of the task, the target location of the task, the priority or urgency of the task.
3. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that, The status information includes the agent status information and the environmental status information; The agent state information includes: the current 3D position coordinates of the agent, the current speed of the agent, the current remaining battery power of the agent, the current available computing resources of the agent, and the progress of the task currently executed by the agent; The environmental state information includes: obstacles, weather, no-fly zones, or the road traffic conditions for the unmanned vehicle to travel.
4. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that, The specific calculation of the matching attention weight is as follows: Pass the attention weight of the edge and negative infinity through the indicator function, and pass the obtained value through the softmax function to obtain the matching attention weight between the task and the agent; The indicator function includes: when there is an edge connection between nodes, the value is the attention weight of the edge; when there is no edge connection between nodes, the value is 0.
5. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that, The calculation formula for the comprehensive Q value is: Based on the matching attention weight and the individual Q value of the agent, calculate the comprehensive Q value ; Among them, represents the comprehensive Q value corresponding to the i th agent; represents the state at the t th moment; represents the action selected at the t th moment; M represents the number of agents; represents the matching attention weight between the task and the m th agent; represents the individual Q value of the m th agent; represents the network parameters of the m th agent.
6. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 1, characterized in that The training process of the double deep Q network includes: Define the state, action, and immediate reward; Create two deep Q networks with the same structure, which are used as the current network and the target network respectively; The current network is based on the state information and task information, and selects actions using the ε-greedy strategy; the ε-greedy strategy randomly selects actions with probability ε or selects the corresponding actions in the state where the agent has the maximum comprehensive Q value; Initialize the experience replay pool; The agent executes the selected action, updates the state, and calculates the immediate reward; Take the state before the update, the selected action, the immediate reward, and the state after the update as experience, and store it in the experience replay pool; When the amount of experience in the experience replay pool meets the preset requirements, randomly sample a batch of experience from the experience replay pool, use the current network to select the optimal action in the updated state, and use the target network to calculate the evaluation value of the optimal action; Pass the optimal action selected by the current network in the updated state through the current network to obtain the Q value of the current network; Based on the Q value of the current network and the evaluation value of the optimal action calculated by the target network, calculate the mean squared error loss function; minimize the mean squared error loss function and backpropagate to update the Q value of the current network; Every fixed number of steps, perform a soft update on the network parameters of the target network based on the network parameters in the Q value of the current network.
7. The method for dynamic resource allocation of an agent driven by graph attention and double Q network according to claim 6, wherein The calculation formula for the immediate reward is: ; ; ; ; Among them, represents the immediate reward at the t moment; represents the task completion reward; represents the path optimization reward; represents the resource utilization reward; represents the first weight coefficient; represents the second weight coefficient; represents the third weight coefficient; K represents the total number of tasks; represents whether the k th task is completed. If completed, the value is 1; if not completed, the value is 0; represents the delay tolerance of the k th task; M represents the number of agents; represents the path length of the i th agent; represents the energy consumption of the i th agent in path planning; represents the resources actually used by the i th agent; represents the total available resources of the i th agent; represents the idle time of the i th agent.
8. The intelligent agent dynamic resource allocation method driven by graph attention and double Q network according to claim 6, characterized in that, The update formula for the Q value of the current network is: ; Among them, represents the Q value after the current network update; represents the t state at time represents the t action at time represents the network parameters after the current network update; represents the Q value before the current network update; represents the network parameters before the current network update; represents the learning rate; represents the t immediate reward at time represents the discount factor; represents the Q value of the target network; represents the network parameters of the target network; represents the t state at time represents the optimal action in the state of the current network and its Q value.
Citation Information
Patent Citations
Internet of vehicles multi-agent edge computing content cache decision-making method based on graph attention
CN116634396A
Marine wireless network resource allocation method based on attention mechanism
CN119136249A