Calculation unloading and resource allocation method and system based on unmanned aerial vehicle trajectory optimization
By adopting computational offloading and resource allocation methods based on drone trajectory optimization in the industrial Internet of Things, the problems of poor real-time performance and limited data processing capabilities in traditional equipment monitoring and maintenance methods are solved, and efficient and intelligent equipment monitoring and predictive maintenance are achieved.
Patent Information
- Application Number
- CN202510661849.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the industrial Internet of Things, traditional equipment monitoring and maintenance methods have problems such as poor real-time performance, limited coverage and limited data processing capabilities, which are difficult to meet the needs of modern industrial production for efficient and intelligent maintenance.
The calculation offloading and resource allocation methods based on drone trajectory optimization are adopted to improve the transmission efficiency, real-time processing capabilities and prediction and maintenance of industrial equipment data by intelligently optimizing the cruise trajectory of the drone and dynamically adjusting the bandwidth resource allocation and task processing strategies.
Effectively reduce the average information age of the system, ensure the real-time and accuracy of equipment status monitoring, meet the requirements of the industrial Internet of Things for low latency and high reliability, and optimize the bandwidth allocation and flight trajectory of the drone, and improve the execution efficiency of predicted and maintenance tasks.
Smart Images

Figure CN120186683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of industrial Internet of Things and air-ground integrated network, and particularly relates to a method and system for computing offloading and resource allocation for optimizing the trajectory of an unmanned aerial vehicle in the scenario of monitoring and maintaining Internet of Things devices. Background Art
[0002] With the rapid development of the fifth-generation mobile communication technology and the popularization of intelligent terminals, the industrial Internet of Things (IIoT) has been increasingly widely applied in fields such as intelligent manufacturing, smart grid, and intelligent transportation. Remote monitoring and predictive maintenance of industrial equipment have become important means to improve production efficiency, reduce maintenance costs, and reduce equipment failures. Traditional methods of equipment monitoring and maintenance rely on manual inspections or data collection based on fixed sensors, which have problems such as poor real-time performance, limited coverage, and limited data processing capabilities, and are difficult to meet the requirements of modern industrial production for efficient and intelligent maintenance.
[0003] Traditional mobile computing faces challenges such as limited computing power, insufficient storage space, and high energy consumption. To improve the intelligent level of equipment monitoring, remote monitoring systems based on cloud computing have been widely used. Mobile Cloud Computing (MCC) provides an efficient solution for reconciling the contradiction between delay-sensitive, computationally intensive applications and the limited resources of terminals with its excellent computing performance and flexible resource allocation capabilities. MCC migrates computing tasks from the terminal to the cloud, organically combining the powerful computing and storage capabilities of the cloud with the portability and flexibility of mobile terminals. However, the MCC architecture highly depends on network connections, and data needs to be transmitted through multiple layers of routing to the core network. This long-distance transmission will generate a relatively high delay, and it is difficult to meet application scenarios with high real-time requirements, such as fault prediction and anomaly detection of key equipment. In addition, storing a large amount of industrial data in the cloud also faces security risks such as data leakage and privacy infringement.
[0004] To overcome the limitations of MCC, Multi-access Edge Computing (MEC) has emerged as the times require. MEC effectively reduces the delay of data transmission and task processing by sinking cloud computing capabilities to the network edge, making it closer to users and data sources, and has become an important supplement to MCC. In the MEC architecture, users can selectively offload tasks to a powerful MEC server for processing. In this process, the offloaded computing tasks do not need to access the core network. Therefore, the delay of data transmission and task processing can be effectively reduced, and the computing and communication loads of the core network can be reduced.
[0005] However, traditional MEC servers are usually fixedly deployed at base stations or edge gateways, making it difficult to cover a wide range of industrial devices, especially in scenarios such as remote factories, oil fields, and wind farms. Therefore, introducing an air-ground integrated network (AGIN) and unmanned aerial vehicles (UAVs) as mobile edge computing nodes can significantly improve the flexibility and coverage of device monitoring and maintenance. Under the AGIN architecture, UAVs can serve as edge computing nodes, carrying computing and communication devices to perform inspection tasks in industrial environments and collect and process device status data in real time. Terminal devices or sensors can choose to process data locally or offload computing tasks to UAVs, which perform real-time data analysis and anomaly detection and transmit the processing results to the cloud data center to achieve collaborative optimization of remote monitoring and predictive maintenance. As the core computing platform, the cloud can receive complex computing tasks from UAVs and use its powerful computing capabilities for in-depth analysis to support UAVs and edge devices.
[0006] Computing offloading and resource allocation are two important research topics in MEC systems. Computing offloading mainly focuses on optimizing the distribution of tasks, that is, how to reasonably allocate computing tasks between industrial sensors and edge computing nodes to improve the real-time performance of device status monitoring, reduce computing and communication delays, and optimize energy consumption. Resource allocation focuses on computing resources, communication resources, storage resources, etc. in the system, and studies how to improve the overall efficiency of the system by dynamically adjusting the allocation strategies of these resources.
[0007] In the industrial Internet of Things MEC system, leveraging the mobility advantage of UAVs can improve the overall system performance and meet the needs of dynamic mobile users for low-latency and high-reliability services. However, in the UAV-assisted cloud-edge-end computing offloading architecture, the mobility of UAVs poses new challenges to computing offloading and resource allocation. The trajectory of UAVs directly affects the communication quality between them and terminal and cloud centers, thereby affecting the efficiency of task offloading and the rationality of resource allocation. In the UAV-assisted cloud-edge-end computing offloading architecture, traditional static computing offloading and resource allocation algorithms are difficult to adapt to the dynamically changing network environment of UAVs, which may lead to problems such as increased task delays, increased energy consumption, or low resource utilization. Therefore, it is crucial to design computing offloading and resource allocation algorithms based on UAV trajectory optimization. Summary of the Invention
[0008] Object of the Invention: The object of the present invention is to propose a method and system for computing offloading and resource allocation based on UAV trajectory optimization. By intelligently optimizing the cruise trajectory of UAVs, dynamically adjusting the bandwidth resource allocation and task processing strategies, the transmission efficiency, real-time processing ability of industrial device data, and the accuracy of predictive maintenance are improved, thereby ensuring the safe operation of key devices, reducing maintenance costs, and increasing production efficiency.
[0009] Technical solution: To achieve the above-mentioned invention objective, the present invention adopts the following technical solution: In a first aspect, the present invention provides a computing offloading and resource allocation method based on UAV trajectory optimization, including the following steps: Based on a multi-UAV-assisted cloud-edge-end computing offloading architecture, under the constraints of computing and communication resources, flight speed, and observation range, an optimization problem model with the motion decision of UAVs, the decision of the task offloading ratio of user terminals, the decision of computing and offloading to the cloud by UAV edge servers, and the decision of bandwidth ratio allocation by the cloud center for all UAVs as decision variables and the minimization of the average age of information of all user terminals as the optimization objective; Convert the optimization problem model into a Markov decision process model, construct three types of heterogeneous agents for user terminals, UAVs, and the cloud center, and based on the heterogeneous multi-agent advantage strategy-value algorithm, and adopt a federated update mechanism among the same type of agents for offline distributed training and online execution.
[0010] Further, the age of information of the user terminal is the current time minus the generation time of the latest completed task of the user terminal; the motion decision of the UAV is represented by the abscissa and ordinate in the horizontal direction of the UAV and the corresponding displacement changes; the decision of the task offloading ratio of the user terminal is represented by the ratio of the amount of tasks executed locally to the amount of tasks in the user data buffer; the decisions of computing and offloading to the cloud by the UAV edge server are respectively represented by a Boolean variable, including executing in the UAV side of the task queue, offloading to the cloud center, and not yet executed and not yet offloaded; the decision of the bandwidth ratio allocation by the cloud center is represented by a bandwidth allocation vector.
[0011] Further, the global state space of the Markov decision process model is the sum of the observation spaces of all agents, including the location information of user terminals, data cache space information, the location information of UAVs, data queue information, and the bandwidth allocation information of the cloud center.
[0012] Further, the observation space of each user terminal agent includes the location information and data cache space information of the user terminal, and the action space is the decision variable of the task offloading ratio; the observation space of the UAV agent includes the location information and data queue information of the UAV, as well as the location information of user terminals within its coverage range and the bandwidth allocation information of the cloud center for it, and the action space is the motion decision variable, the decision variable of computing and offloading to the cloud by the UAV edge server; the observation space of the cloud center agent includes the location information and data queue information of all UAVs and the bandwidth allocation information, and the action space is the decision variable of the bandwidth ratio allocation.
[0013] Further, in the heterogeneous multi-agent advantage strategy-value algorithm, the policy network of the user terminal agent takes the user terminal data cache space information and location information as inputs and outputs local computing decisions, and the value network takes the data cache space information and computing decisions as inputs to obtain reward values; the policy network of the cloud center agent takes the state vectors of all drones as inputs and outputs bandwidth allocation policies, and the value network takes the state vectors of all drones and bandwidth allocation policies as inputs to obtain reward values.
[0014] Further, in the heterogeneous multi-agent advantage strategy-value algorithm, the policy network of the drone agent processes the information matrix of the terminal through a convolutional neural network to obtain the motion decision of the drone, and obtains task execution and task offloading decisions based on the data queue information and bandwidth allocation information through a multi-layer perceptron network; wherein the information matrix of the terminal is a three-dimensional matrix, the first two dimensions represent the grid of the observable range centered on the drone, and the third dimension information includes the buffer data size and age of information of the user terminals within the observable range. The value network of the drone agent processes the drone state map through a convolutional neural network, extracts the drone position and its task buffer state information, and combines the bandwidth allocation information to input into a multi-layer perceptron network, and merges the outputs of the convolutional neural network and the multi-layer perceptron network and then obtains the reward value through a fully connected layer.
[0015] Further, the offline distributed training of the heterogeneous multi-agent adopts a federated update mechanism, specifically: at fixed time intervals, the same type of federated agents exchange their neural network parameters; after exchanging the parameters, each federated agent retains its own network parameters according to the set weights and performs weighted mixing with the parameters of other agents to form a global model.
[0016] In a second aspect, the present invention provides a computing offloading and resource allocation system based on drone trajectory optimization, including: A problem modeling module, which is used to, based on a multi-drone-assisted cloud-edge-end computing offloading architecture, under the constraints of computing and communication resources, flight speed, and observable range, take the drone motion decision, the user terminal task offloading ratio decision, the computing and offloading to the cloud decision of the drone edge server, and the bandwidth ratio allocation decision of the cloud center for all drones as decision variables, and minimize the average age of information of all user terminals as an optimization objective to optimize the problem model. A deep reinforcement learning module, which is used to transform the optimization problem model into a Markov decision process model, construct three types of heterogeneous agents for the user terminal, the drone, and the cloud center, and perform offline distributed training and online execution based on the heterogeneous multi-agent advantage strategy-value algorithm and adopting a federated update mechanism among the same type of agents.
[0017] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method for computing offloading and resource allocation based on UAV trajectory optimization are implemented.
[0018] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the method for computing offloading and resource allocation based on UAV trajectory optimization are implemented.
[0019] Beneficial effects: The present invention proposes a method for computing offloading and resource allocation based on UAV trajectory optimization, which is applicable to device monitoring and predictive maintenance scenarios in industrial Internet of Things. For a multi-UAV-assisted edge computing architecture, this method uses UAVs as mobile edge computing nodes to process data collected by multiple user terminals (sensors or device controllers), and under constraints such as computing and communication resources, flight speed, and observation range, it realizes efficient data offloading and resource optimization configuration. This method adopts a heterogeneous multi-agent advantage strategy-value algorithm, enabling UAVs to adaptively optimize computing task allocation and resource scheduling in complex dynamic environments, and through joint federated updates, enhancing the collaborative ability of similar industrial devices and supporting distributed online decision-making. The algorithm designed by the present invention can effectively reduce the average age of information of the system, ensure the real-time and accuracy of device status monitoring, and meet the requirements of industrial Internet of Things for low latency and high reliability. At the same time, this method can also optimize the bandwidth allocation and flight trajectory of UAVs, improve the execution efficiency of predictive maintenance tasks, and enhance the intelligent operation and maintenance ability of industrial systems. Description of the Drawings
[0020] Figure 1 It is a schematic diagram of the scenario of computing offloading and resource allocation based on UAV trajectory optimization in an embodiment of the present invention.
[0021] Figure 2 It is a schematic diagram of the training process of various agents in an embodiment of the present invention.
[0022] Figure 3 It is a comparison result diagram of the average age of information of different algorithms in an embodiment of the present invention.
[0023] Figure 4 It is a comparison result diagram of the amount of tasks completed by the cloud center under different algorithms in an embodiment of the present invention.
[0024] Figure 5 It is a schematic diagram of the bandwidth allocation of different UAVs under the FHMA2C algorithm in an embodiment of the present invention.
[0025] Figure 6 It is a distribution diagram of an original UE and UAV example in an embodiment of the present invention.
[0026] Figure 7 This is a UAV motion trajectory diagram exemplified in the embodiments of the present invention. Detailed implementation manners
[0027] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0028] Combined with Figure 1 In the scenario shown, the embodiments of the present invention disclose a computing offloading and resource allocation method based on UAV trajectory optimization. For the device monitoring and predictive maintenance system in the industrial Internet of Things environment, a UAV-assisted cloud-edge-end computing offloading architecture is introduced to solve problems such as the difficulty in ensuring the freshness of device status data and the complexity of the design of heterogeneous device cooperation strategies in the industrial field. This method considers minimizing the average age of information of all device status data under constraints such as computing and communication resources, flight speed, and observation range. To solve this non-convex mixed-integer optimization problem, a heterogeneous multi-agent reinforcement learning algorithm FHMA2C (Federated-HMA2C) combined with federated learning is designed. This algorithm is based on the heterogeneous multi-agent advantage strategy-value algorithm (Heterogeneous Multi-Agent Advantage Actor-Critic Algorithm, HMA2C) framework, solves the problem of the design of mixed strategies, and promotes the cooperation between agents through federated parameter updates. Through offline distributed training and online execution, each agent is allowed to make online decisions in a complex industrial environment.
[0029] A computing offloading and resource allocation method based on UAV trajectory optimization according to an embodiment of the present invention. First, based on a multi-UAV-assisted cloud-edge-end computing offloading architecture, under the constraints of computing and communication resources, flight speed, and observation range, an optimization problem model with the minimum system age of information as the optimization objective, where the decision variables are UAV motion decision, user terminal task offloading ratio decision, computing and offloading to the cloud decision of the UAV edge server, and bandwidth ratio allocation decision of the cloud center for all UAVs; then, the optimization problem model is transformed into a Markov decision process model, three types of heterogeneous agents are constructed for the user terminal, UAV, and cloud center, and based on the heterogeneous multi-agent advantage strategy-value algorithm, and a federated update mechanism is adopted among the same type of agents for offline distributed training and online execution. Specifically, in specific implementation, the system age of information can be represented by the average age of information of all user terminals, and the age of information of a user terminal can be represented by the current time minus the generation time of the latest completed task of the user terminal; the UAV motion decision can be represented by the abscissa and ordinate of the UAV in the horizontal direction and the corresponding displacement changes; the user terminal task offloading ratio decision can be represented by the ratio of the amount of tasks executed locally to the amount of tasks in the user data buffer; the computing and offloading to the cloud decisions of the UAV edge server can be represented by a Boolean variable respectively, including executing at the UAV end of the task queue, offloading to the cloud center, and not executed and not offloaded temporarily; the bandwidth ratio allocation decision of the cloud center can be represented by a bandwidth allocation vector.
[0030] The following describes the specific steps of the embodiment of the present invention in detail in combination with a specific system model.
[0031] Step (1), establish a multi-UAV-assisted cloud-edge-end computing offloading architecture, including a terminal device model, an MEC server model, and a cloud control center model; the specific steps are as follows: (1a) Design a multi-UAV-assisted cloud-edge-terminal computing offloading network architecture, which includes multiple industrial sensors or device controllers, namely user terminals (UEs); multiple unmanned aerial vehicles (UAVs); and a cloud center. Under limited computing and communication resources, UEs can execute some tasks locally or offload them to UAVs. UAVs are equipped with rich computing resources and can provide computing offloading services for UEs, using their strong computing resources to accelerate task processing. The cloud center is the core computing engine of the entire architecture and can receive complex tasks from UAVs, such as long-term device health status assessment, large-scale industrial data analysis, etc. It can quickly complete task processing using its powerful computing ability. In addition, the cloud center or UAVs will further forward the processed task results to the corresponding UEs to achieve real-time device status adjustment, intelligent maintenance decision-making, or alarm triggering, ensuring the closed-loop execution of industrial monitoring tasks; In the embodiment of the present invention, the target area is set as a 1000-meter × 1000-meter square urban environment. The entire area is divided into 10,000 location points, and each location point represents a 10-meter × 10-meter small area. The area contains 30 ground UEs, 4 UAVs, and 1 cloud central control center.
[0032] (1b) During the task execution process, the continuous total execution time is divided into T time slots. For the convenience of analysis and calculation, it is assumed that the time of each time slot is small enough. Therefore, it can be considered that the position of the UAV remains fixed in each time slot. In the embodiment of the present invention, the total number of time slots is set to 200.
[0033] (1c) Establish a terminal device model. There are N UEs in the architecture, and the user set is represented as . The UEs are randomly distributed on the ground. In time slot t, the position of each UE is represented by three-dimensional coordinates as . Among them, and respectively represent the abscissa and ordinate of the user in the map, and the height is fixed at 0, that is, located on the ground. In each time slot, each ground UE will randomly and independently generate compute-intensive tasks. In time slot t, the task generated by UE n can be represented as . Among them, represents the data volume of the task, that is, the scale of data required for task transmission and calculation; represents the age of information of the task, which is used to measure the timeliness of the task. To support local computing of tasks, each terminal is equipped with CPU computing resources, and its computing ability is fixed at , and in the embodiment of the present invention, it is set to 0.2 Kb / time slot.
[0034] (1d) Establish a MEC server model. There are U UAVs equipped with MEC servers deployed in the target area to provide computing offloading services for ground UEs, and their set is represented as 。The initial positions of the UAVs are randomly distributed within the target area. In time slot t, the position of UAV u can be expressed as , where and represent the abscissa and ordinate of the UAV in the horizontal plane respectively, and H is the fixed flight altitude of the UAV, which is set to 150 m in the embodiments of the present invention. In time slot t, the movement of UAV u is expressed as , where and represent the changes in the abscissa and ordinate displacements of the UAV in the horizontal direction respectively. The position of UAV u in the next time slot can be expressed as . To support the computational offloading of tasks, the computing capacity of the MEC server for each UAV u is fixed at , which is set to 20 Kb / slot in the embodiments of the present invention. In the embodiments of the present invention, the movement radius of the UAV is set to 30 m, the observation radius of the UAV is 300 m, and the service radius of the UAV is 200 m.
[0035] (1e), establish a cloud control center model. The cloud center plays a role in task processing and resource scheduling in the architecture. In time slot t, the fixed position of the cloud center is expressed as , where in the embodiments of the present invention, is set. Due to limited communication resources, a bandwidth allocation mechanism is adopted to provide communication services for all UAV-cloud center links. In time slot t, the corresponding bandwidth allocation vector is defined as .
[0036] Step (2), based on the above architecture, establish an industrial equipment status monitoring task model, a computational offloading model, and a communication model; the specific steps are as follows: (2a), exemplarily, a random task model is established in the embodiments of the present invention. The amount of data generated by tasks follows a Poisson distribution, and the average generation rate of each task in each time slot is , which is set to 1 Kb / slot in the embodiments of the present invention. The task generation process can be expressed as ; where k is the size of the task data volume, and the random variable represents the generation situation of the task. When G = 1, it means that a task is generated in the current time slot; when G = 0, it means that no task is generated in the current time slot. G follows a Bernoulli distribution, and the probability of generating a task is p , which is set to 0.3 in the embodiments of the present invention. G can be expressed as ; when a task is generated, they will first be stored in the data buffer of UE n. Subsequently, these tasks will be executed locally or offloaded to the UAV for data processing in a first-in, first-out manner.
[0037] (2b) Establish a user computing offloading model. Define as the proportion of the local execution task volume to the task volume in the user data buffer, . The task volume of local computing by UE n and the task volume offloaded from UE n to the UAV are respectively expressed as ; where, represents the task backlog in the data buffer of UE n at time slot t.
[0038] (2c) Establish a UAV computing offloading model. Each UAV maintains multiple first-in-first-out queues for storing tasks offloaded from different UEs. These tasks are queued in the queues waiting to be processed, and can either be executed locally on the UAV or further offloaded to the cloud center for processing. The computing offloading decisions of the terminals may last for multiple time slots, but the task scheduling and resource allocation decisions of the UAV are updated in each time slot. Each execution time slot of the UAV t only allows processing one queue for edge computing, that is, the UAV u will allocate its computing resources to one of all the task queues of this UAV to achieve efficient scheduling of task processing. The local computing decision and computing offloading decision of the i-th queue of UAV u are respectively represented by and , , , when , it means that the i-th queue of UAV u is executed at the UAV side in time slot t, otherwise, it has not been executed yet; when , it means that the i-th queue of UAV u is offloaded to the cloud center for computing in time slot t, otherwise, it has not been offloaded for computing yet. Therefore, the execution decision of the UAV and the decision to offload to the cloud center satisfy: ; where, represents the total number of all queues of UAV u.
[0039] Correspondingly, the task processing update process of the UAV at time slot t can be expressed as ; where, represents the task backlog in the data buffer of the UAV t at time slot u ; represents the task data volume offloaded from the UAV to the cloud center; represents the task data volume processed by the UAV.
[0040] Due to the powerful computing ability of the cloud center, it can complete tasks with extremely low latency, and its computing time can be ignored.
[0041] (2d), establish a communication model. There are two communication links in the network architecture, namely from UE to UAV edge and from UAV edge to cloud center. The communication of both links can be modeled as an air-to-ground channel. At time slot t, the loss between two entities of the air-to-ground channel includes the line-of-sight loss part and the non-line-of-sight loss part, and can be expressed as ; where, and are the approximate probabilities of line-of-sight loss and non-line-of-sight loss: ; where, a and b are parameters related to the environment, and in the embodiments of the present invention, they are set to 9.61 and 0.61, is the elevation angle.
[0042] and respectively represent line-of-sight loss and non-line-of-sight loss, and can be respectively expressed as ; where, f represents the carrier frequency, and in the embodiments of the present invention, it is set to 2.4 GHz; c represents the speed of light; and represent the environmental parameters in the case of line-of-sight loss and non-line-of-sight loss, and in the embodiments of the present invention, they are set to 1 and 20; represents the Euclidean distance between both ends of the link.
[0043] Step (3), based on the above model, construct an optimization objective for minimizing the age of information of the system, and determine the constraint conditions; including the following specific steps: (3a), in order to ensure the timeliness of data in the system, introduce the age of information AoI to measure the freshness of data. The age of information of UEn at time slot t is defined as the current time minus the generation time of the latest completed task of UE, and can be expressed as ; where, represents the generation time of the latest completed task of UE n at time slot t. During simulation, the calculation of the age of information includes: a. Queuing time, that is, the waiting time of the UE task from generation to being scheduled. When the data packet is waiting in the buffer of the sensor to be collected or processed, its age of information will increase with time. Specifically, it is expressed as the waiting time of the UE task in the UE data buffer / UAV task queue.
[0044] b. UAV collection time and UAV offloading time. The time taken by the UAV to collect data from the UE is called the UAV collection time, and the time taken by the UAV to offload the data packet to the cloud center is called the UAV offloading time. Generally speaking, it is the time to output the task from the sending end (UE / UAV) to the channel, specifically expressed as , where R represents the transmission rate of the sending end (UE / UAV). The simulation sets the UE transmission rate to be 8000 bps. The UAV transmission rate is further calculated according to the communication model. First, calculate the loss calculated according to the communication model, and calculate the signal-to-noise ratio as , represents the transmission power of the sending end, represents the noise power. Then, calculate the UAV transmission rate according to the Shannon formula as , where represents the allocated communication bandwidth.
[0045] c. Data processing delay. The UE task offloading includes two parts: the task volume locally computed by UE n and the task volume offloaded from UE n to the UAV. The data processing delay is the time taken for the data packet to be processed at the UE / UAV. The UE local data processing delay can be expressed as , and the UAV data processing delay can be expressed as .
[0046] The system AoI is the average age of information of all UEs , which can be expressed as .
[0047] (3b). Establish an optimization problem model. The problem studied can be described as: minimizing the system average AoI under the constraints of computing and communication resources, flight speed, and observation range, etc. The objective function and constraints are as follows ; where M represents the motion decision of the UAV, C represents the task offloading ratio decision of the UE, E and O represent the computing and offloading to the cloud decisions of the UAV edge server, B represents the bandwidth ratio allocation decision of the cloud center. The objective function is to minimize the system average AoI defined in the formula. The constraints C1 and C2 ensure that the flight range of the UAV is limited to the target area, X max represents the length of the target area, Y maxrepresents the width of the target area, R move represents the maximum moving distance of the UAV within one time slot; the constraint C3 restricts the task offloading ratio of the UE to ensure the rationality of the offloading ratio; the constraints C4 and C5 represent the edge computing constraints and offloading constraints of the UAV data cache queue; the constraint C6 restricts the bandwidth allocation of the cloud center to ensure that the bandwidth resource allocation does not exceed the total amount; the constraint C7 represents that the transmission power of the UAV does not exceed the maximum power P max , and avoid overload. Through the collaborative optimization of the above decision variables and constraints, the system can efficiently achieve the overall optimization of user task offloading, UAV movement, edge computing, and cloud resource allocation.
[0048] Step (4), transform the optimization problem into a Markov Decision Process (MDP) model; it includes the following specific steps: (4a), determine the agents. In the MDP, the UE, UAV, and a cloud center are regarded as three types of heterogeneous agents. The UE is mainly responsible for generating and offloading computing tasks. The UAV performs computing and task forwarding at the edge layer, while the cloud center provides powerful computing capabilities and global coordination functions. These agents have different observation spaces and action spaces and can make decisions based on their unique observation spaces and action spaces. All agents share the same reward function to achieve the optimization of global performance.
[0049] (4b), determine the global state space. The global state space is the sum of the observation spaces of all agents, including the location information of the UE, data cache space information, location information of the UAV, data queue information, and bandwidth allocation information, etc., and can be expressed as .
[0050] (4c), determine the agent observation spaces. The observation space of each UE agent includes the UE location and data cache space information, and can be expressed as .
[0051] The observation space of each UAV includes the location information of the UAV, the data queue information of the UAV, the location information of the terminals within its coverage, and bandwidth allocation information, and can be expressed as , where are all UEs within the coverage of UAV u.
[0052] The observation space of the cloud center includes the location information and data queue information of all UAVs and bandwidth allocation information, and can be expressed as .
[0053] (4d), determine the action space of the agents. For each time slot t , agents of each type act differently. The actions of the UE agent n can be expressed as . The actions of the UAV agent u can be expressed as . The actions of the cloud center agent can be expressed as .
[0054] (4e), determine the cumulative reward. Since the goal is to minimize the average AoI of the MEC system, the reward is regarded as a penalty, and the penalty is defined as the average age of information of all UEs at time slot t, expressed as . The cumulative reward can be expressed as , where represents the discount factor, usually taking the value , T is the size of the total number of observed time slots, used to balance the importance of the current reward and future rewards. The discount factor is set to 0.86 in the simulation.
[0055] Step (5), according to the MDP modeling, design the heterogeneous agent network structure, and design the Heterogeneous Multi-Agent Advantage Actor-Critic Algorithm (HMA2C) architecture, as Figure 2 shown, which is the HMA2C algorithm training model diagram provided by the embodiment of the present invention; it includes the following specific steps: (5a), each agent is equipped with a policy network and a value network, and interacts with the environment to learn the optimal policy. The online policy network takes the local observation of the agent as input and outputs a probability distribution of an action, representing the selection probability of each possible action in the current state. The online value network evaluates the value of the selected action in the current state, helps the agent judge the quality of the selected action in the given state, and thus provides feedback information for action selection. The input of the online value network of each agent includes the input and output of the online policy network.
[0056] (5b), network structure design: design the policy network and value network structures of three types of agents (UE, UAV, cloud center). The UE policy network design uses a multi-layer perceptron (MLP) as the basic architecture, and takes the UE data cache space information and location information as input. Through the fully connected layer, the local computing decision output (scalar value) is calculated, indicating whether the terminal chooses local computing or offloads tasks; the UE value network design takes the terminal data cache space information and computing decision as input, calculates the corresponding reward value, and measures the quality of the computing decision.
[0057] The terminal information matrix is defined as a three-dimensional matrix. The first two dimensions represent the grid of the observable range centered on the UAV, represented in coordinate form. The size of the third dimension is 2, representing respectively the buffer data size of the UEs within the observable range and the age of information of the UEs within the observable range. The third dimension of the terminal information matrix can be expressed as .
[0058] The UAV policy network is designed as a hybrid neural network combining a convolutional neural network (CNN) and an MLP. The CNN part processes the terminal information matrix (including the geographical location of the terminal relative to the UAV, the buffer data size, and the age of information), extracts the regions with larger data volume and higher age of information through convolutional layers, ensuring that key data is processed preferentially; the MLP part combines the buffer state of the UAV and the currently allocated bandwidth to generate a data scheduling vector, optimizing the data transmission and task offloading strategies; specifically, the terminal information matrix is input into the CNN network, flattened through a Flatten layer, and passed through another fully connected layer to output the movement decision of the UAV. The MLP inputs the buffer state of the UAV and the currently allocated bandwidth and outputs the task execution and task offloading decisions. The final action output of the UAV includes the movement decision, the calculation task execution, and the task offloading decision.
[0059] The UAV value network is designed as a hybrid neural network combining a CNN and an MLP. The UAV state map is defined as a three-dimensional matrix. The first two dimensions represent the map coordinates, and the size of the third dimension is 2, including respectively the total amount of data calculated by the UAV currently and the average age of information of the data calculated by the UAV currently. The UAV state map represents the data processing situation of the UAV. The CNN processes the UAV state map, extracting the current position of the UAV and its task buffer state information; the MLP combines information such as bandwidth and the buffer state of the calculation task to evaluate the value of the decisions on UAV task execution and task offloading. The outputs of the two networks are merged through a concatenate layer, and the merged data passes through a fully connected layer to output a single value, which is the reward value for the UAV.
[0060] The cloud center policy network is designed as a fully connected neural network, taking the state vectors of all UAVs (including UAV positions, task buffer states, etc.) as inputs and outputting a bandwidth allocation policy to coordinate the communication resource allocation among UAVs and ensure the overall network load balance. The cloud center value network is designed as an MLP, taking the state vectors of all UAVs and the bandwidth allocation policy as inputs, calculating the global reward value of the allocation policy, and optimizing the overall resource scheduling.
[0061] (5c), the target policy network and the target value network are constructed. They have the same network structure and initialized network parameters as the online network. The target network adopts a soft update method, and its parameters are copied from the online network at a certain period, so as to make the policy update more stable and reliable. Among them, the parameter updates of the target policy network and the target value network can be expressed as ; Among them, represents the parameters of the target policy network in the (t + 1)-th time slot; represents the parameters of the target policy network in the t-th time slot; represents the parameters of the online policy network in the t-th time slot; represents the parameters of the target value network in the (t + 1)-th time slot; represents the parameters of the target value network in the t-th time slot; represents the parameters of the online value network in the t-th time slot; is the soft update parameter of the network, representing the update period or update frequency of the network. In the embodiment of the present invention, it is set to 0.8. It is set that the target network is updated once every 6 training cycles.
[0062] (5d), each agent explores actions based on its own policy network using - greedy policy. The exploration probability is set to exponential decay. When the exploration probability is less than a certain value or the experience pool is less than a certain value, actions are randomly selected, otherwise actions are selected according to the policy network of each agent respectively Among them, represents the action of each agent, represents the state of each agent, represents the action space. In the embodiment of the present invention, the exploration rate is set to 0.2.
[0063] (5e), after all agents have executed their respective actions, each agent can obtain a unified reward , the environment transfers to the next state , each agent stores the explored experience into its own experience pool. The capacity of the experience pool is limited. When the capacity is full, the new record will replace the earliest record.
[0064] (5f), the agent performs experience replay and samples data from the experience pool for training. To ensure the synchronization of the learning process and the interaction with the environment, in each training cycle, each agent samples data from the experience buffer in batches of size B. To balance the utilization of new data and old data, the latest 0.25×B data plus randomly selected 0.75×B data will be used for training. In the embodiment of the present invention, the batch size is set to 64.
[0065] After taking out the sample in (5g), calculate the online value network loss function and update the value network parameters. The parameter update of the online value network can be expressed as ; where represents the learning rate of the value network, and the simulation is set to 0.002; is the loss function of the value network and can be expressed as ; where D represents the number of samples in each mini-batch sampling, is the Q value calculated by the online value network, while is the target Q value calculated by the target value network, , the target action is the action result output by the target policy network based on and the current network parameters, ; represents the target policy network function.
[0066] In (5h), calculate the online policy network loss function and update the policy network parameters. The parameter update of the online policy network can be expressed as ; where represents the learning rate of the policy network, which is set to 0.001 in the embodiments of the present invention; is the loss function of the online policy network and can be expressed as .
[0067] Step (6), based on the above architecture, introduce the federated update mechanism, design the FHMA2C algorithm FHMA2C based on joint federated update, and perform offline distributed training and online execution on the model; the specific steps are as follows: In (6a), draw on the idea of federated learning and introduce the federated update mechanism. Allow multiple agents to jointly optimize the global model by local model training and periodically sharing model parameters without sharing data, while retaining the characteristics of distributed training and achieving global coordination. Every fixed time time slots, the neural network parameters are exchanged between the same type of federated agents. In the embodiments of the present invention, is set to 6 time slots.
[0068] In (6b), then, each federated agent will retain its own network parameters according to the set weight and perform weighted mixing with the parameters of other agents to form a global model. Promote the entire system to converge towards the global optimal solution. Taking the policy network of the user terminal as an example, the federated update process can be expressed as: ; where A set matrix representing all UE network parameters and represents the federated update weight. Then, the federated parameter update weight of the UAV agent is . In the embodiment of the present invention, is set to 0.5.
[0069] The following conducts simulation experiments in combination with the Python and TensorFlow machine learning platforms, and at the same time compares the embodiments of the present invention with the baseline algorithm to verify the beneficial effects of the present invention.
[0070] In Figure 3 , a comparison result graph of the average age of information of different algorithms provided by the embodiments of the present invention is described. Two other algorithms are simulated and compared, namely: the Heterogeneous Multi-Agent Advantage Actor-Critic (HMAAC) algorithm, and HMA2C has the same architecture but does not adopt federated learning; the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, and the policy of each agent is trained based on the traditional centralized DDPG algorithm. It can be seen from the result graph that the average age of information of the FHMA2C algorithm remains at a low level, and the convergence speed is fast. After convergence, the average age of information is the smallest. It can be seen that the FHMA2C algorithm has superior stability and efficiency in this air-ground integrated MEC scenario, can quickly adapt to environmental changes, and effectively reduce information latency.
[0071] In Figure 4 , a comparison result graph of the amount of tasks completed by the cloud center under different algorithms provided by the embodiments of the present invention is described. The total amount of tasks completed increases linearly with the training cycle. The FHMA2C method has the highest growth rate throughout the training process. Its task completion amount is always higher than other comparison algorithms throughout the training process. After 5000 training cycles, the task completion amount of the cloud center of the FHMA2C algorithm is about 1.94×104, the task completion amount of the cloud center of the HMAAC algorithm is about 1.82×104, and the task completion amount of the cloud center of the MADDPG algorithm is about 1.74×104. Therefore, the task completion amount of the FHMA2C algorithm is about 7% higher than that of the HMAAC algorithm and about 11% higher than that of the MADDPG algorithm, fully reflecting the significant advantages of the FHMA2C algorithm in task processing and providing stronger support for efficient task allocation and processing.
[0072] In Figure 5It describes the schematic diagram of bandwidth allocation for different UAVs under the FHMA2C algorithm provided by the embodiments of the present invention. During the training period from 0 to 1000, the fluctuation range of bandwidth allocation is relatively large, and the allocation ratios of each UAV vary significantly within the range of 0 to 0.5. This is because the resource competition among UAVs is relatively significant in the initial stage of training. After 1000 training cycles, the bandwidth allocation begins to stabilize, and the allocation ratios of each UAV mainly concentrate in the range of 0.2 to 0.3, showing an obvious convergence trend and demonstrating high fairness in bandwidth allocation.
[0073] In Figure 6 it describes a distribution diagram of original UEs and UAVs in an example of the embodiments of the present invention. The positions of UEs are marked with blue triangles, the initial positions of UAVs are marked with orange dots, the observation radius of the UAV is represented by a blue circle, the service radius is represented by a green circle, and the numbers 1 to 4 represent UAV numbers.
[0074] In Figure 7 it describes the trajectory changes of a UAV during the iterative process. The trajectories of each UAV are represented by lines of different colors. It can be seen that the UAVs pass through as many UE areas as possible and avoid unnecessary path overlaps as much as possible, reflecting the coverage efficiency and path optimization effect of the system.
[0075] According to the description of the present invention, those skilled in the art should not find it difficult to see that a heterogeneous multi-agent reinforcement learning algorithm FHMA2C designed by the present invention can effectively ensure the freshness of data transmission in the MEC system and can achieve reasonable bandwidth allocation and trajectory optimization of edge devices.
[0076] Based on the same inventive concept, the embodiments of the present invention also disclose a computing offloading and resource allocation system based on UAV trajectory optimization, including: a problem modeling module, which is used to, based on a multi-UAV-assisted cloud-edge-end computing offloading architecture, under the constraints of computing and communication resources, flight speed, and observation range, take the UAV motion decision, the user terminal task offloading ratio decision, the computing and offloading to the cloud decision of the UAV edge server, and the bandwidth ratio allocation decision of the cloud center for all UAVs as decision variables, and minimize the average age of information of all user terminals as an optimization objective to establish an optimization problem model; a deep reinforcement learning module, which is used to transform the optimization problem model into a Markov decision process model, construct three types of heterogeneous agents for the user terminal, UAV, and cloud center, based on the heterogeneous multi-agent advantage strategy-value algorithm, and adopt a federated update mechanism among the same type of agents to perform offline distributed training and online execution.
[0077] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the calculation offloading and resource allocation method based on UAV trajectory optimization are implemented.
[0078] An embodiment of the present invention also discloses a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the calculation offloading and resource allocation method based on UAV trajectory optimization are implemented.
[0079] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the steps of the method of the present invention are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server. Where the present invention is not described in detail, it is common knowledge to those skilled in the art.
Claims
1. A method for computing offloading and resource allocation based on UAV trajectory optimization, characterized in that, It includes the following steps: Based on the multi-UAV-assisted cloud-edge-end computing offloading architecture, under the constraints of computing and communication resources, flight speed, and observation range, an optimization problem model with the motion decision of UAVs, the decision of the task offloading ratio of user terminals, the decision of computing and offloading to the cloud by UAV edge servers, and the decision of the bandwidth ratio allocation of the cloud center to all UAVs as decision variables and the minimization of the average age of information of all user terminals as the optimization goal; Transform the optimization problem model into a Markov decision process model, construct three types of heterogeneous agents for user terminals, UAVs, and the cloud center, and perform offline distributed training and online execution based on the heterogeneous multi-agent advantage strategy-value algorithm and adopting a federated update mechanism among agents of the same type.
2. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 1, characterized in that, The age of information of a user terminal is the current time minus the generation time of the latest completed task of the user terminal; the motion decision of the UAV is represented by the abscissa and ordinate in the horizontal direction of the UAV and the corresponding displacement changes; the decision of the task offloading ratio of the user terminal is represented by the ratio of the amount of tasks executed locally to the amount of tasks in the user data buffer; the decisions of computing and offloading to the cloud by the UAV edge server are each represented by a Boolean variable, including executing in the UAV end of the task queue, offloading to the cloud center, and not yet executed and not yet offloaded; the decision of the bandwidth ratio allocation of the cloud center is represented by a bandwidth allocation vector.
3. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 1, characterized in that, The global state space of the Markov decision process model is the sum of the observation spaces of all agents, including the location information and data cache space information of user terminals, the location information and data queue information of UAVs, and the bandwidth allocation information of the cloud center.
4. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 1, characterized in that, The observation space of each user terminal agent includes the location information and data cache space information of the user terminal, and the action space is the decision variable of the task offloading ratio; the observation space of the UAV agent includes the location information and data queue information of the UAV, as well as the location information of user terminals within its coverage range and the bandwidth allocation information of the cloud center to it, and the action space is the motion decision variable and the decision variables of computing and offloading to the cloud by the UAV edge server; the observation space of the cloud center agent includes the location information and data queue information of all UAVs and the bandwidth allocation information, and the action space is the decision variable of the bandwidth ratio allocation.
5. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 4, characterized in that, In the heterogeneous multi-agent advantage strategy-value algorithm, the policy network of the user terminal agent takes the user terminal data cache space information and location information as inputs and outputs the local computing decision, and the value network takes the data cache space information and computing decision as inputs to obtain the reward value; the policy network of the cloud center agent takes the state vectors of all UAVs as inputs and outputs the bandwidth allocation policy, and the value network takes the state vectors of all UAVs and the bandwidth allocation policy as inputs to obtain the reward value.
6. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 4, characterized in that, In the heterogeneous multi-agent advantage strategy-value algorithm, the policy network of the UAV agent processes the information matrix of the terminal through a convolutional neural network to obtain the motion decision of the UAV, and obtains the task execution and task offloading decisions through a multi-layer perceptron network according to the data queue information and bandwidth allocation information; wherein the information matrix of the terminal is a three-dimensional matrix, the first two dimensions represent the grid of the observable range centered on the UAV, and the information of the third dimension includes the buffer data size and the age of information of the user terminals within the observable range. The value network of the UAV agent processes the UAV state map through a convolutional neural network, extracts the UAV position and its task buffer state information, and combines the bandwidth allocation information and inputs it into a multi-layer perceptron network. After merging the outputs of the convolutional neural network and the multi-layer perceptron network, the reward value is obtained through a fully connected layer.
7. The method for computing offloading and resource allocation based on UAV trajectory optimization according to claim 1, characterized in that, The offline distributed training of the heterogeneous multi-agent adopts a federated update mechanism, specifically: at fixed time intervals, the same type of federated agents exchange their neural network parameters; after exchanging the parameters, each federated agent retains its own network parameters according to the set weights and performs weighted mixing with the parameters of other agents to form a global model.
8. A system for computing offloading and resource allocation based on UAV trajectory optimization, characterized in that, It includes: A problem modeling module, which is used to optimize the problem model based on the multi-UAV-assisted cloud-edge-terminal computing offloading architecture, with the UAV motion decision, the user terminal task offloading ratio decision, the computing and offloading to the cloud decision of the UAV edge server, and the bandwidth ratio allocation decision of the cloud center for all UAVs as decision variables, and the minimization of the average age of information of all user terminals as the optimization goal. A deep reinforcement learning module, which is used to transform the optimization problem model into a Markov decision process model, construct three types of heterogeneous agents for user terminals, UAVs and cloud centers, and perform offline distributed training and online execution based on the heterogeneous multi-agent advantage strategy-value algorithm and adopting a federated update mechanism among the same type of agents.
9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for computing offloading and resource allocation based on UAV trajectory optimization according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for computing offloading and resource allocation based on UAV trajectory optimization according to any one of claims 1-7.
Citation Information
Patent Citations
Battlefield dynamic data acquisition timeliness optimization method and device and computer equipment
CN115509759A
Edge computing cooperation method based on hybrid strategy in federated mode
CN116737391A
MEC system multi-agent cooperation framework integrating AoI and internal excitation
CN119788702A
Task unloading method of multi-unmanned aerial vehicle assisted MEC system based on Safe MARL
CN120010949A
Cited By
Double-time-scale optimization method for assisting task unloading and caching in cooperation of multiple unmanned aerial vehicles
CN121541943A