Task-oriented energy efficiency optimization method and device for dynamic position deployment of unmanned aerial vehicle

By constructing a multi-model UAV mobile edge computing system and adopting distributed matching game theory and multi-objective reward function optimization strategies, the problems of user association conflict and energy efficiency in UAV systems under dynamic high-density service scenarios are solved, and system-level load balancing and efficient task processing are achieved.

CN122363246APending Publication Date: 2026-07-10PUTIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PUTIAN UNIV
Filing Date
2026-02-27
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing UAV mobile edge computing systems face problems such as user association conflicts, uneven computing load, and insufficient system energy efficiency optimization in dynamic high-density service scenarios. Especially in complex scenarios where the location of ground equipment changes dynamically and task requirements are generated randomly, traditional methods are difficult to adapt to the dynamic fluctuations of the environment in real time, resulting in the inability to guarantee system energy efficiency and service quality.

Method used

A mobile edge computing system model for unmanned aerial vehicles (UAVs) is constructed, which includes continuous domain kinematics, air-to-ground communication transmission, latency calculation, and energy consumption model. A multi-agent deep deterministic policy gradient algorithm is used to generate deterministic action vectors. Task association is achieved through distributed matching game. Latency, energy consumption, and task completion are optimized through a multi-objective reward function. The policy network is trained and updated in combination with a centralized network model.

Benefits of technology

It achieves system-level load balancing, has real-time decision-making capabilities in highly dynamic environments, significantly improves the overall energy efficiency of the system, and can optimize the flight trajectory and task offloading of UAV swarms in complex application scenarios, thereby improving the utilization of computing resources and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122363246A_ABST
    Figure CN122363246A_ABST
Patent Text Reader

Abstract

The application provides a task energy efficiency optimization-oriented unmanned aerial vehicle dynamic position deployment control method and device, and relates to the technical field of unmanned aerial vehicle communication control.The application constructs an unmanned aerial vehicle mobile edge computing system model, collects a current environment state at each discrete time step to generate a local observation vector, and generates a deterministic action vector; then, a distributed matching game with a ground mobile device is performed to complete task association; then, position updating is performed, a multi-target reward function is calculated, and centralized network model training is performed to iteratively update a Critic evaluation network and an Actor strategy network; when the model training reaches a preset requirement, the training is terminated, and a task energy efficiency optimal unmanned aerial vehicle dynamic position deployment control strategy is obtained. The application can realize deep coupling optimization of physical trajectories and offloading decisions, effectively solve user association conflicts and uneven computing load problems in a dynamic scene, and significantly improve system comprehensive energy efficiency while guaranteeing service quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of unmanned aerial vehicle (UAV) communication control and mobile edge computing technology, and more specifically, to a method and apparatus for dynamic location deployment control of UAVs oriented towards mission energy efficiency optimization. Background Technology

[0002] Unmanned aerial vehicles (UAVs) possess advantages such as high mobility, flexible deployment, and line-of-sight communication, enabling them to efficiently complete ground task collection and on-site unloading services, thus overcoming the limitations of traditional fixed edge nodes. They hold broad application prospects in fields such as smart agriculture, smart cities, and emergency communications. However, UAV-assisted mobile edge computing systems still face core bottlenecks: UAVs have limited onboard energy, significantly constraining their continuous operation capability by energy consumption; furthermore, the wireless environment, user location, and task requirements are all dynamic. How to jointly optimize UAV flight trajectories, resource allocation, and user association strategies to maximize system energy efficiency in complex scenarios is a problem that urgently needs to be solved in this field.

[0003] In existing technical solutions, joint control methods based on deep reinforcement learning are commonly used to address the task offloading and trajectory optimization problems of multi-UAV assisted mobile edge computing systems. The core of this approach is to model the system as a Markov decision process and use Deep Deterministic Policy Gradient (DDPG) or Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithms to learn the UAV control policy. When dealing with the critical issue of associating ground equipment with a specific UAV, existing methods generally employ heuristic rules based on proximity or maximum received signal strength. This means that ground equipment is associated with the UAV that is currently closest or has the strongest signal by default, and then task offloading and computation processing are performed.

[0004] However, this existing technical solution has significant limitations. First, it lacks a hybrid decision-making mechanism that can effectively decouple continuous motion control from discrete user association. Simple association rules often ignore the game-theoretic competition in the user association process, easily leading to uneven load distribution among drone swarms, causing some drones to be overloaded while others are idle. Second, existing methods often face training difficulties when the motion space is too large, and the optimization objective focuses on a single performance indicator, making it difficult to find the optimal balance between task collection rate, service latency, and energy consumption. Especially in complex scenarios where the location of ground equipment changes dynamically and task requirements are generated randomly, traditional methods struggle to adapt to the dynamic fluctuations of the environment in real time, resulting in the inability to guarantee the overall energy efficiency and service quality of the system.

[0005] In view of the above, this application is hereby submitted. Summary of the Invention

[0006] The present invention aims to provide a dynamic location deployment control method and device for UAVs with task-oriented energy efficiency optimization, in order to solve the problems of user association conflict, uneven computing load and insufficient system energy efficiency optimization faced by multi-UAV mobile edge computing systems in dynamic high-density service scenarios.

[0007] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:

[0008] A dynamic positioning and deployment control method for unmanned aerial vehicles (UAVs) optimized for mission energy efficiency, applied to the UAV, includes: S1, Construct and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model; S2: At each discrete time step, the current environmental state is collected to generate a local observation vector, and a deterministic action vector is generated through the Actor policy network. S3, based on the new state generated by the deterministic action vector, perform a distributed matching game with the ground mobile device to complete the task association; S4. After completing the task association, perform position update according to the deterministic action vector, calculate the multi-objective reward function including time delay, energy consumption and task completion amount, and perform centralized network model training, iteratively update the Critic evaluation network and Actor policy network. S5. When the number of training iterations of the model reaches a preset threshold or the multi-objective reward function converges to a stable value, training is terminated, and the dynamic position deployment control strategy of the UAV with optimal task energy efficiency is obtained.

[0009] Preferably, the UAV mobile edge computing system model consists of multiple ground mobile devices, multiple rotary-wing UAVs, and ground base stations; The continuous domain kinematic model is a model of the motion of an aerial computing platform composed of multiple rotary-wing UAVs moving continuously in a two-dimensional horizontal plane. Its motion is controlled by displacement vectors, and its position is updated using the following formula: ; ; in, Indicates drone In the next step Two-dimensional horizontal coordinates; Indicates drone At time step Two-dimensional horizontal coordinates; For drones At time step The flight distance is used to control the range of movement; For drones At time step The heading angle is used to control the direction of movement; , These are the cosine function and the sine function, respectively.

[0010] Preferably, the air-to-ground communication transmission model calculates the three-dimensional spatial distance, communication angle, and signal path loss between the rotorcraft UAV and the ground base station using quantitative calculations, and combines this with Shannon's formula to calculate the real-time data transmission rate at different UAV locations. The formula is as follows: ; ; ; ; in, For drones The three-dimensional spatial distance from the ground base station; For drones Communication angle with ground base stations; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; To fix the flight altitude of the drone; For signal path loss; The path loss index; This represents the average loss difference between line-of-sight and non-line-of-sight transmission. The free space path loss constant at a specific frequency; This is the elevation angle deviation constant; This is the additional loss constant for line-of-sight transmission; For drones At time step Data transmission rate; This refers to the bandwidth of the wireless channel; This refers to the drone's transmission power. The power of the Gaussian white noise is given.

[0011] Preferably, the latency calculation model is the total processing latency of the UAV, including local computation latency and offloading transmission latency, and the formula is: ; ; ; in, For drones At time step Total processing latency; For drones At time step Local computation latency; For drones At time step The offloading transmission delay; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The number of CPU cycles required to process a single computing task; The computing frequency of the UAV's onboard processor; The queuing time for a mission in the drone queue; For drones At time step Number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

[0012] Preferably, the energy consumption model represents the total energy consumption of the UAV at a time step, including flight propulsion energy consumption, local computing energy consumption, and communication transmission energy consumption, as shown in the formula: ; ; ; ; ; in, For drones At time step Total energy consumption; Energy consumption for drone flight propulsion; Calculate the power consumption of the drone locally; Energy consumption for drone communication transmission; For drone power; The horizontal flight speed of the drone is determined by the flight distance. The discrete time step; , These are the blade profile induced power and the induced power constant, respectively. This refers to the tip velocity of the rotor blades; The average induced velocity of the drone while it is hovering; This refers to the drag coefficient of the drone fuselage. air density; For rotor realism; The rotor disk area; The effective energy consumption capacitance coefficient of the processor; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The computing frequency of the UAV's onboard processor; This refers to the drone's transmission power. For drones At time step The number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

[0013] Preferably, when generating local observation vectors, a multi-agent deep deterministic strategy gradient algorithm is used to complete the local state observation of the UAV; The local observation vector includes the UAV's current position features, payload features, and environmental perception features, and its expression is: ; in, For drones At time step The local observation vector; For the unmanned aerial vehicle Current horizontal coordinate; The maximum coordinate value preset by the system; For drones At time step The current length of the task queue; For drones At time step The number of tasks collected in the current time slot; For drones With the The relative position distance of ground mobile devices in an unmatched state; This represents the total number of currently unmatched ground mobile devices. The deterministic action vector consists of the flight heading angle, flight distance, and mission unloading ratio coefficient, and its expression is: ; in, For drones At time step Deterministic action vectors; For drones At time step Flight distance; For drones At time step The heading angle; For drones At time step The task offloading ratio coefficient is used to dynamically adjust the allocation of tasks between local computing and backhaul base stations.

[0014] Preferably, when performing distributed matching games with ground mobile devices, the Gale-Shapley algorithm is used for iterative matching game progression, specifically as follows: Ground mobile devices in an unmatched state initiate an association request to the drone at the top of their preference list; The drone retains the current optimal set of requests based on its own computing service capacity limit and rejects the rest. The rejected ground mobile device initiates a request to the next candidate drone in its preference list until all ground mobile devices in the system reach a stable matching state and complete the task association. The preference list sorts the drones within the coverage area based on a rating function, where a lower rating indicates a higher degree of preference. The expression is as follows: ; ; ; in, For ground mobile equipment For drones Overall score; For drones At time step The current length of the task queue; For drones With ground base stations The Euclidean distance; Indicates drone ; For drones With ground mobile equipment The Euclidean distance; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; For ground mobile equipment The horizontal coordinates.

[0015] Preferably, the formula for the multi-objective reward function is: ; in, For drones At time step The target reward value; For drones At time step Total processing latency; For drones At time step Total energy consumption; For drones At time step The number of tasks collected in the current time slot; , , These are the corresponding weighting coefficients.

[0016] Preferably, during centralized network model training, the local observation vectors and deterministic action vectors of all UAVs are integrated to obtain the global state. ; When updating the Critic evaluation network, the global state, the UAV deterministic action vector, the UAV target reward value, and the experience tuple of the next global state are stored in the shared experience replay pool. M experience samples are randomly sampled from the shared experience replay pool. The mean squared error loss function is used, and the difference between the Q-value predicted by the current network and the target Q-value calculated based on the target network is minimized using gradient descent to iteratively optimize the Critic evaluation network parameters. The formula for the mean squared error loss function is: ; ; in, This is the mean square error loss; The number of batch samples collected for the experience playback pool; This represents the current global state. For the target Actor policy network based on The output is a deterministic action vector; The target Q value; For the target Actor policy network based on the next global state The output is a deterministic action vector; Discount factor; The current Critic evaluation network parameters; The network parameters are evaluated by the Critic. The current Critic evaluation network predicts the Q-value; The Q-value of the target Critic evaluation network is used; When updating the Actor policy network, based on the M sampled empirical samples, the gradient of the Actor policy network's objective function is calculated, and the Actor policy network parameters are updated along the gradient ascent direction. The formula for the gradient of the Actor policy network's objective function is: ; in, Describe the policy objective function The gradient; This represents the Q-value of the Critic network output. Regarding the action The gradient; This represents the gradient of the Actor policy network output with respect to the network parameters; For Actor policy networks based on global state The output is a deterministic action vector; The action output by the Actor policy network with respect to its parameters The gradient; For drones The Actor policy network function; These are the network parameters for the current Actor policy; Actor policy network parameters The updated formula is: ; in, is the learning rate of the Actor network.

[0017] The present invention also provides a dynamic positioning and deployment control device for unmanned aerial vehicles (UAVs) optimized for mission energy efficiency, comprising: The model building unit is used to build and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model. The action generation unit is used to collect the current environmental state at each discrete time step to generate a local observation vector, and generate a deterministic action vector through the Actor policy network; The task association unit is used to perform a distributed matching game with the ground mobile device based on the new state generated by the deterministic action vector to complete the task association. The model training unit is used to perform position updates based on the deterministic action vector after completing task association, calculate a multi-objective reward function including latency, energy consumption and task completion amount, and perform centralized network model training, iteratively updating the Critic evaluation network and the Actor policy network. The optimal strategy unit is used to terminate training when the number of training iterations of the model reaches a preset threshold or the multi-objective reward function converges to a stable value, thereby obtaining the UAV dynamic position deployment control strategy with the best task energy efficiency.

[0018] The present invention also provides a dynamic location deployment control device for unmanned aerial vehicles (UAVs) with optimized mission energy efficiency, including a processor and a memory. The memory stores a computer program that can be executed by the processor to implement the dynamic location deployment control method for UAVs with optimized mission energy efficiency as described above.

[0019] The present invention also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a processor of the device on which the computer-readable storage medium resides, implement the above-described task-oriented energy-efficiency optimized UAV dynamic position deployment control method.

[0020] In summary, compared with the prior art, the present invention has the following beneficial effects: This invention enables system-level load balancing. By introducing a task association strategy based on matching game theory, this invention directly incorporates the real-time task queuing load of rotary-wing UAVs into the association decision criteria, effectively solving the problem of task contention in overlapping areas covered by multiple UAVs, avoiding local congestion, and improving the global utilization of computing resources.

[0021] This invention provides real-time decision-making capabilities in highly dynamic environments. It utilizes the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to directly output control variables in a continuous space, avoiding the accuracy loss caused by discretization. Combined with a decentralized execution architecture, this enables UAVs to achieve millisecond-level responses based on local observations, eliminating the need for high-frequency communication and reducing control latency.

[0022] This invention can significantly improve the overall energy efficiency of the system. By jointly optimizing latency, energy consumption, and workload through a multi-objective reward function, this invention guides the UAV swarm to automatically seek Pareto optimal solutions between "moving closer to the user to reduce transmission energy consumption" and "reducing maneuvering flight energy consumption," as well as between "local computation" and "offload and backhaul," thereby significantly reducing the energy consumption per unit task while ensuring service quality.

[0023] This invention breaks through the limitations of traditional single-dimensional optimization by performing real-time joint optimization of flight trajectory control in the physical domain and task offloading ratio in the digital domain, enabling the system to better adapt to complex application scenarios such as smart agriculture, emergency rescue, and smart cities. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of a task-oriented energy efficiency optimization method for dynamic location deployment control of unmanned aerial vehicles (UAVs).

[0026] Figure 2 This is a schematic diagram of a mobile edge computing system for drones provided in Example 1.

[0027] Figure 3 This is a diagram of the joint action generation architecture based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm provided in Example 1.

[0028] Figure 4 This is a simulation diagram of the dynamic flight trajectory and service effect of multiple UAVs provided in Example 1.

[0029] Figure 5 This is a schematic diagram of a UAV dynamic location deployment control device for mission energy efficiency optimization provided in Embodiment 2.

[0030] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0032] Example 1 Embodiment 1 of the present invention provides a dynamic location deployment control method for UAVs with task energy efficiency optimization, which can be implemented by a dynamic location deployment control device for UAVs with task energy efficiency optimization (hereinafter referred to as the control device), and in particular, executed by one or more processors within the control device.

[0033] In this embodiment, the control device may be an electronic device equipped with a processor, which carries a computer program for the task-oriented energy-efficient UAV dynamic position deployment control method and the computer program can be executed, such as a computer, smartphone, smart tablet, workstation, etc., without limitation.

[0034] This invention provides a dynamic positioning and deployment control method and system for UAVs with optimized mission energy efficiency. Its core logic lies in achieving precise control of multi-rotor UAV swarms in complex dynamic environments through the coupling of a multi-agent deep reinforcement learning framework and a matching game mechanism. In practical applications, such as smart agriculture environmental monitoring or post-disaster emergency communication, a large number of ground-based mobile devices are distributed on the ground. These devices typically have limited computing power and battery life, requiring the offloading of computational tasks to rotorcraft UAVs covering them.

[0035] like Figure 1 As shown, a dynamic location deployment control method for UAVs oriented towards mission energy efficiency optimization includes steps S1 to S5.

[0036] S1, construct and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model.

[0037] In step S1, the system first performs a detailed model of the mobile edge computing system.

[0038] like Figure 2 As shown, this embodiment constructs a heterogeneous network system consisting of multiple mobile devices (MDs), multiple unmanned aerial vehicles (UAVs), and a base station (BS). The system operation mode is set to discrete time step, and the basic physical parameters (such as UAV flight altitude H, base station horizontal coordinates, wireless channel bandwidth W, etc.) and constraint parameters (such as the maximum flight distance of UAVs in a single time step, the maximum task processing capacity φ, etc.) of all devices are initialized.

[0039] In the continuous domain kinematics modeling, each rotorcraft UAV operates in three-dimensional space, but the dynamic adjustment of its deployment position is mainly reflected on the two-dimensional horizontal plane, with its height H kept at a fixed value to simplify control complexity. At discrete time steps t, the horizontal coordinates of rotorcraft UAV i are updated through its output motion commands.

[0040] Specifically, the controller obtains the flight distance output by the Actor network. with flight heading angle The lateral and longitudinal displacement increments are calculated using trigonometric functions to determine the new position at time t+1.

[0041] The position update formula is: ; ; in, Indicates drone In the next step Two-dimensional horizontal coordinates; Indicates drone At time step Two-dimensional horizontal coordinates; For drones At time step The flight distance is used to control the range of movement; For drones At time step The heading angle is used to control the direction of movement; , These are the cosine function and the sine function, respectively.

[0042] This continuous spatial coordinate update mechanism avoids the deployment accuracy loss caused by discrete grid partitioning, allowing the drone swarm to obtain the optimal communication line of sight by fine-tuning its position.

[0043] In terms of communication transmission modeling, since the UAV operates at high altitude, the link between it and the ground base station (BS) and ground mobile equipment (MD) is mainly affected by line-of-sight path loss. The system calculates the three-dimensional spatial distance and elevation angle between the UAV and the base station in real time and incorporates them into a specific air-to-ground channel model. This is achieved by introducing a path loss exponent. Loss difference In addition, parameters such as elevation angle deviation constant accurately characterize the attenuation characteristics of the signal during propagation.

[0044] Based on the calculated path loss value, the system uses Shannon's formula to determine the current real-time transmission rate.

[0045] The formulas for three-dimensional spatial distance, communication angle, signal path loss, and real-time data transmission rate are as follows: ; ; ; ; in, For drones The three-dimensional spatial distance from the ground base station; For drones Communication angle with ground base stations; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; To fix the flight altitude of the drone; For signal path loss; The path loss index; This represents the average loss difference between line-of-sight and non-line-of-sight transmission. The free space path loss constant at a specific frequency; This is the elevation angle deviation constant; This is the additional loss constant for line-of-sight transmission; For drones At time step Data transmission rate; This refers to the bandwidth of the wireless channel; This refers to the drone's transmission power. The power of the Gaussian white noise is given.

[0046] The data transmission rate directly determines the efficiency of transmitting mission data from the relay UAV back to the core network.

[0047] In computational latency modeling, the system divides the total processing latency into two dimensions: local processing and offload transmission. Local computational latency is limited by the computing frequency of the UAV's onboard processor and the queuing status of the current task queue. If the UAV is currently heavily loaded, newly accessed tasks will experience a long queuing wait time. Offload transmission latency depends on the ratio of the amount of data to be processed to the transmission rate achievable under the current wireless channel bandwidth. Through this two-layer latency modeling, the system provides a quantitative performance indicator for the subsequent reward function.

[0048] The formula for the delay calculation model is: ; ; ; in, For drones At time step Total processing latency; For drones At time step Local computation latency; For drones At time step The offloading transmission delay; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The number of CPU cycles required to process a single computing task; The computing frequency of the UAV's onboard processor; The queuing time for a mission in the drone queue; For drones At time step Number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

[0049] In the UAV energy consumption modeling section, this embodiment considers the physical motion characteristics of the rotary-wing UAV in detail. Flight propulsion energy consumption is characterized by a complex power model, which represents the total energy consumption of the UAV at each time step, including flight propulsion energy consumption, local computing energy consumption, and communication transmission energy consumption, and integrates blade profile induced power, induced power, and fuselage drag. The formula is: ; ; ; ; ; in, For drones At time step Total energy consumption; Energy consumption for drone flight propulsion; Calculate the power consumption of the drone locally; Energy consumption for drone communication transmission; For drone power; The horizontal flight speed of the drone is determined by the flight distance. The discrete time step; , These are the blade profile induced power and the induced power constant, respectively. This refers to the tip velocity of the rotor blades; The average induced velocity of the drone while it is hovering; This refers to the drag coefficient of the drone fuselage. air density; For rotor realism; The rotor disk area; The effective energy consumption capacitance coefficient of the processor; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The computing frequency of the UAV's onboard processor; This refers to the drone's transmission power. For drones At time step The number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

[0050] When a drone accelerates to approach a user, its propulsion power increases significantly; while hovering, although the horizontal displacement is zero, it still consumes power to maintain lift. Local computing power consumption is proportional to the processor's effective capacitance and the square of the computing frequency. Communication transmission power consumption is the cumulative transmission power over the transmission duration. By monitoring these three energy efficiency components in real time, the system guides the drone to achieve a balance between service quality and energy consumption through data feedback.

[0051] This step provides a unified mathematical model and calculation basis for the subsequent location updates, action execution, and latency / energy consumption quantification of the drone, giving all decision-making and execution processes a quantitative standard.

[0052] Existing technologies mostly employ a single model or a fragmented multi-model design, failing to achieve full-link model linkage of motion, communication, latency, and energy consumption. This makes it impossible to accurately depict the chain effect of UAV position changes on mission energy efficiency. This step integrates the four core models according to the actual execution logic, realizing full-link quantitative mapping of position-communication-latency-energy consumption, enabling the UAV to accurately assess the energy efficiency of location deployment.

[0053] S2 collects the current environmental state at each discrete time step to generate a local observation vector, and generates a deterministic action vector through the Actor policy network.

[0054] In step S2, the system generates joint actions based on a multi-agent deep deterministic policy gradient algorithm. This embodiment employs a centralized training and decentralized execution architecture.

[0055] During the training phase, a global Critic network acquires the state and action information of all UAVs to evaluate the merits of the current strategy. During the execution phase, each rotorcraft UAV only needs to run its own Actor network, which extracts features and performs nonlinear mapping to output a deterministic action vector that conforms to physical constraints.

[0056] like Figure 3 As shown, each rotary-wing UAV initializes a user preference list, calculates the user's utility function, and then returns the result to be stored in the experience replay pool. The UAV acquires a local observation vector at time t, which is normalized to ensure the uniformity of the magnitudes of different physical quantities. The observation vector contains current position features (such as its own current position coordinates), load features (such as the remaining capacity of the task queue and the total number of tasks collected in the current time slot), and environmental perception features (such as the distribution heat of unserved mobile devices in the surrounding area as perceived by sensors), and its expression is: ; in, For drones At time step The local observation vector; For the unmanned aerial vehicle Current horizontal coordinate; The maximum coordinate value preset by the system; For drones At time step The current length of the task queue; For drones At time step The number of tasks collected in the current time slot; For drones With the The relative position distance of ground mobile devices in an unmatched state; This represents the total number of currently unmatched ground mobile devices.

[0057] After receiving local observation vectors, the Actor network (i.e., the online policy network) processes them through a multi-layer fully connected neural network to output a three-dimensional continuous action vector, i.e., a deterministic action vector. The first component of this vector is the flight heading angle, which is linearly transformed to the range of zero to twice pi after being mapped using a hyperbolic tangent activation function. The second component is the flight distance, mapped to the range of zero to the maximum single-step displacement. The third component is a task unloading ratio coefficient. This ratio coefficient is a key decision variable, determining what percentage of the ground-based task data collected by the UAV is processed on a local edge server and what percentage is relayed to the base station via a wireless link.

[0058] Its expression is: ; in, For drones At time step Deterministic action vectors; For drones At time step Flight distance; For drones At time step The heading angle; For drones At time step The task offloading ratio coefficient is used to dynamically adjust the allocation of tasks between local computing and backhaul base stations.

[0059] This combined output of physical domain trajectory and digital domain resource allocation enables integrated control at the system level.

[0060] This step enables distributed autonomous perception and decision-making on the UAV, allowing the UAV to independently complete "state acquisition-action generation" based on its own perceived local environmental information, without relying on global information or central nodes, thus laying the decision-making foundation for mission energy efficiency optimization.

[0061] S3, based on the new state generated by the deterministic action vector, performs a distributed matching game with the ground mobile device to complete the task association.

[0062] In step S3, the system executes a task association strategy based on matching game theory to resolve task competition conflicts in overlapping areas covered by multiple machines.

[0063] When multiple rotary-wing UAVs simultaneously cover a single ground-based mobile device (MD), the traditional principle of proximity access often leads to some UAVs in the central location being overloaded, while UAVs in the peripheral locations remain idle.

[0064] This embodiment guides association by constructing a multi-dimensional scoring function. Based on the calculated scoring function, drones within the coverage area are ranked, and a preference list is constructed. The expression for this list is: ; ; ; in, For ground mobile equipment For drones Overall score; For drones At time step The current length of the task queue; For drones With ground base stations The Euclidean distance; For drones With ground mobile equipment The Euclidean distance; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; For ground mobile equipment The horizontal coordinates.

[0065] When constructing its preference list, the Ground Mobile Device (MD) not only considers the spatial distance to the drones but also obtains the drone's task queue length and the transmission cost from the drone to the base station in real time. Through this comprehensive scoring mechanism, the task flow automatically tilts towards drones with good link quality and light computational load.

[0066] When performing distributed matching games with ground mobile devices, the Gale-Shapley algorithm is used for iterative matching games.

[0067] Unmatched ground-based mobile devices (MDs) send requests to the highest-rated drone in their preference list. Upon receiving multiple requests, rotary-wing drones (UAVs), based on their computational capacity limits, temporarily retain the best applicant and reject the rest. Rejected devices do not abandon service but instead send new requests to the next candidate in the preference list. This cyclical process of proposal and rejection continues until all devices have found satisfactory partners or can no longer make new matches, thus achieving global load balancing at the discrete association decision level.

[0068] This step resolves the issues of overlapping coverage and task contention caused by the dynamic positioning of multiple drones, ensuring that the dynamic deployment of drones is adapted to task allocation. It achieves stable task association between drones and the Management Device (MD), guaranteeing that the service range and capacity of each drone match its new state, avoiding overload or resource idleness. Task association is completed through a distributed game theory approach, eliminating the need for a central node for unified scheduling, which aligns with the distributed execution framework characteristics of drones.

[0069] S4. After completing the task association, perform position update based on the deterministic action vector, calculate a multi-objective reward function including time delay, energy consumption and task completion amount, and perform centralized network model training, iteratively updating the Critic evaluation network and Actor policy network.

[0070] In step S4, the system drives policy evolution through multi-objective reward design and network model training. The design logic of the multi-objective reward function is to penalize high latency and high energy consumption behaviors while rewarding high task processing volume.

[0071] The formula for the multi-objective reward function is: ; in, For drones At time step The target reward value; For drones At time step Total processing latency; For drones At time step Total energy consumption; For drones At time step The number of tasks collected in the current time slot; , , These are the corresponding weighting coefficients.

[0072] By adjusting the weighting coefficients, the system exhibits different optimization preferences. In the initial stage when the battery is fully charged, the task weights can be increased. To improve throughput; and during the power alarm phase, an energy consumption penalty is added to force the drone to adopt a more energy-efficient flight and computing mode.

[0073] During network training, the local observation vectors and deterministic action vectors of all UAVs are integrated to obtain the global state. .

[0074] The system stores the experience tuples (including global state, UAV deterministic action vector, UAV target reward value and next global state) generated in each time slot into the shared experience replay pool.

[0075] When updating the Critic network, a mini-batch of samples M is randomly drawn from the pool. The mean squared error loss function is used, and gradient descent is employed to minimize the difference between the Q-value predicted by the current network and the target Q-value calculated based on the target network. Figure 3 The TD error in the evaluation network is updated using the backpropagation algorithm to update the weights of the evaluation network.

[0076] The formula for the mean squared error loss function is: ; ; in, This is the mean square error loss; The number of batch samples collected for the experience playback pool; This represents the current global state. For the target Actor policy network based on The output is a deterministic action vector; The target Q value; For the target Actor policy network based on the next global state The output is a deterministic action vector; Discount factor; The current Critic evaluation network parameters; Critic evaluates network parameters for the target network. The current Critic evaluation network predicts the Q-value; The Q-value of the target Critic evaluation network is used.

[0077] When updating the Actor network, the gradient of the objective function of the Actor policy network is calculated using the policy gradient theorem. The parameters are then adjusted along the gradient ascent direction given by the Critic network so that the actions performed by the UAV can obtain higher cumulative rewards.

[0078] The formula for the gradient of the objective function of the Actor policy network is: ; in, Describe the policy objective function The gradient; This represents the Q-value of the Critic network output. Regarding the action The gradient indicates the direction for action improvement; This represents the gradient of the Actor policy network output with respect to the network parameters; For Actor policy networks based on global state The output is a deterministic action vector; The action output by the Actor policy network with respect to its parameters The gradient; For drones The Actor policy network function; These are the network parameters for the current Actor policy.

[0079] Actor policy network parameters The updated formula is: ; in, is the learning rate of the Actor network.

[0080] To ensure training stability, the system also introduces a soft update mechanism, in which the parameters of the target network slowly follow the updates of the online network at a very small ratio, effectively preventing oscillations and non-convergence phenomena common in reinforcement learning.

[0081] This step translates the UAV's decision-making instructions into actual action execution and task processing, and accurately quantifies the execution effect, providing real-world data for reward calculation. A multi-objective reward function quantifies the energy efficiency of this decision, achieving feedback from "decision-execution-effect quantification," allowing the UAV to learn from the execution results. Centralized network training iteratively updates the Critic and Actor networks, resolving the game theory problem of multi-UAV collaboration, and optimizing the "state-action" mapping rules of the Actor policy network, ensuring that subsequent action generation better aligns with the goal of "optimal task energy efficiency."

[0082] S5. When the number of training iterations of the model reaches a preset threshold or the multi-objective reward function converges to a stable value, training is terminated, and the dynamic position deployment control strategy of the UAV with optimal task energy efficiency is obtained.

[0083] When any termination condition is triggered, the centralized network model training is immediately terminated, and the parameters of the finally converged Actor policy network are solidified to form a dynamic location deployment control strategy for UAVs. This strategy contains a complete decision-making logic of "state perception - action generation - location update - task scheduling".

[0084] The solidified optimal control strategy is distributed to all UAVs and localized. In practical applications, the UAVs can directly perform autonomous perception, decision-making and execution at discrete time steps based on this strategy without the need for further network training.

[0085] In actual operation, the initially randomly distributed swarm of rotary-wing UAVs gradually learned the distribution pattern of ground mobile equipment (MD) through continuous training and interaction.

[0086] like Figure 4 As shown, after 7000 rounds of reinforcement learning training for UAVs, the dynamic positioning deployment and mission service effect of three UAVs in a two-dimensional mission area are visualized. The red curve represents the trajectory of UAV 0, the green curve represents the trajectory of UAV 1, and the blue curve represents the trajectory of UAV 2. Dots represent the MD position, triangles represent the BS position, and both the x-coordinate and y-coordinate represent horizontal distances.

[0087] The drones autonomously plan smooth flight paths and dynamically deploy above areas with high task density. Once tasks in a certain area are completed, the drones predict the next task hotspot based on environmental perception characteristics and adjust their flight course in advance to migrate. At the same time, by dynamically adjusting the offloading ratio, the drones tend to process tasks locally when backhaul bandwidth is limited, and tend to utilize the powerful computing power of the base station when local computing pressure is too high.

[0088] This embodiment ensures that control commands are based on the most realistic environmental feedback by automatically recalculating path loss and updating transmission rates through real-time monitoring of wireless environment changes. The Actor network, through a specific activation function design, guarantees the legality and smoothness of output actions. By introducing weight factors to prioritize high-priority tasks, the system's resilience in emergency scenarios is improved. During training and updates, experience replay and soft update techniques significantly enhance the algorithm's robustness when handling high-dimensional, non-stationary environmental data.

[0089] This method, which integrates the discrete coordination capabilities of game theory into a deep reinforcement learning continuous control framework, achieves deep coupling between physical layer mobility and application layer computing tasks through the collaborative operation of multiple steps. In smart city inspection scenarios, the system reduces the latency of video data transmission by more than 30% by optimizing the deployment location of drones, and extends the overall endurance of the drone swarm by 20% through reasonable task allocation. In disaster relief scenarios, when ground communication base stations are damaged, the system quickly establishes an aerial computing platform and ensures real-time processing of critical data at the rescue site through dynamic position adjustments, gaining valuable time for rescue decision-making.

[0090] Furthermore, the present invention exhibits excellent scalability. When the system needs to connect more rotary-wing drones to expand the coverage area, due to the decentralized execution distributed architecture, the newly added drones only need to load a pre-trained Actor network model to cooperate with the existing drone swarm through local observation, without requiring hardware upgrades or reconstruction of the entire control center. This self-organizing and adaptive characteristic gives the present invention significant technical advantages and practical value in handling large-scale IoT device access.

[0091] In summary, compared with the prior art, the present invention has the following beneficial effects: First, it achieves system-level load balancing. By introducing a task association strategy based on matching game theory, this invention directly incorporates the real-time task queuing load of rotary-wing UAVs into the association decision criteria. Compared to traditional association methods that rely solely on distance or signal strength, this technical solution automatically guides task flows from UAVs with higher loads to those with relatively abundant resources through a multi-round negotiation mechanism. This effectively eliminates task contention conflicts in overlapping areas covered by multiple UAVs, significantly improving the overall utilization rate of network computing resources.

[0092] Secondly, it possesses real-time decision-making capabilities in highly dynamic environments. This invention utilizes a multi-agent deep deterministic policy gradient algorithm to directly output control variables within a continuous action space, completely resolving the control accuracy loss problem caused by the discretized action space. Coupled with a decentralized execution architecture, each rotorcraft UAV only needs to acquire local observation information to generate flight and unloading decisions within milliseconds, eliminating the need for high-frequency synchronous communication with the ground central controller. This significantly reduces system control signaling overhead and decision latency, enabling rapid response to random movements of ground equipment and sudden fluctuations in mission requirements.

[0093] Third, it significantly improves the overall energy efficiency of the system. This invention constructs a refined energy efficiency optimization evaluation system by designing a multi-objective reward function that integrates latency, energy consumption, and workload. This system guides the UAV swarm to automatically seek Pareto optimal solutions between "moving closer to the user to reduce access transmission energy consumption" and "reducing maneuvering to reduce propulsion energy consumption," and between "local computing processing" and "offloading to the base station for backhaul processing." While ensuring the service quality of ground equipment, it significantly reduces the average energy consumption per unit task and extends the operating time of the UAV swarm.

[0094] Fourth, it achieves deep coupling optimization of physical domain trajectory and digital domain decision-making. This invention overcomes the limitations of traditional methods that optimize trajectory planning and resource allocation step-by-step, integrating the flight heading, displacement distance, and task unloading ratio of the rotary-wing UAV into a unified action space for real-time joint control. This integrated optimization mode allows the deployment of physical locations to closely follow the distribution of computing tasks, ensuring a high degree of matching between communication link quality and computing resource supply, and providing more reliable technical support for complex application scenarios such as smart agriculture, emergency rescue, and smart cities.

[0095] Fifth, it enhances the system's robustness and scalability. The distributed architecture based on multi-agent reinforcement learning allows for large-scale refactoring of the overall control logic when increasing or decreasing the number of drones. The agents, through cooperative strategies learned during centralized training, exhibit strong synergy during execution. Even if some drones deplete their power and leave service, the remaining drones can autonomously adjust their coverage and offloading strategies based on local perception, ensuring the continuity of edge computing services.

[0096] In summary, this invention provides an efficient, flexible, and highly adaptive edge computing control scheme for unmanned aerial vehicles (UAVs) by integrating the continuous control capabilities of deep reinforcement learning with the discrete coordination capabilities of game theory, effectively solving the problems of resource conflicts and energy efficiency bottlenecks in dynamic scenarios.

[0097] Example 2 like Figure 5 As shown, the second embodiment of the present invention also provides a dynamic positioning and deployment control device for unmanned aerial vehicles (UAVs) optimized for mission energy efficiency, comprising: The model building unit is used to build and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model. The action generation unit is used to collect the current environmental state at each discrete time step to generate a local observation vector, and generate a deterministic action vector through the Actor policy network; The task association unit is used to perform a distributed matching game with the ground mobile device based on the new state generated by the deterministic action vector, and complete the task association. The model training unit is used to perform position updates based on the deterministic action vector after completing task association, calculate a multi-objective reward function including latency, energy consumption and task completion amount, and perform centralized network model training, iteratively updating the Critic evaluation network and the Actor policy network. The optimal strategy unit is used to terminate training when the number of training iterations of the model reaches a preset threshold or the multi-objective reward function converges to a stable value, thereby obtaining the UAV dynamic position deployment control strategy with the best task energy efficiency.

[0098] Example 3 The third embodiment of the present invention also provides a dynamic location deployment control device for unmanned aerial vehicles (UAVs) with optimized mission energy efficiency, which includes a memory and a processor. The memory stores a computer program, which can be executed by the processor to implement the dynamic location deployment control method for UAVs with optimized mission energy efficiency as described above.

[0099] Example 4 The fourth embodiment of the present invention also provides a computer-readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, they implement the above-described task-oriented energy-efficiency optimized UAV dynamic position deployment control method.

[0100] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A dynamic positioning and deployment control method for unmanned aerial vehicles (UAVs) oriented towards mission energy efficiency optimization, applied to the UAV terminal, characterized in that, include: Construct and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model; At each discrete time step, the current environmental state is collected to generate a local observation vector, and a deterministic action vector is generated through an Actor policy network. Based on the new state generated by the deterministic action vector, a distributed matching game with the ground mobile device is performed to complete the task association; After completing the task association, position updates are performed based on the deterministic action vector, a multi-objective reward function including latency, energy consumption and task completion amount is calculated, and centralized network model training is performed to iteratively update the Critic evaluation network and the Actor policy network. Training is terminated when the number of training iterations reaches a preset threshold or the multi-objective reward function converges to a stable value, thus obtaining the UAV dynamic location deployment control strategy with optimal mission energy efficiency.

2. The method for dynamic positioning and deployment control of unmanned aerial vehicles (UAVs) based on mission energy efficiency optimization according to claim 1, characterized in that... The UAV mobile edge computing system model consists of multiple ground mobile devices, multiple rotary-wing UAVs, and ground base stations; The continuous domain kinematic model is a model of the motion of an aerial computing platform composed of multiple rotary-wing UAVs moving continuously in a two-dimensional horizontal plane. Its motion is controlled by displacement vectors, and its position is updated using the following formula: ; ; in, Indicates drone In the next step Two-dimensional horizontal coordinates; Indicates drone At time step Two-dimensional horizontal coordinates; For drones At time step The flight distance is used to control the range of movement; For drones At time step The heading angle is used to control the direction of movement; , These are the cosine function and the sine function, respectively.

3. The method for dynamic positioning and deployment control of unmanned aerial vehicles (UAVs) oriented towards mission energy efficiency optimization according to claim 2, characterized in that... The air-to-ground communication transmission model quantifies the three-dimensional spatial distance, communication angle, and signal path loss between the rotorcraft UAV and the ground base station, and uses Shannon's formula to calculate the real-time data transmission rate at different UAV locations. The formula is as follows: ; ; ; ; in, For drones The three-dimensional spatial distance from the ground base station; For drones Communication angle with ground base stations; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; To fix the flight altitude of the drone; For signal path loss; The path loss index; This represents the average loss difference between line-of-sight and non-line-of-sight transmission. The free space path loss constant at a specific frequency; This is the elevation angle deviation constant; This is the additional loss constant for line-of-sight transmission; For drones At time step The data transmission rate; This refers to the bandwidth of the wireless channel; This refers to the drone's transmission power. The power of the Gaussian white noise is given.

4. The method for dynamic positioning and deployment control of unmanned aerial vehicles (UAVs) oriented towards mission energy efficiency optimization according to claim 3, characterized in that... The latency calculation model is the total processing latency of the UAV, including local computation latency and offloading transmission latency, and the formula is: ; ; ; in, For drones At time step Total processing latency; For drones At time step Local computation latency; For drones At time step The offloading transmission delay; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The number of CPU cycles required to process a single computing task; The computing frequency of the UAV's onboard processor; The queuing time for a mission in the drone queue; For drones At time step Number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

5. A dynamic positioning and deployment control method for UAVs based on mission energy efficiency optimization according to claim 4, characterized in that... The energy consumption model represents the total energy consumption of the UAV at a time step, including flight propulsion energy consumption, local computing energy consumption, and communication transmission energy consumption. The formula is as follows: ; ; ; ; ; in, For drones At time step Total energy consumption; Energy consumption for drone flight propulsion; Calculate the power consumption of the drone locally; Energy consumption for drone communication transmission; For drone power; The horizontal flight speed of the drone is determined by the flight distance. The discrete time step; , These are the blade profile induced power and the induced power constant, respectively. This refers to the tip velocity of the rotor blades; The average induced velocity of the drone while it is hovering; This refers to the drag coefficient of the drone fuselage. air density; For rotor realism; The rotor disk area; The effective energy consumption capacitance coefficient of the processor; This represents the maximum task processing capacity of the drone within a single time step. For drones At time step The number of local computing tasks; The computing frequency of the UAV's onboard processor; This refers to the drone's transmission power. For drones At time step The number of unloaded transfer tasks; The data size for a single computation task; For drones At time step The data transmission rate.

6. The method for dynamic positioning and deployment control of unmanned aerial vehicles (UAVs) oriented towards mission energy efficiency optimization according to claim 1, characterized in that... When generating local observation vectors, a multi-agent deep deterministic strategy gradient algorithm is used to complete the local state observation of the UAV. The local observation vector includes the UAV's current position features, payload features, and environmental perception features, and its expression is: ; in, For drones At time step The local observation vector; For the unmanned aerial vehicle Current horizontal coordinate; The maximum coordinate value preset by the system; For drones At time step The current length of the task queue; For drones At time step The number of tasks collected in the current time slot; For drones With the The relative position distance of ground mobile devices in an unmatched state; This represents the total number of currently unmatched ground mobile devices. The deterministic action vector consists of the flight heading angle, flight distance, and mission unloading ratio coefficient, and its expression is: ; in, For drones At time step Deterministic action vectors; For drones At time step Flight distance; For drones At time step The heading angle; For drones At time step The task offloading ratio coefficient is used to dynamically adjust the allocation of tasks between local computing and backhaul base stations.

7. A dynamic positioning and deployment control method for UAVs oriented towards mission energy efficiency optimization according to claim 1, characterized in that... In performing distributed matching games with ground mobile devices, the Gale-Shapley algorithm is used for iterative matching game progression, specifically as follows: Ground mobile devices in an unmatched state initiate an association request to the drone at the top of their preference list; The drone retains the current optimal set of requests based on its own computing service capacity limit and rejects the rest. The rejected ground mobile device initiates a request to the next candidate drone in its preference list until all ground mobile devices in the system reach a stable matching state and complete the task association. The preference list sorts the drones within the coverage area based on a rating function, where a lower rating indicates a higher degree of preference. The expression is as follows: ; ; ; in, For ground mobile equipment For drones Overall score; For drones At time step The current length of the task queue; For drones With ground base stations The Euclidean distance; Indicates drone ; For drones With ground mobile equipment The Euclidean distance; For drones Horizontal coordinates; The horizontal coordinates of the ground base station; For ground mobile equipment The horizontal coordinates.

8. A dynamic positioning and deployment control method for UAVs based on mission energy efficiency optimization according to claim 5, characterized in that... The formula for the multi-objective reward function is: ; in, For drones At time step The target reward value; For drones At time step Total processing latency; For drones At time step Total energy consumption; For drones At time step The number of tasks collected in the current time slot; , , These are the corresponding weighting coefficients.

9. A dynamic positioning and deployment control method for unmanned aerial vehicles (UAVs) oriented towards mission energy efficiency optimization according to claim 8, characterized in that... During centralized network model training, the local observation vectors and deterministic action vectors of all UAVs are integrated to obtain the global state. ; When updating the Critic evaluation network, the global state, the UAV deterministic action vector, the UAV target reward value, and the experience tuple of the next global state are stored in the shared experience replay pool. M experience samples are randomly sampled from the shared experience replay pool. The mean squared error loss function is used, and the difference between the Q-value predicted by the current network and the target Q-value calculated based on the target network is minimized using gradient descent to iteratively optimize the Critic evaluation network parameters. The formula for the mean squared error loss function is: ; ; in, This is the mean square error loss; The number of batch samples collected for the experience playback pool; This represents the current global state. For the target Actor policy network based on The output is a deterministic action vector; The target Q value; For the target Actor policy network based on the next global state The output is a deterministic action vector; Discount factor; The current Critic evaluation network parameters; The network parameters are evaluated by the Critic. The current Critic evaluation network predicts the Q-value; The Q-value of the target Critic evaluation network is used; When updating the Actor policy network, based on the M sampled empirical samples, the gradient of the Actor policy network's objective function is calculated, and the Actor policy network parameters are updated along the gradient ascent direction. The formula for the gradient of the Actor policy network's objective function is: ; in, Describe the policy objective function The gradient; This represents the Q-value of the Critic network output. Regarding the action The gradient; This represents the gradient of the Actor policy network output with respect to the network parameters; For Actor policy networks based on global state The output is a deterministic action vector; The action output by the Actor policy network with respect to its parameters The gradient; For drones The Actor policy network function; These are the network parameters for the current Actor policy; Actor Policy Network Parameters The updated formula is: ; in, is the learning rate of the Actor network.

10. A dynamic positioning and deployment control device for unmanned aerial vehicles (UAVs) optimized for mission energy efficiency, characterized in that, include: The model building unit is used to build and load a UAV mobile edge computing system model that includes a continuous domain kinematics model, an air-to-ground communication transmission model, a time delay calculation model, and an energy consumption model. The action generation unit is used to collect the current environmental state at each discrete time step to generate a local observation vector, and generate a deterministic action vector through the Actor policy network; The task association unit is used to perform a distributed matching game with the ground mobile device based on the new state generated by the deterministic action vector to complete the task association. The model training unit is used to perform position updates based on the deterministic action vector after completing task association, calculate a multi-objective reward function including latency, energy consumption and task completion amount, and perform centralized network model training, iteratively updating the Critic evaluation network and the Actor policy network. The optimal strategy unit is used to terminate training when the number of training iterations of the model reaches a preset threshold or the multi-objective reward function converges to a stable value, thereby obtaining the UAV dynamic position deployment control strategy with the best task energy efficiency.