Digital twin task optimization method and system based on unmanned aerial vehicle mobile communication system

By constructing a channel model between drones and end users in smart cities and optimizing task hierarchical processing, drones are scheduled as mobile edge computing nodes, solving the problem of edge server overload, improving system performance and resource utilization, and enhancing system flexibility and reliability.

CN121056911BActive Publication Date: 2026-07-31GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2025-08-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In smart cities, edge server overload leads to a decline in service quality and low resource utilization. How can we use drone-assisted computing to rationally allocate resources to alleviate overload and improve system performance and resource utilization?

Method used

By constructing a channel model between UAVs and end users, optimizing task hierarchical processing, scheduling UAVs as mobile edge computing nodes, and combining the Lyapunov optimization framework and the genetic-hybrid space multi-agent near-end optimization algorithm, the long-term average total latency of the system is optimized to ensure communication link stability and task processing efficiency.

Benefits of technology

It effectively alleviates the overload problem of edge servers, improves the processing efficiency and system reliability of digital twin tasks, enhances the flexibility and scalability of the system, reduces network congestion, and improves task response speed and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056911B_ABST
    Figure CN121056911B_ABST
Patent Text Reader

Abstract

This application discloses a digital twin task optimization method and system based on an unmanned aerial vehicle (UAV) mobile communication system. The method includes: when the edge server is overloaded, scheduling a UAV as a mobile edge computing node to travel to a designated area and associate with the end user to deploy a digital twin and process tasks; constructing a channel model to obtain the uplink communication rate and uplink transmission delay, determining the computation delay, prioritizing tasks to obtain queuing time, and calculating the long-term average total latency of the system as the optimization objective; minimizing the total latency by solving the model to optimize the task processing process. This invention effectively solves the server overload problem, reduces task latency, and improves the overall system performance, resource utilization, processing efficiency, and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital twin technology, and in particular to a method and system for optimizing digital twin tasks based on unmanned aerial vehicle (UAV) mobile communication systems. Background Technology

[0002] Digital twin technology has broad application prospects in smart cities, but it also faces several key challenges. First, latency needs to be considered to ensure real-time simulation. The construction of a digital twin relies on massive data transmission and computational analysis, requiring a low-latency network to ensure the digital twin model can be updated and synchronized with the physical entity's state in real time. Second, maintenance is another crucial aspect. Maintaining the digital twin model requires sufficient computing and communication resources from the wireless network to ensure continuous updates and operation. Addressing these challenges hinges on the proper deployment of digital twins on servers. A well-planned server deployment provides robust computing power for data processing and model maintenance, fundamentally optimizing network latency and ensuring the smooth implementation and full effectiveness of digital twin technology in smart cities.

[0003] In smart city applications, the overload problem of edge servers is becoming increasingly prominent. With the continuous increase in the number of users and the growing complexity of service requests, the load imbalance of edge servers is becoming more severe. In service hotspots, due to the high density of users and concentrated service demands, edge servers may experience overload, leading to decreased service quality or even service interruptions. To achieve coverage of most areas, a large number of edge servers need to be deployed, which not only increases deployment costs but may also lead to resource waste. Furthermore, during off-peak service periods, the reduced number of service requests lowers server resource utilization, also resulting in resource waste. Therefore, how to rationally allocate and utilize edge server resources while meeting user needs and avoiding overload is one of the major challenges facing the application of digital twin technology in smart cities. To address this issue, drone-assisted computing is a promising technology that has emerged in recent years. Given that future drones will be equipped with more powerful onboard computers, larger-capacity data storage units, and more advanced communication modules, these drones, with their abundant underutilized resources, can be scheduled as general-purpose computing nodes to undertake computing tasks. Specifically, the powerful resources provided by a large number of idle drones can be integrated and utilized to alleviate traffic pressure during peak hours. This approach eliminates the need for additional server deployment, effectively improving the utilization of idle resources and significantly reducing server load and resource consumption. Furthermore, the flexibility and maneuverability of drones enable rapid deployment and adjustment in different scenarios, further enhancing the system's adaptability and reliability. However, digital twins have stringent real-time requirements, and different end-user digital twin models have varying demands for network service quality. Therefore, it is necessary to classify digital twin tasks and, in conjunction with drones, alleviate the overload problem of edge servers. Summary of the Invention

[0004] Based on this, the objective of this invention is to address the aforementioned technical problems by providing a digital twin task optimization method and system based on a drone mobile communication system. This invention, through reasonable resource allocation and hierarchical processing of digital twin tasks, can effectively alleviate the overload problem of edge servers in hotspot areas, thereby improving the overall system performance and resource utilization. To achieve the above-mentioned objective, the first aspect of this application provides a digital twin task optimization method based on a drone mobile communication system, comprising:

[0005] The drone mobile communication system includes drones, several end users, several drones, an edge server, and a base station. When the edge server is overloaded, drones are dispatched as mobile edge computing nodes to a designated area to associate with end users, deploy digital twins of the end users, and process corresponding digital twin tasks. The digital twin task optimization method based on the drone mobile communication system includes the following steps:

[0006] Construct a channel model between the UAV and the end user in the system to establish a correlation, and derive the uplink communication rate from the end user to the UAV from the channel model;

[0007] During the process of transmitting mission data from the end user to the drone, since the downlink transmission delay is much smaller than the uplink transmission delay, the uplink transmission delay from the end user to the drone is calculated only by the uplink communication rate from the end user to the drone.

[0008] Determine the computational latency of the digital twin tasks corresponding to the end users processed by the drone;

[0009] Prioritize the digital twin tasks generated from the digital twin to determine the queuing time for each task while it is waiting to be processed.

[0010] The long-term average total system latency during the entire process of the UAV processing the digital twin task corresponding to the terminal user is determined based on the uplink transmission latency, computation latency, and queuing time.

[0011] The long-term average total latency of the system is used as the optimization objective model, and the optimization objective model is solved to minimize the long-term average total latency of the system in order to optimize the digital twin task.

[0012] Preferably, before a drone is dispatched by a base station to a designated area to associate with an end user to deploy a digital twin and handle corresponding digital twin tasks, the conditions that must be met are that the compensation provided by the base station to the drone and the energy of the drone are both greater than or equal to the total energy consumption generated during the drone's service, i.e., satisfying the following utility model:

[0013]

[0014]

[0015] in, For the utility of drones, The payment that a base station provides to a drone after it accepts a scheduling request from the base station. For the power of drones, For drones Total energy consumption during service:

[0016]

[0017] Among them, drones The average speed and distance traveled towards the target area were respectively and Flight power is , , and These are the average computing power, average communication power, and average hovering power of the drone when providing services. The hovering time is the time during which the drone communicates and computes.

[0018] Preferably, the drone and the end user are associated using the following communication model:

[0019] End users With drones The communication link between them uses line-of-sight (LoS) communication between drones. and end users The channel model between them is:

[0020]

[0021] in, This represents the channel power gain at a reference distance of 1m for the UAV. and end users Distance between for:

[0022]

[0023] Among them, drones The hovering position is , For the hovering altitude of the drone, end user The position is

[0024] End users To drones The uplink communication rate is:

[0025]

[0026] in, End users To drones The mission's transmit power, Represents Gaussian noise power. This represents the channel bandwidth.

[0027] Preferably, the queuing order during the queuing waiting time is determined by the task priority, and the formula for calculating the task priority is:

[0028]

[0029] in, Indicates drone exist Always available computing resources The computing resources required for the task For the maximum acceptable delay, For the task backlog queue, The maximum queue length. The backlog queue impact factor is expressed as follows:

[0030]

[0031] The expression for calculating the length of the task backlog queue is:

[0032]

[0033] in, Throughout the entire system time The amount of task data arriving at each drone in the queue does not exceed the amount of task data that the drone can process, that is:

[0034]

[0035] in, Time to arrive drone The task data volume is drones The amount of task data that can be processed is .

[0036] Preferably, the long-term average total latency of the system is calculated using the following expression:

[0037]

[0038] in, Represents a binary variable, when the end user With drones When associated, ,otherwise, , For end users To drones The uplink transmission delay is calculated using the following expression:

[0039]

[0040] in, For end users Transmitted to drone The amount of task data;

[0041] For drones The computational latency for processing the digital twin task corresponding to the end user is calculated using the following expression:

[0042]

[0043] in, For drones Assigned to end users Computing resources The number of CPU cycles required to process each unit of data;

[0044] This represents the queuing time for a task while it awaits processing, specifically the queuing time for digital twin tasks waiting to be processed in the drone task queue. It is recursively derived using the following formula:

[0045]

[0046] in, Indicates the priority of the task.

[0047] Preferably, the optimization objective is expressed as:

[0048]

[0049] In this context, solving the objective model to minimize the long-run average total delay of the system is formulated as optimization problem P1:

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057] Among them, constraints The requirement is that each end-user's digital twin can only be deployed on a single drone to ensure the exclusivity and efficiency of data processing and interaction; constraints Regulations for deployment on drones The number of digital twins on the device must not exceed its pre-set maximum allowed number. This avoids performance degradation or system failure due to exceeding the load-bearing capacity; constraints Indicates the allocation to the deployment of drones The computing resources of the digital twin on the device must not exceed those of the drone. Maximum computing resources This is to ensure that drones do not experience lag or crashes due to excessive allocation of computing resources when handling tasks; constraints Clearly in drones The total delay must not exceed that of the drone. Longest hovering time To ensure that all necessary data processing and interaction tasks can be completed during drone hovering, maintaining the system's real-time performance and stability; constraints This ensures that the revenue generated by the drone during its service period is greater than or equal to zero; constraint This ensures the stability of the task queue.

[0058] Preferably, the optimization problem P1 is solved by decomposing the long-term constraint problem into a single-slot problem using Lyapunov functions, specifically as follows:

[0059] Define the task backlog queue vector ,Right now:

[0060]

[0061] The Lyapunov function is defined as:

[0062]

[0063] in, If and only if When it is a zero vector, it satisfies ;

[0064] In the optimization problem, a Lyapunov drift term plus a penalty term is introduced, expressed as follows:

[0065]

[0066] in, It is a control parameter that balances the optimization objective function and the stability of the task queue, and ;

[0067] The objective function is re-characterized as:

[0068]

[0069] The upper bound of the conditional expectation of the Lyapunov drift function is expressed as:

[0070]

[0071] Through the The optimization problem is decomposed into minimizing the upper bound plus penalty term of the Lyapunov drift term in each time slot, while also removing the constant term from the upper bound of the Lyapunov drift function. After removing the objective function, the new optimization objective function is obtained as follows:

[0072]

[0073] The optimization problem P1 is ultimately defined as P2:

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081] Based on the Lyapunov optimization framework, the long-term resource-constrained optimization problem is transformed into a deterministic upper bound optimization problem for each time slot.

[0082] Preferably, the entire process of solving the optimization objective model to minimize the long-term average total latency of the system is transformed into a Markov decision process model. The association strategy, transmission power, and computing resource allocation between the UAV and the end user are optimized through the Genetic-Hybrid MAPPO multi-agent proximal optimization algorithm. In the Markov decision process model, each UAV acts as an agent, and the set of state spaces of the UAV is defined as S(t). The state space includes the signal-to-noise ratio at time t and the priority of the task.

[0083] Signal-to-noise ratio and task priority As a state, the signal-to-noise ratio reflects the network's communication status, the user's location information, and the UAV's location information; task priority reflects the UAV's computing resource status and basic information about the digital twin task. Represented as:

[0084]

[0085] To minimize the long-term average total latency of the global system, the agent needs to select the optimal action based on the dynamic environment. Therefore, to maximize the reward, the agent needs to optimize the association strategy between the drone and the user. Transmission power Selection and computing resources The allocation defines the action as:

[0086]

[0087] The reward function is defined as a negative of the current system latency to optimize the long-term average total latency of the system. The reward function is defined as follows:

[0088] .

[0089] Preferably, the genetic-hybrid spatial multi-agent proximal optimization algorithm GA-HybridMAPPO includes:

[0090] Hybrid motion space processing, using an actor-critic network to hierarchically output discrete motion. and continuous action , ;

[0091] Genetic algorithm embedding utilizes binary encoding to generate association strategy chromosomes, and constructs a fitness function to evaluate the quality of chromosomes. The fitness function is defined as follows:

[0092]

[0093] in, Representing chromosomes, i.e., the association strategy;

[0094] The collaborative optimization genetic algorithm (GA) and the hybrid space multi-agent proximal optimization algorithm are used. In the initial stage, each agent triggers the GA optimization mechanism, generating a set of high-quality association policies which are injected into the discrete action space of the agent network for selection. Simultaneously, the continuous action space of the agent network generates continuous actions based on the state. In subsequent stages, while maintaining the initial association policies, the MAPPO policy optimization is performed, focusing on continuous action space decision-making. The specific process is as follows:

[0095] Initialize all conditions to their default values;

[0096] exist At that time, a set of discrete actions is generated using a genetic algorithm (GA) and injected into the discrete space. Then, the best discrete action is selected from the discrete action space. Based on... Prioritize all tasks, and then generate a series of actions using a function in a contiguous space based on the state and task priority queue.

[0097] exist At that time, keep Real-time related decisions, updating task priorities Based on the state and task priority queue, a continuous sequence of actions is generated using functions in a contiguous space.

[0098] The combined actions are derived from the above steps. ,implement Receive a reward and the next state These samples are then stored in the experience pool, and the loss function of the actor network is updated.

[0099] Through multiple iterations, the long-term average total latency of the system is minimized.

[0100] To achieve the purpose of the invention, a second aspect of this application provides a digital twin task optimization system based on an unmanned aerial vehicle (UAV) mobile communication system, applied to the digital twin task optimization method based on an UAV mobile communication system described in the above technical solution. The system includes:

[0101] K drones are used as mobile edge computing nodes to associate with end users when edge servers are overloaded, in order to deploy digital twins of end users and perform digital twin tasks.

[0102] Base stations are used for scheduling and refueling drones;

[0103] The association module is used to construct a channel model between the drone and the end user in the system for association.

[0104] The priority calculation module is used to prioritize the digital twin tasks generated from the digital twin.

[0105] The system latency calculation module is used to calculate the uplink transmission latency from the terminal user to the drone based on the uplink communication rate from the terminal user to the drone, determine the calculation latency of the drone processing the digital twin task corresponding to the terminal user, and determine the queuing time of the digital twin task when waiting for processing by prioritizing the digital twin task.

[0106] The optimization module is used to take the long-term average total latency of the system as the optimization target model, and solve the optimization target model to minimize the long-term average total latency of the system in order to optimize the digital twin task.

[0107] Compared with the prior art, the beneficial effects of this invention are:

[0108] This invention effectively alleviates edge server overload by scheduling drones as mobile edge computing nodes, significantly improving the processing efficiency and system reliability of digital twin tasks. By constructing a channel model between the drone and end users, it accurately calculates uplink communication rate and uplink transmission latency, ensuring the stability of the communication link. Combining the computational latency of digital twin tasks with priority-based queuing time, it optimizes the task processing order, thereby minimizing the system's long-term average total latency, improving overall system performance and resource utilization, accelerating task response speed, and enhancing the end-user experience. Furthermore, this method improves the system's flexibility and scalability, enabling it to adapt to dynamic environmental changes and providing an efficient, low-latency solution for digital twin applications in drone mobile communication systems. Attached Figure Description

[0109] Figure 1 This is a flowchart illustrating the steps of a digital twin task optimization method based on an unmanned aerial vehicle (UAV) mobile communication system in one embodiment.

[0110] Figure 2 This is a schematic diagram of a drone-assisted edge digital twin scenario in one embodiment;

[0111] Figure 3 This is a schematic diagram illustrating the priority of digital twin tasks in one embodiment.

[0112] 1-End user, 2-Edge server, 3-Drone, 4-Base station. Detailed Implementation

[0113] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. The following embodiments are used to illustrate the invention but are not intended to limit its scope.

[0114] Example 1

[0115] Embodiment 1 of this application provides a digital twin task optimization method based on an unmanned aerial vehicle (UAV) mobile communication system, such as... Figure 1 As shown, it includes:

[0116] S1: The UAV mobile communication system includes UAVs, several terminal users, several UAVs, an edge server, and a base station. When the edge server is overloaded, UAVs are dispatched as mobile edge computing nodes to a designated area to associate with terminal users, deploy digital twins of the terminal users, and process corresponding digital twin tasks. The digital twin task optimization method based on the UAV mobile communication system includes the following steps:

[0117] S2: Construct a channel model between the UAV and the end user in the system to establish a correlation, and derive the uplink communication rate from the end user to the UAV from the channel model. During the process of the end user transmitting task data to the UAV, since the downlink transmission delay is much smaller than the uplink transmission delay, the uplink transmission delay from the end user to the UAV is calculated only by the uplink communication rate from the end user to the UAV.

[0118] S3: Determine the computation latency of the digital twin task corresponding to the end user processed by the drone;

[0119] S4: Prioritize the digital twin tasks generated by the digital twin to determine the queuing time for the digital twin tasks while waiting for processing;

[0120] S5: Determine the long-term average total system latency when the UAV processes the digital twin task corresponding to the terminal user throughout the entire process based on the uplink transmission latency, calculation latency, and queuing waiting time;

[0121] S6: Use the long-term average total latency of the system as the optimization objective model, and solve the optimization objective model to minimize the long-term average total latency of the system in order to optimize the digital twin task.

[0122] like Figure 2 As shown, the drone-assisted edge digital twin scenario mainly consists of two parts. The first part involves the association between drone 3 and end user 1. In areas where edge server 2 is overloaded, drone 3, after receiving a scheduling request from base station 4, will fly to the designated area to provide services to end user 1. End user 1 associates with drone 3 according to the actual situation. At this time, drone 3 will assume the role of edge server 2, building a digital twin network for end user 1 (including deploying the end user's digital twin and processing the digital twin tasks it generates) to provide services. The second part is the hierarchical processing of tasks. Since digital twins have extremely high real-time requirements, to avoid network congestion in the wireless digital twin network, which would prevent the digital twin from reflecting the state of the physical entity in a timely manner, tasks are hierarchically arranged according to the elements of digital twin tasks. Drones will prioritize processing higher-priority tasks, thereby effectively reducing the latency of the entire wireless digital twin network.

[0123] In this embodiment, the wireless digital twin system network mainly covers two key parts. First, regarding the association between end users and drones, after the drone receives a scheduling request, it will act as an edge server to assist in the deployment of the digital twin. In this process, it is necessary to focus on the optimal association between drones and end users, as well as resource allocation, striving to achieve the minimum deployment latency to improve system response speed and user experience. Second, given that drones may face network congestion when receiving a large number of processing tasks, it is necessary to classify user-requested tasks. By rationally prioritizing tasks, important and urgent tasks are processed first, effectively solving the network congestion problem caused by processing a large number of tasks during peak hours, thereby reducing the burden on the edge server and ensuring the stable operation and efficient service of the entire network.

[0124] 1. Association Strategy:

[0125] Assuming a drone The average speed and distance traveled towards the target area were respectively and Flight power is When providing services, the average computing power, average communication power, and average hovering power are respectively... , and And the hovering time is It is easy to see that the communication time and computation time are both equal to the hovering time, therefore the drone The total energy consumption during the service period is:

[0126]

[0127] Therefore, when the drone's energy satisfy Only then will a drone be able to accept a request for service. The utility can be expressed as:

[0128]

[0129] in The payment a base station provides to a drone after it accepts a request from the base station. (Definition) ,Right now This means that the reward received by the drone will always be greater than or equal to the energy consumed.

[0130] Time of the system Divided into In a time slot At any time, for terminal devices Its digital twin can be represented as ,in It is the operational data required to run a digital twin. For the behavior model of digital twins, Represents the real-time state of a physical entity. Indicates in time slot middle Status updates. After the digital twin is deployed on the drone platform, it needs to be dynamically synchronized and continuously maintained based on the real-time status data of end users to ensure the real-time mapping accuracy between the digital twin system and the physical entity.

[0131] In real-world scenarios, limited resources are often a challenge. In such situations, the deployment and subsequent update tasks of different digital twins vary in urgency, requiring comprehensive consideration and rational resource allocation. Prioritizing tasks with high urgency is crucial to ensure the efficient progress and successful completion of critical missions. This involves considering terminal devices... In the time slot Triggered digital twin model service requests (including initial model deployment and dynamic state updates) are uniformly defined as digital twin tasks, denoted as... This task represents the critical operational requests throughout the entire lifecycle of the digital twin, and each digital twin task contains four key attributes. ,in Used to distinguish the types of digital twin tasks. This indicates the task of deploying a digital twin. Time indicates the digital twin update task; This represents the data size of the digital twin task; This indicates the computing resources required to perform the digital twin task; This represents the maximum acceptable latency for the digital twin task.

[0132] Upon receiving a dispatch request, the drone will fly to the designated area as required to provide corresponding services to the end user. The first step in this process is establishing a connection between the end user and the drone. Represents a binary variable, when the end user With drones When associated, ,otherwise, To improve communication efficiency between drones and end users, the communication link between them is primarily dominated by the line-of-sight (LoS) channel. In a three-dimensional coordinate system, assuming the drone... The hovering position is ,in For the hovering altitude of the drone, end user The position is From this, we can calculate the drone's... and end users Distance between:

[0133]

[0134] Therefore, drones and end users The channel model between them is:

[0135]

[0136] in This represents the channel power gain at a reference distance of 1m. From this, we can derive the end-user's... To drones The uplink communication rate is:

[0137]

[0138] in End users To drones The mission's transmit power, Represents Gaussian noise power. This represents the channel bandwidth.

[0139] Network latency analysis:

[0140] The service latency of a digital twin system mainly consists of deployment latency and update latency. The deployment phase involves the transmission and loading of the initial model, resulting in a large data volume and thus higher latency; while the update phase only requires synchronizing incremental state data, with lower latency overhead. To establish a unified optimization model, this paper models the entire lifecycle service process of the digital twin as a continuous time-varying task flow and implements a differentiated resource scheduling strategy.

[0141] End users The amount of mission data transmitted to the drone is During data transmission, since the downlink transmission delay is much smaller than the uplink transmission delay, only the uplink transmission delay needs to be considered. Therefore, the end user... To drones The uplink transmission delay is expressed as:

[0142]

[0143] To ensure consistency between the digital twin model and the physical model, the digital twin needs continuous data interaction and information exchange with the physical model. This process is crucial for the dynamic updating of the digital twin model and the real-time control of the physical entity. Simultaneously, the computing resources allocated to users by the server directly affect the update rate of the digital twin. (Drone) The computational latency for processing the digital twin task corresponding to the end user is expressed as follows:

[0144]

[0145] in, For drones Assigned to end users Computing resources The number of CPU cycles required to process each unit of data.

[0146] In a wireless network model, when the end user With drones When correlated, the resulting long-term average total system delay can be expressed as:

[0147]

[0148] in This represents the queuing time for a task while it awaits processing. Assuming the task queuing model is an M / M / 1 model, the average arrival rate of the task is... The average service rate is The average waiting time is The resulting waiting time is the prediction result of queuing theory.

[0149] In practice, once the association strategy between drones and end users is determined, the queuing for digital twin tasks is also determined. The waiting latency of the first digital twin deployment task to be processed in each drone task queue is set to 0. The waiting latency of tasks waiting to be processed in the drone task queue is recursively derived from the following formula:

[0150]

[0151] To describe the long-term stability of the system, a task backlog queue is constructed at the UAV end. Assume that in time slots... In the middle, the drone arrived The workload is In drones The amount of work to be processed is The task backlog queue is represented as The length of the task backlog queue is:

[0152]

[0153] Where the definition To prevent the task queue from growing indefinitely and to ensure the relative stability of the entire system, it is necessary to ensure that the amount of task data arriving at each UAV pair within the system time does not exceed the amount of task data processed.

[0154]

[0155] Then build drones Task prioritization model, for Analyzing the three original attributes will consider several factors that significantly impact task priority. For multiple digital twin tasks on the same UAV in the same time slot, execution will follow these criteria: First, digital twin tasks with lower maximum acceptable latency should be prioritized; second, digital twin tasks requiring more computing resources will have lower processing priority; finally, the longer the task backlog queue, the lower the execution priority of newly arriving digital twin tasks. Therefore, the end user... The generated tasks are in drones The priority of the above processing can be expressed as:

[0156]

[0157] in, Indicates drone exist Always available computing resources; The backlog impact factor is expressed as follows:

[0158]

[0159] in, Indicates drone The maximum task backlog queue length is related to the drone's computing power.

[0160] like Figure 3 As shown, for Time slots, in drones There are multiple tasks waiting to be processed in the task queue. The priority of the tasks will be determined according to the following two principles: (1) Large tasks have a high priority and will be processed first; (2) For tasks in different time slots, a first-come-first-served approach will be adopted, and the tasks in the previous time slot will be processed first before the tasks in the current time slot are processed.

[0161] 2. Optimize the target model

[0162] Once the conditions are met, the drone arrives at the designated area and establishes a connection with the end user. At this point, the drone will act as an edge server, providing various services to the end user. The digital twin generated by the end user will be deployed on the drone, which will also handle the interaction data between the end user and the digital twin. Based on the system characteristics, the long-term average total latency of the system is used as the optimization objective, expressed as:

[0163]

[0164] In this context, the optimization problem can be formulated as:

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] Among them, constraints The requirement is that each end-user's digital twin can only be deployed on a single drone to ensure the exclusivity and efficiency of data processing and interaction; constraints Regulations for deployment on drones The number of digital twins on the device must not exceed its pre-set maximum allowed number. This avoids performance degradation or system failure due to exceeding the load-bearing capacity; constraints Indicates the allocation to the deployment of drones The computing resources of the digital twin on the device must not exceed those of the drone. Maximum computing resources This is to ensure that drones do not experience lag or crashes due to excessive allocation of computing resources when handling tasks; constraints Clearly in drones The total delay must not exceed that of the drone. Longest hovering time This is to ensure that all necessary data processing and interaction tasks can be completed during drone hovering, maintaining the system's real-time performance and stability. Constraints This ensures that the revenue generated by the drone during its service period is greater than or equal to zero; constraint This ensures the stability of the task queue.

[0173] Lyapunov optimization framework:

[0174] Optimization problem It is difficult to solve directly, mainly due to the following challenges:

[0175] (1) Long-term computational resource constraints make the solution difficult. Under resource constraints, the computational resource allocation decision of UAVs has significant cross-time slot dependence. Therefore, it is necessary to take into account the resource allocation strategies of different time slots in order to optimize system performance.

[0176] (2) To achieve the optimal solution, it is necessary to obtain global environmental status information in real time, especially the dynamically evolving UAV task queue and load status, but this is challenging to achieve.

[0177] (3) As the number of end users and drones increases, the complexity of the solution will also increase exponentially. Even if the global environment state is known, it is difficult to solve the above problems.

[0178] To address the dynamic optimization challenges faced by UAVs under long-term resource constraints, this embodiment employs a Lyapunov optimization framework to solve the optimization problem. The core function of this framework is to decompose the complex time-varying optimization problem with long-term resource constraints into a series of deterministic single-slot subproblems based on instantaneous system states and easily solvable online using Lyapunov functions.

[0179] First, define the task backlog queue vector. ,Right now The Lyapunov function can be defined as:

[0180]

[0181] From the definition of the Lyapunov function, we know that If and only if When it is a zero vector, it satisfies Furthermore, the Lyapunov drift function is defined as follows:

[0182]

[0183] Lemma 1: For Lyapunov functions ,when At that time, there exists a constant. and This allows the upper bound of the Lyapunov drift to be expressed as:

[0184]

[0185] Proof: For In terms of the definition:

[0186] 1. When ,but ,at this time , ;

[0187] 2. When ,at this time ,therefore , .

[0188] From the above, we can see that:

[0189]

[0190] For all task backlog queues, we know that:

[0191]

[0192]

[0193]

[0194] By moving the terms of the inequality above and taking the expectation of both sides, we get:

[0195]

[0196]

[0197]

[0198] in Because the number of tasks arriving and the amount of service provided by each queue are bounded, and Therefore, there exists a constant. This ensures that all task queues satisfy the requirements. Therefore, the upper bound of the conditional expectation of the Lyapunov drift function is expressed as:

[0199]

[0200] Will Defined as Then the upper bound of the conditional expectation of the Lyapunov drift function can be transformed into:

[0201]

[0202] Lemma 1 is proved.

[0203] In the Lyapunov optimization framework, by introducing a penalty term, bi-objective optimization is achieved while ensuring system stability. The expression for the Lyapunov drift term plus the penalty term can be obtained as follows:

[0204]

[0205] in It is a control parameter that balances the optimization objective function and the stability of the task queue, and The objective function is redefined as follows:

[0206]

[0207] However, substituting the above objective function into the optimization problem still presents difficulties, requiring knowledge of current and future global information. Therefore, further simplification of the objective function is necessary. This can be achieved through... The optimization problem is decomposed into minimizing the upper bound plus a penalty term of the Lyapunov drift term in each time slot. Simultaneously, the constant term in the upper bound of the Lyapunov drift function is... After removing the objective function, the new optimization objective function is obtained as follows:

[0208]

[0209] The optimization problem is ultimately defined as:

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217] Based on the Lyapunov optimization framework, the long-term resource-constrained optimization problem is transformed into a deterministic upper bound optimization problem for each time slot to ensure the stability of the task queue and further search for the optimal solution.

[0218] 3. Resource allocation method based on the Genetic-HybridMAPPO multi-agent proximal optimization algorithm:

[0219] Multi-Agent Proximal Policy Optimization (MAPPO) Algorithm. In wireless digital twin networks, the association and resource allocation between UAVs and end users are crucial. To efficiently address this issue, the MAPPO algorithm is introduced. This algorithm aims to minimize system latency by solving the association and resource allocation strategies between UAVs and end users. Specifically, the UAV selects the optimal computational resource allocation strategy based on the priority of the digital twin tasks, ensuring that critical tasks can be processed in a timely manner under limited resources. This not only effectively reduces network congestion but also alleviates the overload of edge servers, thereby significantly improving the overall system efficiency and quality of service.

[0220] Genetic Algorithm (GA). The MAPPO algorithm framework processes hybrid action spaces through hierarchical policy networks, but it still faces the problem of inefficient exploration due to combinatorial explosion in high-dimensional discrete decision-making (UAV-user association). To solve the problem of inefficient exploration in high-dimensional discrete action spaces, this invention innovatively embeds a genetic algorithm into the MAPPO training framework. This mechanism utilizes the powerful global search capability of GA to generate a set of association policies that meet the constraints and injects them into the discrete action space of MAPPO. Specifically, GA represents the association relationship between the UAV and the end user through binary encoding, ensuring that the association constraints are met. The fitness function, combined with the critic network in MAPPO, evaluates the merits of the policies and provides optimization direction for GA. Through this collaborative approach, GA and MAPPO complement each other: GA efficiently explores the global optimum in the discrete action space, while MAPPO performs fine-grained resource allocation optimization in the continuous action space.

[0221] (1) MDP Model Construction

[0222] For the deployment of digital twins on drones, assuming each drone acts as an intelligent agent, the state space set of the drones is defined as follows: The state space contains the signal-to-noise ratio and task priority at each time step. The elements of the state space completely describe the information at a given time step, indicating that the future state of the UAV depends only on the current state and is independent of past states, satisfying the Markov property. Therefore, the problem is transformed into a Markov decision process model. In constructing the model, the state space, action space, and reward function are defined as follows:

[0223] 1. Status: For formulaic digital twin relationships and resource allocation problems, the environmental status needs to fully reflect all key information of the drone and end user in any time slot. Signal-to-noise ratio... and task priority As a state, the signal-to-noise ratio (SNR) reflects the network's communication status, user location information, and UAV location information; task priority reflects the UAV's computing resource status and basic information about the digital twin task. It can be represented as:

[0224]

[0225] 2. Actions: To minimize the long-term average total latency of the global system, the agent needs to select the optimal action based on the dynamic environment. Therefore, to maximize rewards, the agent needs to optimize the association strategy between the drone and the user. Transmission power Selection and computing resources The allocation. Actions can be defined as:

[0226]

[0227] 3. Reward Function: In the optimization objective model, the goal is to minimize the system's long-run average total latency. Therefore, the reward function is defined as the negative of the current system latency to optimize the system's latency. The reward function is defined as follows:

[0228]

[0229] (2) MAPPO algorithm with hybrid action space (HybridMAPPO)

[0230] In the constructed MDP model, the given state space and action space are high-dimensional, and the action space is a mixture of discrete and continuous actions. The relationship between drones and end users. It belongs to discrete actions, and the transmission power Selection and computing resources The assignment belongs to continuous actions. To solve these problems, this invention uses a hierarchical actor neural network to solve the mixed action space problem, and combines an improved actor neural network with reinforcement learning to solve the constructed MDP problem. In the actor neural network, consider... Network parameters as discrete layers Network parameters for consecutive layers. The parameters of the speaker network are defined as follows: During optimization, the improved MAPPO algorithm will seek the optimal network parameters to maximize the long-term reward. The long-term reward can be expressed as:

[0231] [ ]

[0232] In the MAPPO architecture, multiple agents operate using the PPO algorithm based on policy gradients and confidence intervals. The actor network encompasses both discrete and continuous spaces, acting as the decision network for MAPPO with mixed policies within the same environment state. For discrete action decisions, each agent receives a state and then generates a stochastic policy using a softmax function. For decision-making involving consecutive actions, its stochastic strategy... It follows a Gaussian distribution, which is generated by the mean of the output of the sigmoid function and the standard deviation of the output of the softplus function.

[0233] In MAPPO with a hybrid strategy, the agent's operation process includes two phases: centralized training and distributed execution. For the agent of a drone, assuming... As a network parameter for critics, the critic network is based on... and Calculate the action value function:

[0234]

[0235] The optimization goal of the critic network is to make its estimate as close as possible to the actual advantage estimate. To achieve this goal, this gap is measured and optimized by minimizing a loss function:

[0236]

[0237] in The advantage function, representing the advantage of the current decision compared to other feasible decisions, is expressed as:

[0238]

[0239] in, and These represent the discount factor and the single-step time error, respectively. Represented as:

[0240]

[0241] During the distributed execution phase, the actor network updates both the continuous and discrete space parameters, and then, based on the centralized calculations by the critics... The function results and local observation output actions. The loss function in the discrete action space of the actor network is defined as:

[0242]

[0243] in, Represents parameters in the discrete action space The scope of the update; Represents a confidence interval The parameters, and ; Clipping function Will The update range is limited to the confidence interval to ensure network stability. Similarly, the loss function for the continuous action space in the actor network is defined as:

[0244] .

[0245] (3) Genetic Algorithm

[0246] Although the HybridMAPPO framework handles the hybrid action space through a hierarchical policy network, it still faces the problem of inefficient exploration due to combinatorial explosion in high-dimensional discrete decision-making (UAV-user association). To address the challenge of inefficient exploration in high-dimensional discrete action spaces, this paper innovatively embeds a genetic algorithm into the MAPPO training framework. This mechanism utilizes the global search capability of GA to generate a set of association policies that meet the constraints and injects them into the discrete action space of the actor network, forming an effective complementarity between genetic algorithms and reinforcement learning.

[0247] First, we need to define chromosome encoding. In the discrete phase, the association between the drone and the end user is a binary decision; therefore, the chromosome uses binary encoding, with each gene bit representing an association decision. At the same time, it is necessary to ensure that the constraints are met during coding. .

[0248] The fitness function of a genetic algorithm is used to evaluate the quality of a chromosome. To guide the GA search for high-value strategies, for a given chromosome, it is combined with the current state and input into the critic network. The resulting value estimate is then converted into a positive value using the Softplus function, which is used as the fitness value. Therefore, the fitness function is defined as the result of applying the Softplus function to the critic network's value estimate:

[0249]

[0250] in This refers to chromosomes, i.e., association strategies.

[0251] Based on the initial population, the fitness value calculated using the fitness function yields the probability of each individual being selected. A roulette wheel selection method is then used to choose a subset of individuals as superior ones. Simultaneously, bit-flip mutation is employed, with the crossover probability set between 0.6 and 0.95 and the mutation probability between 0.001 and 0.05. Through crossover and mutation, new offspring are generated. These newly generated offspring, along with their fitness values, are incorporated into the original population, forming a new population. This process is then iteratively repeated to optimize the entire population.

[0252] Example 2

[0253] Embodiment 2 of this application, based on Embodiment 1, further illustrates the Genetic-HybridMAPPO multi-agent proximal optimization algorithm, as follows:

[0254] The GA-HybridMAPPO framework achieves collaborative optimization of genetic algorithms and the hybrid MAPPO strategy through a two-stage mechanism. During continuous action decision-making, resource allocation needs to consider the state of the task priority queue. Throughout the algorithm's execution, each stage triggers the GA optimization mechanism, generating a set of high-quality association strategies that are injected into the discrete action space of the actor network for selection. Simultaneously, the actor network's continuous action space generates continuous actions based on the state. In subsequent stages, the initial association strategy is maintained, and MAPPO strategy optimization is performed, focusing on continuous action space decision-making.

[0255] The specific process is as follows:

[0256] 1) All conditions in the system are set to default values;

[0257] 2) In Time: Use GA to generate a set of discrete actions and inject it into the discrete space, then select the best discrete action from the discrete action space; according to Prioritize all tasks, and then generate a series of actions using a function in a contiguous space based on the state and task priority queue.

[0258] 3) In At that time, keep Real-time related decisions, updating task priorities Based on the state and task priority queue, a continuous sequence of actions is generated using functions in a contiguous space.

[0259] 4) Based on the combined actions obtained above ,Will Perform the task and receive a reward. and the next state These samples are then stored in the experience pool, and the loss function of the actor network is updated.

[0260] 5) Minimize the long-term average total latency of the system through multiple iterations.

[0261] Example 3

[0262] This embodiment 3, based on embodiments 1 and 2, provides a digital twin task optimization system based on an unmanned aerial vehicle (UAV) mobile communication system, used in the digital twin task optimization method based on an UAV mobile communication system described in embodiments 1 and 2. The system includes:

[0263] K drones are used as mobile edge computing nodes to associate with end users when edge servers are overloaded, in order to deploy digital twins of end users and perform digital twin tasks.

[0264] Base stations are used for scheduling and refueling drones;

[0265] The association module is used to construct a channel model between the drone and the end user in the system for association.

[0266] The priority calculation module is used to prioritize the digital twin tasks generated from the digital twin.

[0267] The system latency calculation module is used to calculate the uplink transmission latency from the terminal user to the drone based on the uplink communication rate from the terminal user to the drone, determine the calculation latency of the drone processing the digital twin task corresponding to the terminal user, and determine the queuing time of the digital twin task when waiting for processing by prioritizing the digital twin task.

[0268] The optimization module is used to take the long-term average total latency of the system as the optimization target model, and solve the optimization target model to minimize the long-term average total latency of the system in order to optimize the digital twin task.

[0269] In each time slot Internally, the system is set by The system consists of a drone and a base station. The association module, priority calculation module, system latency calculation module, and optimization module are generally integrated within the drone. The system also includes... Each end user (such as devices and mobile devices in the Internet of Things) and edge servers. When the edge servers are overloaded, the drone, under the condition of satisfying the utility model, will accept the scheduling request of the base station, go to the designated area to associate with the end user, provide as a mobile edge computing node to deploy the digital twin of the end user and process the digital twin tasks generated by the digital twin.

[0270] In summary, this invention provides a method and system for optimizing digital twin tasks based on a drone mobile communication system. This invention leverages the flexibility of drones, sending scheduling requests to them via base stations and using them as mobile edge service computing nodes. By establishing a connection between end users and drones, the GA-HybridMAPPO algorithm is used to solve for the optimal connection between end users and drones and the optimal resource allocation strategy for drones. The drones will assist edge servers in deploying digital twins generated by end users, thereby effectively alleviating the overload problem of edge servers. To meet the stringent real-time requirements of digital twins, the digital twin tasks generated by end users are processed in a hierarchical manner. In the specific implementation process, based on three major influencing factors—the data size of the digital twin task, the computing resources required to complete the task, and the maximum allowable latency—and by introducing a backlog queue influence factor, tasks are ranked and arranged in a hierarchical manner. Drones will prioritize processing higher-priority tasks, thereby effectively reducing the latency of the entire wireless digital twin network and avoiding network congestion problems that could prevent the digital twin from reflecting the state of the physical entity in a timely manner.

Claims

1. A method for digital twin task optimization based on a UAV mobile communication system, characterized in that, The drone mobile communication system includes drones, several end users, several drones, an edge server, and a base station. When the edge server is overloaded, drones are dispatched as mobile edge computing nodes to a designated area to associate with end users, deploy digital twins of the end users, and process corresponding digital twin tasks. The digital twin task optimization method based on the drone mobile communication system includes the following steps: A channel model is constructed between the UAV and the end user in the system to establish a correlation. The uplink communication rate from the end user to the UAV is obtained from the channel model. During the process of the end user transmitting task data to the UAV, since the downlink transmission delay is much smaller than the uplink transmission delay, the uplink transmission delay from the end user to the UAV is calculated only by the uplink communication rate from the end user to the UAV. Determine the computational latency of the digital twin tasks corresponding to the end users processed by the drone; Prioritize the digital twin tasks generated from the digital twin to determine the queuing time for each task while it is waiting to be processed. The long-term average total system latency during the entire process of the UAV processing the digital twin task corresponding to the terminal user is determined based on the uplink transmission latency, computation latency, and queuing time. The long-term average total latency of the system is used as the optimization objective model, and the optimization objective model is solved to minimize the long-term average total latency of the system in order to optimize the digital twin task.

2. The method of claim 1, wherein, Before a drone can be dispatched by a base station to a designated area to associate with an end user, deploy a digital twin, and handle corresponding digital twin tasks, it must meet the following conditions: the compensation provided by the base station to the drone and the energy of the drone must both be greater than or equal to the total energy consumption generated during the drone's service, i.e., satisfying the following utility model: wherein, utility of the UAV, reward provided by the base station to the UAV after the base station accepts the scheduling request of the UAV, energy of the UAV, the UAV total energy consumption generated during the service period: in, Indicates drone The average speed of flight to the target area This indicates the current flight distance of the drone as it heads towards the target area. Indicates the flight power of the drone. This represents the average computing power of the drone while providing services. This indicates the average communication power of the drone while providing services. This indicates the average hovering power of the drone while providing services. The hovering time is the time during which the drone communicates and computes.

3. The method of claim 1, wherein, The drone and the end user are associated using the following communication model: End users With drones The communication link between them uses line-of-sight (LoS) communication between drones. and end users Channel model between for: wherein, represents the channel power gain at a 1 m reference distance, the unmanned aerial vehicle and the end user distance is: wherein the hovering position of the drone is , , the hovering height of the drone, the position of the end user is ​ End user to the drone uplink communication rate is: wherein is a terminal user to a drone mission transmit power, represents a Gaussian noise power, represents a channel bandwidth.

4. The method of claim 3, wherein, The queuing order during the waiting time is determined by the task priority, and the formula for calculating the task priority is: in, Indicates drone exist Always available computing resources The computing resources required for the task. For the maximum acceptable delay, For the task backlog queue, The maximum queue length. The backlog queue impact factor is expressed as follows: The expression for calculating the length of the task backlog queue is: wherein, the amount of data of the queue task that each unmanned aerial vehicle reaches in the whole system time does not exceed the amount of data of the task that the unmanned aerial vehicle can process, that is, the amount of data of the queue task that each unmanned aerial vehicle reaches in the whole system time does not exceed the amount of data of the task that the unmanned aerial vehicle can process, that is, wherein, the time instant at which the drone arrives the amount of task data for the drone the amount of task data that the drone is able to process .

5. The method of claim 4, wherein, The formula for calculating the long-term average total delay of the system is as follows: in, Represents a binary variable, when the end user... With drones When associated, ,otherwise, , For end users To drones The uplink transmission delay is calculated using the following expression: in, For end users Transmitted to drone The amount of task data; For drones The computational latency for processing the digital twin task corresponding to the end user is calculated using the following expression: in, For drones Assigned to end users Computing resources The number of CPU cycles required to process each unit of data; This represents the queuing time for a task while it awaits processing, specifically the queuing time for digital twin tasks waiting to be processed in the drone task queue. It is recursively derived using the following formula: in, Indicates the priority of the task.

6. The method according to any one of claims 1-5, characterized in that, The optimization objective model is expressed as: In this context, solving the objective model to minimize the long-run average total delay of the system is formulated as optimization problem P1: in, For drones Total energy consumption generated during service; constraints The requirement is that each end-user's digital twin can only be deployed on a single drone to ensure the exclusivity and efficiency of data processing and interaction; constraints Regulations for deployment on drones The number of digital twins on the device must not exceed its pre-set maximum allowed number. This avoids performance degradation or system failure due to exceeding the load-bearing capacity; constraints Indicates the allocation to the deployment of drones The computing resources of the digital twin on the device must not exceed those of the drone. Maximum computing resources This is to ensure that drones do not experience lag or crashes due to excessive allocation of computing resources when handling tasks; constraints Clearly in drones The total delay must not exceed that of the drone. Longest hovering time To ensure that all necessary data processing and interaction tasks can be completed during drone hovering, maintaining the system's real-time performance and stability; constraints This ensures that the revenue generated by the drone during its service period is greater than or equal to zero; constraint This ensures the stability of the task queue.

7. The method according to claim 6, characterized in that, The optimization problem P1 is solved by decomposing the long-term constraint problem into a single-slot problem using Lyapunov functions, specifically: Define the task backlog queue vector ,Right now: The Lyapunov function is defined as: in, If and only if When it is a zero vector, it satisfies ; In the optimization problem, a Lyapunov drift term plus a penalty term is introduced, expressed as follows: in, It is a control parameter that balances the optimization objective function and the stability of the task queue, and ; The objective function is re-characterized as follows: The upper bound of the conditional expectation of the Lyapunov drift function is expressed as: Through the The optimization problem is decomposed into minimizing the upper bound plus penalty term of the Lyapunov drift term in each time slot, while also removing the constant term from the upper bound of the Lyapunov drift function. After removing the old function, the new optimization objective function is obtained as follows: The optimization problem P1 is ultimately defined as P2: Based on the Lyapunov optimization framework, the long-term resource-constrained optimization problem is transformed into a deterministic upper bound optimization problem for each time slot.

8. The method according to claim 7, characterized in that, The process of solving the optimization objective model to minimize the long-term average total latency of the system is transformed into a Markov decision process model. The association strategy between the UAV and the end user, the transmission power, and the allocation of computing resources are optimized through the Genetic-Hybrid MAPPO algorithm. In the Markov decision process model, each UAV acts as an agent. The set of state spaces of the UAV is defined as S(t), which includes the signal-to-noise ratio at time t and the priority of the task. Signal-to-noise ratio and task priority As a state, the signal-to-noise ratio reflects the network's communication status, the user's location information, and the UAV's location information; task priority reflects the UAV's computing resource status and basic information about the digital twin task. Represented as: To minimize the long-term average total latency of the global system, the agent needs to select the optimal action based on the dynamic environment. Therefore, to maximize the reward, the agent needs to optimize the association strategy between the drone and the user. Transmission power Selection and computing resources The allocation defines the action as: The reward function is defined as a negative of the current system latency to optimize the long-term average total latency of the system. The reward function is defined as follows: 。 9. The method according to claim 8, characterized in that, The genetic-hybrid spatial multi-agent proximal optimization algorithm GA-HybridMAPPO includes: Hybrid motion space processing, using an actor-critic network to hierarchically output discrete motion. and continuous action , ; Genetic algorithm embedding utilizes binary encoding to generate association strategy chromosomes, and constructs a fitness function to evaluate the quality of chromosomes. The fitness function is defined as follows: in, Representing chromosomes, i.e., the association strategy; The collaborative optimization genetic algorithm (GA) and the hybrid space multi-agent proximal optimization algorithm are used. In the initial stage, each agent triggers the GA optimization mechanism, generating a set of high-quality association policies which are injected into the discrete action space of the agent network for selection. Simultaneously, the continuous action space of the agent network generates continuous actions based on the state. In subsequent stages, while maintaining the initial association policies, the MAPPO policy optimization is performed, focusing on continuous action space decision-making. The specific process is as follows: Initialize all conditions to their default values; exist At that time, a set of discrete actions is generated using a genetic algorithm (GA) and injected into the discrete space. Then, the best discrete action is selected from the discrete action space. Based on... Prioritize all tasks, and then generate a series of actions using a function in a contiguous space based on the state and task priority queue. exist At that time, keep Real-time related decisions, updating task priorities Based on the state and task priority queue, a continuous sequence of actions is generated using functions in a contiguous space. The combined actions are derived from the above steps. ,implement Receive a reward and the next state These samples are then stored in the experience pool, and the loss function of the actor network is updated. Through multiple iterations, the long-term average total latency of the system is minimized.

10. A digital twin task optimization system based on an unmanned aerial vehicle (UAV) mobile communication system, applied to the digital twin task optimization method based on an UAV mobile communication system as described in any one of claims 1-9, characterized in that, The system includes: K drones are used as mobile edge computing nodes to associate with end users when edge servers are overloaded, in order to deploy digital twins of end users and perform digital twin tasks. Base stations are used for scheduling and refueling drones; The association module is used to construct a channel model between the drone and the end user in the system for association. The priority calculation module is used to prioritize the digital twin tasks generated from the digital twin. The system latency calculation module is used to calculate the uplink transmission latency from the terminal user to the drone based on the uplink communication rate from the terminal user to the drone, determine the calculation latency of the drone processing the digital twin task corresponding to the terminal user, and determine the queuing time of the digital twin task when waiting for processing by prioritizing the digital twin task. The optimization module is used to take the long-term average total latency of the system as the optimization target model, and solve the optimization target model to minimize the long-term average total latency of the system in order to optimize the digital twin task.