Multi-uav cooperative auxiliary task unloading and cache double-time-scale optimization method

By employing a dual-timescale optimization method for multi-drone collaborative assisted task offloading and caching, combined with DQN and DTD3 algorithms, the problems of service resource dependence and cross-timescale in drone-assisted edge computing are solved, minimizing system energy consumption and latency, and improving user experience.

CN121541943BActive Publication Date: 2026-04-21JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-01-20
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing drone-assisted edge computing does not fully consider the dependence of computing tasks on service resources, ignores the cross-timescale issues of edge service cache updates and drone flight decisions, and has a single optimization objective, making it difficult to achieve a balance between latency and energy consumption, resulting in poor adaptability.

Method used

A dual-timescale optimization method for multi-UAV collaborative auxiliary task offloading and caching is adopted. By constructing optimization objective functions for the first and second time slots, the UAV service caching and flight trajectory decision are optimized respectively. Combined with DQN and the Diffused Dual Delay Deep Deterministic Policy Gradient Algorithm (DTD3) for joint optimization, the system energy consumption and latency are minimized.

Benefits of technology

The system achieves joint optimization of drone service caching, trajectory, computing resource allocation, and offloading decisions across two time scales, reducing system energy consumption and latency, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541943B_ABST
    Figure CN121541943B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of wireless network communication technology and discloses a dual-timescale optimization method for multi-UAV collaborative auxiliary task offloading and caching. The method includes: setting the update time slot for UAV service cache decision as a first time slot, and the update time slots for UAV flight trajectory decision, computing resource allocation decision, and user offloading decision as a second time slot; wherein the first time slot is composed of multiple second time slots; constructing a first optimization objective function and a second optimization objective function; in each first time slot, obtaining the user task's requirements for various service resources, and updating the UAV service cache decision according to the first optimization objective function; in each second time slot, obtaining the UAV location, user task, and UAV service cache, and updating the UAV flight trajectory decision, computing resource allocation decision, and user offloading decision according to the second optimization objective function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless network communication technology, and specifically relates to a dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching. Background Technology

[0002] With the rapid development of IoT and next-generation communication technologies, various mobile applications have experienced explosive growth, leading to a surge in user demand for computing resources. However, mobile devices have limited computing resources, making it difficult to meet users' requirements for high computing power and low latency. Against this backdrop, mobile edge computing is considered a promising solution, allowing users to offload computing tasks to edge servers with more computing resources, thereby achieving lower task latency and meeting the computing needs of edge users. However, traditional base station-style edge servers have high hardware installation costs and are difficult to install in some environments. Using drones to carry edge servers effectively solves this problem by leveraging the flexibility and low cost of drones. However, many challenges remain in drone-assisted edge computing.

[0003] Existing research often fails to consider the dependence of real-world computing tasks on service resources. For example, tasks require additional software libraries or data resources, and the cross-timescale issues inherent in real-time decision-making, such as edge service cache updates and UAV flight decisions, are often overlooked. Furthermore, existing research primarily focuses on single or multiple UAVs independently assisting edge computing, neglecting collaboration between them. The optimization objectives in current research are often single-objective optimizations focusing solely on latency or energy consumption, without simultaneously considering both. This is crucial for user devices with limited energy and high latency requirements. Additionally, existing solutions to optimization problems often employ convex optimization and game theory, which present challenges in terms of computational complexity and poor adaptability to dynamic environments. Summary of the Invention

[0004] The purpose of this invention is to provide a dual-timescale optimization method for multi-UAV collaborative auxiliary task unloading and caching. It optimizes the user's computational task's dependence on service caching and the inconsistency of the timescale of optimization variables. It can jointly optimize UAV service caching, UAV trajectory, computational resource allocation, and unloading decisions on both timescales, thereby minimizing system energy consumption and latency and improving user experience.

[0005] The technical solution provided by this invention is as follows:

[0006] A dual-time-scale optimization method for multi-UAV collaborative assisted task unloading and caching includes:

[0007] The update time slot for drone service cache decision is set as the first time slot, and the update time slots for drone flight trajectory decision, computing resource allocation decision and user unloading decision are set as the second time slot;

[0008] The first time slot is composed of multiple second time slots;

[0009] Construct the first and second optimization objective functions;

[0010] In each first time slot, the user task's demand for various service resources is obtained, and the service cache decision for the UAV is updated according to the first optimization objective function;

[0011] In each second time slot, the drone location, user task, and drone service cache are acquired, and the drone flight trajectory decision, computing resource allocation decision, and user offload decision are updated according to the second optimization objective function.

[0012] The first optimization objective function is:

[0013] ;

[0014] The second optimization objective function is:

[0015] ;

[0016] In the formula, Representing drones, Indicates the total number of drones; Representative service resource types, This indicates the total number of service resource types; Indicates the first time slot. This represents the total number of the first time slots; Indicates drone Cache service resources within the communication coverage area Hit rate in the first time slot, Indicates the second time slot. The weighting factor for time delay is... , The second time slot Total latency and total energy consumption of the internal system This indicates the total number of second time slots. These represent user uninstallation decisions, drone caching decisions, drone flight trajectory decisions, and computing resource allocation decisions, respectively.

[0017] Preferably, drones Cache service resources within the communication coverage area The formula for calculating the hit rate in the first time slot is:

[0018] ;

[0019] in, Indicates the drone in the second time slot Cache service resources logical value, Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks, Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks; The number of small time slots contained in each large time slot.

[0020] Preferably, the formula for calculating the total system delay within the second time slot is:

[0021]

[0022] in:

[0023] ;

[0024] In the formula, For users to select the latency for associated drone unloading calculations, This indicates the logical value that the user selected to uninstall the associated drone; Allow users to select the local computation latency. This indicates that the user has selected a locally calculated logical value; Users can select the latency for assisting drone unloading calculations. This indicates the logical value representing the user's choice to assist with drone unloading; Indicates user, This indicates the total number of users. This indicates the second time slot.

[0025] Preferably, the formula for calculating the latency of the user-selected drone unloading computation is:

[0026] ;

[0027] In the formula, This indicates the transmission latency from the user to the associated drone. This indicates the transmission latency from the associated drone to the collaborating drone. This indicates the computational latency of the collaborative drone.

[0028] Preferably, the formula for calculating the total energy consumption of the system within the second time slot is:

[0029] ;

[0030] in, This indicates the energy consumption for user task transmission and computation. On behalf of users, This indicates the total number of users. This indicates the energy consumption of a drone during flight, representing the drone's... This indicates the total number of drones.

[0031] Preferably, in the first time slot, the service caching decision for the UAV is determined by the DQN model;

[0032] The reward function during training of the DQN model is described above. Set to:

[0033] ;

[0034] in, Indicates drone Cache service resources within the communication coverage area Hit rate in the first time slot, Representing drones, Representative service resource types, This indicates the total number of service resource types.

[0035] Preferably, the dual-timescale optimization method for multi-UAV collaborative assisted task unloading and caching further includes:

[0036] Construct a dual-delay deep deterministic policy gradient algorithm model, which includes an action network and a comment network;

[0037] The action network of the dual-delay deep deterministic policy gradient algorithm model is replaced with a diffusion model to obtain the diffusion dual-delay deep deterministic policy gradient algorithm model.

[0038] Using UAV location, user task, and UAV service cache as state variables, and human-machine flight trajectory decision, computing resource allocation decision, and user unloading decision as action variables, the diffusion dual-delay deep deterministic policy gradient algorithm model is trained to obtain the optimal diffusion dual-delay deep deterministic policy gradient algorithm model.

[0039] Among them, the reward function for model training Set to:

[0040] ;

[0041] In the formula, Indicates the second time slot. The weighting factor for time delay is... , The second time slot Total latency and total energy consumption of the internal system;

[0042] In the second time slot, the drone's location, user tasks, and drone service cache are acquired, and the drone flight trajectory decision, computing resource allocation decision, and user offload decision are obtained through the optimal diffusion dual-delay deep deterministic strategy gradient algorithm model.

[0043] Preferably, the drone's service cache decision is updated in the first second time slot within each first time slot.

[0044] The beneficial effects of this invention are:

[0045] The present invention provides a dual-timescale optimization method for multi-UAV collaborative auxiliary task unloading and caching. This method addresses the dependency of user computing tasks on service caching and the inconsistency of optimization variable timescales. It can jointly optimize UAV service caching, UAV trajectory, computing resource allocation, and unloading decisions on both timescales, thereby minimizing system energy consumption and latency and improving user experience.

[0046] The present invention provides a dual-timescale optimization method for unloading and caching of multi-UAV collaborative auxiliary tasks. It uses the Diffusion Dual-Delay Deep Deterministic Policy Gradient (DTD3) algorithm to optimize UAV trajectory, computing resource allocation, and unloading decision. The diffusion model is used as the policy network of the DTD3 algorithm, which enhances the action expression ability of the model and improves the quality of policy decision. Attached Figure Description

[0047] Figure 1 This is a system model diagram of the multi-UAV collaborative auxiliary task unloading described in this invention.

[0048] Figure 2 This is a flowchart of the dual-timescale optimization method for multi-UAV collaborative auxiliary task unloading and caching described in this invention.

[0049] Figure 3 This is a diagram of the dual-timescale optimization model described in this invention.

[0050] Figure 4 This is a diagram of the DTD3 algorithm model architecture described in this invention.

[0051] Figure 5 This is a graph showing the system consumption of each algorithm in the experimental examples of this invention as a function of the UAV's cache capacity.

[0052] Figure 6 This is a graph showing the system consumption of each algorithm in the experimental examples of this invention as a function of the number of drones.

[0053] Figure 7 This is a graph showing the system consumption of each algorithm in the experimental examples of this invention as a function of the number of users. Detailed Implementation

[0054] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.

[0055] like Figure 1 As shown, in the scenario of drone-assisted mobile edge computing, edge users constantly generate computing tasks that depend on specific (types) of service resources. With its flexibility and good line-of-sight communication with edge users, the drone can cache a certain amount of service resources and carry an edge server to assist edge users in performing task computing. However, how to reduce the computing latency and energy consumption of edge user tasks is a major challenge.

[0056] This invention takes into account that in real-world scenarios, the computing tasks of edge users are dependent on specific service resources, and that the update actions of edge service caches are not in the same time dimension as the trajectory decision, computing resource allocation decision, and unloading decision of drones. Existing methods have difficulty solving multivariate optimization problems across time dimensions.

[0057] To address the above problems, this invention provides a dual-timescale optimization method for multi-UAV collaborative auxiliary task unloading and caching, such as... Figure 2 As shown, the specific implementation process of the present invention is as follows.

[0058] I. Constructing a Multi-UAV Assisted Task Offloading and Service Caching Model

[0059] like Figure 3 As shown, in a multi-UAV collaborative assisted task unloading system, the total task cycle is divided into... T Hourly time slot (second time slot), each hourly time slot (second time slot) t The duration is Each large time slot (first time slot) is composed of It consists of hourly time slots (second time slots). In real-world scenarios, the update frequency of edge service caches is far less than the frequency of edge resource allocation and collaboration decisions. This is because edge service caches need to be downloaded from the remote cloud; frequent updates would incur significant latency and energy costs. Furthermore, service caches can tolerate a certain level of latency without requiring frequent updates. While updating the cache only requires one cycle per hour slot, drone trajectory, computing resource allocation, and collaborative decision-making are all highly real-time decisions that require frequent updates within each slot. Therefore, the scenario model divides the decision-making into two time scales. One type involves cache decision-making at a large time scale, where at the beginning of each large time slot (the first time slot), the drone needs to update the service cache it carries based on the cache decision. The other type involves drone trajectory, computing resource allocation, and offloading decisions at a small time scale, where at the beginning of each small time slot (the second time slot), the drone flies based on the trajectory decision and then offloads auxiliary computing tasks based on the computing resource allocation and offloading decisions.

[0060] In the scene model, there are M A marginal user, with U Each drone assists the user in performing edge task computation, denoted as follows: In each hourly slot t Each user generates a computation task (referred to as a user task), represented by a triple. ,in, Indicates user The size of the generated task, Indicates user The computing resources required for the generated task Indicates user The types of service resources required by the generated task. In real-world scenarios, user task computation typically requires other resources, such as software library function resources and data resources. Task computation can only be performed on devices that possess the necessary computing service resources. To simplify the problem, a common... There are various service resources, each of which can be represented as: The user device possesses all service resources, while the drone, due to its limited memory capacity, can only store a portion of these resources. The drone's total capacity is... The size of each service resource is , ; in the first hour slot of each large time slot At that time, the caching decision corresponding to the drone can be expressed as: ,and Since service caching decisions are made only in the first hourly ... . Indicates in time slot drones u Service resources are cached. Conversely, uncached service resources... Therefore, the total size of service resources cached on the drone cannot exceed the total memory capacity of the drone. This constraint can be expressed as follows: ; This refers to a collection of drones.

[0061] In each hourly slot (second hourly slot), each drone Flight decisions and computational resource allocation decisions are made at a fixed flight altitude. The flight decision can be represented as... Let represent the flight speed vector of the UAV. The computational resource allocation decision can be represented as... , Indicates drone Assigned to edge users The proportion of computing resources allocated to the drone. Therefore, the drone's flight speed cannot exceed its maximum speed. It can be expressed by the formula as follows: Furthermore, the sum of the computing resources allocated to all users for drones is less than or equal to 1, which can be expressed by the formula: .

[0062] In every hourly gap, every edge user The decision to uninstall can be represented as: . The associated drone for edge users is a drone, and the computing tasks are ultimately offloaded to the device. , This can be the user themselves or any drone. An associated drone refers to a drone that directly connects and communicates with the user. The user must be within the associated drone's range to associate with it. If the user wants to offload a task to a drone... But users are not using drones Within the communication coverage area, the user can only first unload the task to the associated drone, and then the associated drone will transmit the task to the drone. At this time, the drone Also known as collaborative drones. Therefore, the user must be within the communication range of the associated drone. The constraint can be expressed as: ,in, The radius of the communication range mapped onto the ground for the drone. For the user's x and y coordinates, The x and y coordinates are those of the UAV.

[0063] Furthermore, a user can only have one associated drone within a single hourly slot, and tasks can only be ultimately offloaded to the associated drone, a collaborative drone, or local computation. This constraint can be expressed as follows: Regardless of whether the computation task is performed locally by the user or offloaded to the drone, the offloaded device must have the service resources required by the task. This constraint can be represented as follows: .

[0064] II. Constructing the system's communication, latency, and energy consumption calculation models

[0065] In the hourly slot (second time slot) Inside, the drone makes flight decisions. The change in the position of the drone can be expressed by the formula: ,in, The duration of each hourly slot (second hourly slot), Here is the location of the drone. The drone's flight energy consumption can be expressed as:

[0066] ;

[0067] in, The blade profile in the hovering state is a constant; The induced power during hovering is a constant. This represents the tip velocity of the rotor blades; This represents the average rotor blade induced velocity during hovering. Indicates the fuselage drag ratio; Indicates air density; Indicates the rigidity of the rotor blades; This represents the area of ​​the rotor disk. Furthermore, to ensure flight safety, the location of the drone cannot exceed the flight area, and the distance between drones should be greater than or equal to the safe flight distance. These constraints can be expressed as:

[0068] ;

[0069] In the formula, Represents the maximum x and y coordinates of the flight site. Indicates the safe flight distance between drones. Indicates drone i Second time slot t Location, Indicates drone u Second time slot t The location.

[0070] In the system, considering that the links between users and drones, and between drones themselves, are all line-of-sight links, the communication rate between them can be uniformly expressed as:

[0071] ;

[0072] in, For communication bandwidth, For the transmission power of the device, This represents noise power. Indicates device With equipment j The channel power gain between them, where the device can be either a drone or a user equipment; , For reference, the channel power gain at a distance of 1 meter. Indicates device j The location.

[0073] In the time slot Within the system, users can choose to perform computational tasks locally or offload them to a linked drone or a collaborating drone. Latency can be expressed as:

[0074] ;

[0075] in, The logical value for local computation is selected by the user; that is, when the user selects local computation, otherwise The user selects the local computation latency as the local computation latency. ; The logical value for the user to select associated drone unloading calculation, i.e., when the user selects associated drone unloading calculation, ,otherwise The latency for user-selected associated drone unloading calculation is... That is, user transmission latency Computational latency of associated drones sum; The logical value for the user to select assistance with the drone's unloading calculation; when the user selects assistance with the drone's unloading calculation... ,otherwise The latency for users to choose to assist the drone with unloading calculations is... That is, the transmission latency from the user to the associated drone. Transmission latency from associated drones to collaborative drones Computational latency with collaborative drones The sum. Since users can only choose one unloading strategy, when a user selects the associated drone unloading calculation... hour, , The same applies when choosing the other two uninstallation strategies.

[0076] For the above formula, For users Local CPU frequency; Calculate the required resource size for the task; ,in For task size, For the transmission rate between the user and the drone; ,in For the drone's CPU frequency, For drones Assigned to user The proportion of computing resources; , ,in For linking drones With collaborative drones Transmission rate between For drones CPU frequency, For drones Assigned to user The proportion of computing resources.

[0077] Therefore, the time gap Total system latency .

[0078] The energy consumption for task transmission and computation can be expressed as:

[0079] ;

[0080] in, The user selects a logical value for local computing, and the user selects the energy consumption for local computing as the local computing energy consumption. ; The logical value for the user to select the associated drone unloading calculation is [value], and the energy consumption for the user-selected associated drone unloading calculation is [value]. That is, the energy consumption of user task transmission Computing energy consumption of associated drones sum; The logical value for the user to select to assist the drone in unloading calculations; the energy consumption for the user-selected drone unloading calculations is... That is, the energy consumption of user task transmission Related drone transmission energy consumption Computing energy consumption of collaborative drones sum.

[0081] For the above formula, CPU capacitor coefficient; For users Local CPU frequency; In the formula, Transmit power to users, Transmit time for users; In the formula, For the drone's CPU frequency, For drones Assigned to user The proportion of computing resources; In the formula, For drones Transmission power, For the transmission time of the drone; In the formula, For drones CPU frequency, For drones Assigned to user The proportion of computing resources.

[0082] Therefore, the time gap Total system energy consumption This refers to the sum of energy consumption for task transmission, computing, and UAV flight.

[0083] Furthermore, throughout the entire task cycle, the user's transmission and computing energy consumption must not exceed their total energy, which can be expressed as: ,in The user's total energy; the drone's flight energy consumption, transmission energy consumption, and computing energy consumption cannot exceed its total energy. This constraint can be expressed as: ,in, This refers to the total energy of the drone itself.

[0084] III. Determine the optimization variables and optimization objectives

[0085] Throughout the mission lifecycle, service caching decisions are optimized on a large timescale. Specifically, each UAV makes a service caching decision at the beginning of each large time slot (the first time slot), which can be expressed by the formula: The optimization objective is to maximize the sum of the hit rates of various service caches within the communication coverage area of ​​all UAVs in the first time slot, which can be expressed by the formula:

[0086] ;

[0087] in, For drones Cache service resources within the communication coverage area The hit rate within the first time slot can be expressed by the formula: In the formula, This indicates the drone cache service resources in the second time slot. s The logical value is if the drone does not cache service resources within the second time slot. s but The value is 0, indicating that no service resources were cached by the drone during the second time slot. s but =1; Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks, Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks, Represents any type of cached resource; The number of small time slots contained in each large time slot.

[0088] Optimizing drone trajectories, computational resource allocation decisions, and user offloading decisions on a small time scale—that is, each drone making trajectory and computational resource allocation decisions in each hourly time slot (the second time slot)—can be expressed by the following formulas:

[0089] , ;

[0090] Each user makes an uninstallation decision in each hourly slot (second hourly slot), which can be expressed by the formula: . Indicates drone The flight direction and velocity vector, Indicates drone Assigned to user The proportion of computing resources.

[0091] The optimization objective is to minimize the weighted sum of the overall system latency and energy consumption throughout the entire task cycle, which can be expressed by the formula:

[0092] ;

[0093] in, , hour gaps Total system latency and total energy consumption. is the weighting coefficient for time delay, and .

[0094] IV. Modeling the optimization task as a Markov process

[0095] As a preferred embodiment, the present invention uses the DQN algorithm for service cache optimization in the large time slot (first time slot) and the DTD3 algorithm for drone trajectory, computing resource allocation, and offloading decision optimization in the small time slot (second time slot). The solutions to the two optimization sub-problems are based on the deep reinforcement learning framework, so the optimization problem must first be modeled as a Markov decision process.

[0096] Markov decision processes (MDFs) have three key elements: a set of states, a set of actions, and a reward function. In a Markov decision process, the algorithmic agent makes action decisions based on the observed set of environmental states, adapts the environment accordingly, and receives a reward based on the reward function. For large-scale service caching optimization problems, the state set represents the proportion of computational requests for various tasks within each UAV's communication coverage area to the total number of requests, as well as the memory footprint of each service cache. Within each hourly slot, the UAV can only observe user requests within its communication coverage area. Within the second time slot, the drone The set of users within the communication coverage area can be represented as: Then drone Dependency caching service resources within the communication coverage area The number of tasks is: The total number of tasks within the communication coverage area is: Therefore, in the first... Within the first time slot (a large time slot), it includes the first... arrive common The second time slot is used by drones within the larger time slot. Dependency caching within the communication coverage area The total number of tasks is: The total number of tasks within the communication coverage area is: drones Dependency caching within the communication coverage area The proportion of tasks can be expressed as

[0097] ;

[0098] Therefore, the set of states can be represented as:

[0099] ;

[0100] The action set, which is the service cache decision set for each drone, can be represented as:

[0101] ;

[0102] The reward is the sum of the hit rates of various service caches within the communication coverage area of ​​all drones in that large time slot, which can be expressed by the formula: ,in For drones Service cache within communication coverage area The hit rate within a large time slot can be expressed by the formula: .

[0103] For the optimization problem of UAV trajectory, computing resource allocation, and offloading decision within a time slot, the state set is a set of UAV position, user task, and UAV service cache, which can be represented as follows:

[0104] ;

[0105] ;as well as

[0106] .

[0107] The action set, which comprises the trajectory decisions, computational resource allocation decisions, and user offloading decisions for all drones, can be represented as:

[0108] The sum of energy consumption and latency of a small time slot system with a negative reward can be expressed as:

[0109] ;

[0110] in, These are time delay weighting parameters that are determined in advance in the objective function.

[0111] V. Utilize a dual-timescale optimization method to obtain the optimal service caching decision on a large timescale and the optimal UAV trajectory, computing resource allocation, and offloading decision on a small timescale.

[0112] In any given hourly time slot (second time slot), if this hourly time slot (second time slot) happens to be the first hourly time slot (second time slot) within a larger time slot (first time slot), then set the environmental state of the current time slot. As input to the trained large-scale DQN algorithm model, the model outputs the optimal service caching decision for the current time slot. In addition, each hourly slot (second hourly slot) requires the collection of environmental states for that hourly slot (second hourly slot). As input to the action network of the trained small-timescale DTD3 algorithm model, the model outputs the optimal UAV flight decision, resource allocation decision, and user offloading decision for the current time slot. , , When performing actual computational tasks, only the trained algorithm model needs to be used to output the best decision. In addition, the principles of the algorithm model and the specific training process are supplemented below.

[0113] For the DQN algorithm on the large time slot (first time slot), the model consists of two... The network consists of two parts: a training network and a training network. The other is the target network. The network's input is the environment state, and its output is the action decision. During training, the quadruples obtained from the interaction between each time step and the environment are used... Add to the experience replay pool, among which It refers to the environmental condition. It is the action decision made. In the environmental state Make an action The reward received This is the next environment state. After the interaction is complete, if the number of tuples in the experience pool exceeds... Take from the experience replay pool Group interaction data For each set of data, the target network is used to calculate... ,in, The set reward discount factor is then used to minimize the target loss. To update the training network. In addition, each The target network is updated every time step, that is, the parameters of the training network are copied to the target network. Parameters are updated periodically. Training is complete when the set number of training steps is reached.

[0114] This invention improves the dual-delay deep deterministic policy gradient algorithm (TD3) by replacing the action network in TD3 with a diffusion model, thus obtaining the diffusion dual-delay deep deterministic policy gradient algorithm (DTD3).

[0115] For the Diffusion-Dual-Delay Deep Deterministic Policy Gradient (DTD3) algorithm on the second time slot, the algorithm consists of six networks: one training action network, one target action network, two training comment networks, and two target comment networks. The action network of the DTD3 algorithm does not use a traditional MLP network model, but rather an action network model based on a back-diffusion process of a diffusion model, which can generate richer actions and thus explore more efficient actions. The comment network uses a dual-comment network to improve upon the limitations of traditional single-comment networks. The problem of overestimating performance is addressed by decoupling the update processes of the action network and the comment network in the overall update process, delaying the update of the action network, and reducing the training oscillation and instability caused by the simultaneous updates of the action network and the comment network.

[0116] The Denoising Diffusion Model (DDPM) was originally developed as a technique for image generation tasks. The model consists of two processes: forward diffusion and reverse denoising. Forward diffusion involves gradually adding noise to a high-quality image until it becomes a blurred image approximating a standard Gaussian distribution. Reverse denoising, on the other hand, denoises the blurred image until it is restored to its original high-quality state. In this problem, the actions taken by the agent are abstractions of image concepts. The forward process involves gradually adding noise to a high-value action until it becomes a blurred action approximating a standard Gaussian distribution. Reverse denoising denoises the approximating Gaussian distribution action to obtain a high-value action. However, since this invention is a MINLP problem, obtaining the original optimal solution is extremely challenging. Therefore, in the diffusion action network model, this invention only uses the reverse denoising process, while the forward diffusion process serves only as the theoretical basis for the reverse process.

[0117] In the diffusion model, the number of diffusion steps is set to... That is, fuzzy action is ,go through Denoising steps yield high-value original motion. In the noise reduction process, first... We initialize it according to the standard Gaussian distribution, and then we use a parameter of... The network is used as a denoiser. The inputs to the denoiser are the action of the current denoising step, the step size index, and the system state. The output is the action of the previous denoising step, thus restoring the original action after denoising. Specifically, according to the forward diffusion process, the actions of the previous step and the current step follow a Gaussian distribution. ,in Since it represents the identity matrix, we only need to find the mean of the distribution. and variance This will give you the action from the previous step. Right now arrive The cumulative product, of which ,in, The set total number of diffusion steps, These are the preset minimum and maximum rates. It can be transformed and re-represented as: ,in , Right now arrive The cumulative product, That is, the input of the noise canceller network in step i is The output. Finally, according to Reparameterization is performed to obtain the action after the previous denoising step. ,in, The standard Gaussian distribution can be expressed by the formula: .go through Denoising steps are performed to obtain the original motion. As the output of the diffusion action network.

[0118] The architecture diagram of the DTD3 algorithm is as follows: Figure 4 As shown, the overall training process is as follows: First, initialize the parameters of the action network and the comment network as follows: And copy its parameters to the corresponding target network respectively Set a total training time. Each round includes [number] rounds, and each round includes [number] rounds. Each time slot, with a delay update parameter of [number] timeslots. Each time slot first inputs the state into the diffusion action network and then... Denoising is performed to obtain action decisions, and interaction data quadruples are obtained based on the interaction between the decisions and the environment. The data is then placed into the experience replay pool. If the sample size of the experience replay pool exceeds... Take from the experience replay pool Group interaction data Calculate for each sample By minimizing the loss function Updated comments network If the delayed update step count is reached at this point... ,in, t For the current time slot, D To delay the parameter update, according to the formula Calculate the policy gradient, delay updating the action network, and then... Perform target comment network soft updates, based on Perform soft updates to the action network. Training ends after the set number of training epochs. To use the trained model, simply input the environmental state into the diffusion action network, and then... After the reverse diffusion, the action decision is returned.

[0119] Based on the training processes of the large-slot DQN algorithm and the small-slot DTD3 algorithm described above, the overall training process of the dual-slot algorithm can be divided into the following steps:

[0120] 1. Initialize the neural network for the DQN algorithm Initialize the action network of DTD3 Comments on the Internet , ;

[0121] 2. Connect the network , , , The parameters are copied to the corresponding target network respectively. , , , ;

[0122] 3. Initialize the experience recycling pools for DQN and DTD3 algorithms respectively. , .

[0123] 4. Initialize the number of time slots (second time slot) for the training task cycle. 2000, each large time slot (first time slot) contains Hourly time slot, current time slot ;

[0124] 5. If the current time slot is the first smaller time slot (second time slot) of a larger time slot (first time slot), then... Set the value to 1, and set the environmental state to 1. enter Get action To perform the action;

[0125] 6. If the current time slot is the last hour slot (second time slot) of a larger time slot (first time slot), then... The value is 0, and the current environment status is... Earn reward points ,Will Store to From Sample N samples and update the network with the goal of minimizing the loss. and each Each major time slot is updated synchronously. , The periodic update parameters set for the DQN algorithm;

[0126] 7. Set environmental state variables enter , to obtain action Perform an action and receive a reward. The environmental state changed to ,Will Store to .from Take N samples from the input. , Output respectively , get output Small value network Update, and every step , To perform the update, D is the delayed update parameter of the DTD3 algorithm;

[0127] 8. If If the value is 2000, the training ends; otherwise, skip to step 5 and continue training.

[0128] VI. Execute the UAV trajectory, computing resource allocation, caching decisions, and offloading decisions given by the algorithm model.

[0129] In each hourly time slot (second time slot), the UAV flight decision, resource allocation decision, and user offloading decision are based on the output of the DTD3 algorithm model. , , Perform user task offloading, drone flight, and task calculation. If this hourly time slot (second time slot) is the first hourly time slot (second time slot) of a larger time slot (first time slot), make a service caching decision based on the output of the DQN algorithm model. Update the service cache (cache service resources) on the drone.

[0130] Test case

[0131] Simulation experiments were conducted on the method proposed in this invention, and existing caching decision algorithms and reinforcement learning algorithms were integrated into a dual-timescale framework for experimental simulation and performance comparison. Specifically, the dual-timescale optimization method (DQN-DTD3) for multi-UAV cooperative assisted task offloading and caching proposed in this invention was compared with the following five schemes:

[0132] DQN-SAC: Caching decisions are optimized using the DQN algorithm, while drone trajectory, resource allocation, and user offloading decisions are optimized using the Soft Actor-Commentator (SAC) algorithm.

[0133] DQN-DDPG: Caching decisions are optimized using the DQN algorithm, while drone trajectory, resource allocation, and user offloading decisions are optimized using the Deep Deterministic Policy Gradient (DDPG) algorithm.

[0134] DQN-DTD3-NOC: The cache decision is optimized using the DQN algorithm, while the drone trajectory, resource allocation, and user offloading decisions are optimized using the Deep Deterministic Strategy Gradient Algorithm (DTD3). However, unlike other comparison algorithms, this algorithm does not involve cooperation between drones.

[0135] LFU-DTD3: The caching decision is optimized using the Least Frequently Used (LFU) algorithm, while the drone trajectory, resource allocation, and user offloading decisions are optimized using the Diffusion Dual Delay Deep Deterministic Strategy Gradient Algorithm (DTD3).

[0136] LRU-DTD3: The cache decision is optimized using the Least Recently Used (LRU) algorithm, while the drone trajectory, resource allocation, and user offloading decisions are optimized using the Diffusion Dual Delay Deep Deterministic Strategy Gradient Algorithm (DTD3).

[0137] Figure 5 This is the experimental result of the system's latency-weighted energy consumption increasing with the increase of the UAV's cache capacity. Figure 5 It can be seen that the system consumption of the proposed method decreases as the drone's cache capacity increases. This is mainly because the drone can store more service cache, allowing users to offload more tasks to the drone, thus reducing task computation latency. Furthermore, the system consumption of the proposed method is significantly lower than other comparative algorithms. Specifically, when the drone's cache capacity increases from 5 to 25, the system consumption of the proposed algorithm decreases from 152.5 to 82.8. Overall, among the five comparative algorithms, the DQN-SAC algorithm has the lowest system consumption, while the LRU-DTD3 algorithm has the highest. The DQN-DTD3 algorithm proposed in this invention has an average system consumption reduction of 10.6 compared to the DQN-SAC algorithm and 38.2 compared to LRU-DTD3. Figure 6 The experimental results show that the system consumption of the proposed method decreases as the number of drones increases. This is mainly because with a larger number of drones, each drone needs to serve fewer users, providing more computing resources to each user on average, and reducing computational latency. Compared with other methods, when the number of drones is 1, the proposed algorithm performs the same as the DQN-DTD3-NOC algorithm. This is because the cooperative and non-cooperative effects are the same when there is only one drone. When the number of drones is greater than 1, the system consumption of the proposed algorithm is lower than that of DQN-DTD3-NOC. Specifically, the system consumption of the proposed algorithm is reduced by an average of 14.4, 20.2, 12.0, 39.6, and 51.0 compared to DQN-SAC, DQN-DDPG, DQN-DTD3-NOC, LFU-DTD3, and LRU-DTD3, respectively. Figure 7 The experimental results show that system consumption increases with the number of users. It is evident that system consumption increases with the number of users because this is due to the increased computational workload, energy consumption, and computational latency. Furthermore, the system energy consumption of the proposed method is lower than other algorithms under different user counts. Specifically, when the number of users increases from 20 to 60, the system consumption of the proposed algorithm increases from 59.9 to 175.3. Compared with other algorithms, the proposed algorithm's system consumption is reduced by an average of 24.2, 44.0, 51.6, 94.0, and 117.38% compared to DQN-SAC, DQN-DDPG, DQN-DTD3-NOC, LFU-DTD3, and LRU-DTD3, respectively.

[0138] Based on the above three sets of experiments, the proposed method has lower system consumption and better performance than other comparative algorithms, even when the drone cache capacity, number of drones, and number of users change.

[0139] This invention proposes a dual-timescale optimization method for multi-UAV collaborative task offloading and caching. Addressing the issue of mobile edge users having service-dependent computational task requirements but insufficient computing power, this method utilizes multiple UAVs carrying service caches and edge servers to assist users in task computation. By jointly optimizing UAV trajectories, computational resource allocation, service caching, and user offloading decisions, it achieves overall low latency and low energy consumption, thereby improving user experience and system efficiency. The method aims to minimize the weighted sum of overall system latency and energy consumption. Considering that the update frequency of service cache is lower than that of other optimization variables in real-world scenarios, and addressing the inconsistency in the timescales of optimization variables, this method uses the DQN algorithm to optimize the UAV service cache on a large timescale, and the Diffusion Dual-Delay Deep Deterministic Policy Gradient Algorithm (DTD3) to optimize UAV trajectories, computational resource allocation, and offloading decisions on a small timescale. This effectively solves the service dependency problem of computational tasks and the optimization problem across time dimensions. Using a diffusion model as the algorithm's policy network enhances the model's action expression ability, improves the quality of policy decisions, and provides better adaptability to dynamically changing environments.

[0140] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A dual-time-scale optimization method for unloading and caching of multi-UAV collaborative auxiliary tasks, characterized in that, include: The update time slot for the drone service cache decision is set as the first time slot, and the update time slots for the drone flight trajectory decision, computing resource allocation decision, and user unloading decision are set as the second time slot; The first time slot is composed of multiple second time slots; Construct the first and second optimization objective functions; In each first time slot, the user task's demand for various service resources is obtained, and the service cache decision for the UAV is updated according to the first optimization objective function; In each second time slot, the drone location, user task, and drone service cache are acquired, and the drone flight trajectory decision, computing resource allocation decision, and user offload decision are updated according to the second optimization objective function. The first optimization objective function is: ; The second optimization objective function is: ; In the formula, Representing drones, Indicates the total number of drones; Representative service resource types, This indicates the total number of service resource types; Indicates the first time slot. This represents the total number of first time slots; Indicates drone Cache service resources within the communication coverage area Hit rate in the first time slot, Indicates the second time slot. The weighting factor for time delay is... , The second time slot Total latency and total energy consumption of the internal system This indicates the total number of second time slots. These represent user uninstallation decisions, drone caching decisions, drone flight trajectory decisions, and computing resource allocation decisions, respectively. In the first time slot, the service caching decision for the UAV is determined using the DQN model; The reward function during the training of the DQN model is described above. Set to: ; in, Indicates drone Cache service resources within the communication coverage area Hit rate in the first time slot, Representing drones, Representative service resource types, This indicates the total number of service resource types; Construct a dual-delay deep deterministic policy gradient algorithm model, which includes an action network and a comment network; The action network of the dual-delay deep deterministic policy gradient algorithm model is replaced with a diffusion model to obtain the diffusion dual-delay deep deterministic policy gradient algorithm model. Using UAV location, user task, and UAV service cache as state variables, and human-machine flight trajectory decision, computing resource allocation decision, and user unloading decision as action variables, the diffusion dual-delay deep deterministic policy gradient algorithm model is trained to obtain the optimal diffusion dual-delay deep deterministic policy gradient algorithm model. Among them, the reward function for model training Set to: ; In the formula, Indicates the second time slot. The weighting factor for time delay is... , The second time slot Total latency and total energy consumption of the internal system; In the second time slot, the drone's location, user tasks, and drone service cache are acquired, and the drone flight trajectory decision, computing resource allocation decision, and user offload decision are obtained through the optimal diffusion dual-delay deep deterministic strategy gradient algorithm model.

2. The dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching according to claim 1, characterized in that, drones Cache service resources within the communication coverage area The formula for calculating the hit rate in the first time slot is: ; in, Indicates the drone in the second time slot Cache service resources logical value, Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks, Indicates the drone in the second time slot Communication coverage area depends on cache service resources The number of user tasks; The number of small time slots contained in each large time slot.

3. The dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching according to claim 2, characterized in that, The formula for calculating the total system delay within the second time slot is: ; in: ; In the formula, For users to select the latency for associated drone unloading calculations, This indicates the logical value that the user selected to uninstall the associated drone; Allow users to select the local computation latency. This indicates that the user has selected a locally calculated logical value; Users can select the latency for assisting drone unloading calculations. This indicates the logical value representing the user's choice to assist with drone unloading; Indicates user, This indicates the total number of users. This indicates the second time slot.

4. The dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching according to claim 3, characterized in that, The formula for calculating the latency of user-selected drone unloading calculations is as follows: ; In the formula, This indicates the transmission latency from the user to the associated drone. This indicates the transmission latency from the associated drone to the collaborating drone. This indicates the computational latency of the collaborative drone.

5. The dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching according to claim 4, characterized in that, The formula for calculating the total energy consumption of the system in the second time slot is: ; in, This indicates the energy consumption for user task transmission and computation. On behalf of users, This indicates the total number of users. Indicates the energy consumption of drone flight. Representing drones, This indicates the total number of drones.

6. The dual-time-scale optimization method for multi-UAV collaborative auxiliary task unloading and caching according to claim 5, characterized in that, Update the drone's service cache decision in the first second time slot of each first time slot.

Citation Information

Patent Citations

  • Dual-time-scale joint optimization service caching and calculation unloading method of unmanned aerial vehicle assisted MEC

    CN118612797A

  • Assisted edge calculation method for space-air-ground integrated network

    CN120151948A