Unmanned aerial vehicle edge computing network resource allocation method based on diffusion model

By adopting a diffusion model-based resource allocation method in the drone-assisted edge computing network, the system model and data transmission model are built, and combined with the actor network and the target dual critic network, the problem of low resource allocation efficiency in the drone-assisted edge computing network is solved, and efficient support and resource optimization for different AIGC tasks are achieved.

CN119946723AActive Publication Date: 2025-05-06BEIHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411922092.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Existing drone-assisted edge computing networks are difficult to efficiently allocate limited network resources and cannot effectively support different types of artificial intelligence generated content (AIGC) tasks.

Method used

A resource allocation method based on diffusion model is adopted to generate the optimal resource allocation strategy by building a system model, data transmission model and energy consumption model, combining the actor network and the target dual critic network.

Benefits of technology

It realizes efficient support for different types of AIGC tasks, improves service efficiency and resource utilization, and reduces system energy overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946723A_ABST
    Figure CN119946723A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of edge computing, in particular to an unmanned aerial vehicle edge computing network resource allocation method based on a diffusion model, which comprises the following steps: acquiring content generation task information randomly generated by any edge user in a target time slot multi-unmanned aerial vehicle auxiliary edge computing network and service selection of the edge user, constructing a system model and a data transmission model corresponding to each service selection; constructing a service selection model and an energy consumption model; determining an overall optimization target based on the system model, the service selection model, the data transmission model and the energy consumption model; and constructing an actor network and a target double-commentator network based on the overall optimization target and the diffusion model, and performing parameter updating on the actor network and the target double-commentator network to generate an optimal resource allocation strategy. According to the invention, efficient resource allocation can be realized by using limited network resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of edge computing technology, and in particular to a method for allocating network resources for unmanned aerial vehicle edge computing based on a diffusion model. Background Art

[0002] Edge computing network resource allocation refers to providing efficient allocation strategies for limited energy, computing, and communication resources in the network through linear programming, heuristic algorithms, game theory, reinforcement learning, and other methods to meet performance indicators such as energy consumption optimization and computing resource utilization. Traditional network infrastructure is limited by flexibility and cost and cannot meet the increasingly widespread edge terminal computing needs. Drone-assisted edge computing takes advantage of the high flexibility and mobility of drones and deploys computing resources on drones to provide low-latency, high-efficiency computing services for terminal edge devices.

[0003] With its own mobility, flexibility and low cost, the edge computing network assisted by UAVs is an important support means for realizing the integrated network of air, space and ground in the sixth-generation mobile communication system, and is widely used in scenarios such as smart cities and smart industries. However, with the advent of the era of large models represented by ChatGPT, the surge in terminal devices' demand for AI-generated content (AIGC) has further increased the network burden, and the large number and variety of terminal computing tasks have brought further challenges to the edge computing network assisted by UAVs.

[0004] At present, there are few studies on multi-UAV-assisted edge computing networks that provide AIGC services to end users. Most of the research focuses on a single UAV as an AIGC service provider, and does not consider the service quality requirements of different types of AIGC tasks. Therefore, how to use efficient resource allocation under limited network resources to support different types of AIGC tasks has become a technical problem that restricts the development of UAV-assisted edge computing networks. Summary of the invention

[0005] To this end, the present invention provides a UAV edge computing network resource allocation method based on a diffusion model to overcome the problem in the prior art that it is impossible to support different types of AIGC tasks by utilizing efficient resource allocation under limited network resources.

[0006] To achieve the above object, the present invention provides a method for allocating network resources of UAV edge computing based on a diffusion model, comprising:

[0007] Step S1, obtaining the content generation task information randomly generated by any edge user in the multi-UAV-assisted edge computing network in the target time slot and the service selection of the edge user, and constructing a system model and a data transmission model corresponding to each service selection;

[0008] The multi-drone assisted edge computing network includes a plurality of drones, a plurality of edge users, and a plurality of edge servers; the content generation task information includes the task data size, the computing load required for the task, the edge user priority, and the maximum service delay limit that the edge user can tolerate;

[0009] Step S2, calculating the corresponding task completion delay according to the system model and each of the data transmission models, and building a service selection model based on the completion delay of each task and the priority of the content generation task;

[0010] Step S3, constructing an energy consumption model based on the energy consumed by the edge user, the edge server, and the GPU of the drone to complete the content generation task in the target time slot;

[0011] Step S4, determining an overall optimization target based on the system model, the service selection model, the data transmission model and the energy consumption model;

[0012] Step S5, constructing an actor network and a target dual critic network based on the overall optimization goal and the diffusion model, and updating parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

[0013] Furthermore, in the step S1, it includes:

[0014] Step S11, constructing a system model based on the content generation task information;

[0015] Step S12, determining a corresponding channel type based on the service selection of the edge user, and obtaining additional interference corresponding to each channel type;

[0016] Step S13, determining a channel coefficient corresponding to each channel type according to the system model;

[0017] Step S14: constructing a data transmission model corresponding to each channel type according to each of the channel coefficients and the additional interference corresponding to each of the channel types.

[0018] Furthermore, in the step S2, it includes:

[0019] Step S21, constructing a service selection sequence based on the service selection of the edge user, and designing corresponding service constraints;

[0020] Step S22, calculating the task completion delay corresponding to each service selection based on the service selection sequence and the system model;

[0021] Step S23: constructing a service selection model according to the completion delay of each task and the priority of the content generation task.

[0022] Furthermore, in the step S22, it includes:

[0023] The task calculation delay and task transmission delay corresponding to each service selection are determined based on the system model, the service selection sequence and the corresponding data transmission model, and the task completion delay corresponding to each service selection is calculated according to the task calculation delay and the task transmission delay.

[0024] Furthermore, in the step S23, it includes:

[0025] The utility function corresponding to each task priority is determined according to the task completion delay and the content generation task priority, and a service selection model is constructed according to each utility function.

[0026] Furthermore, in the step S3, it includes:

[0027] Based on the GPU power of the edge user, the edge server and the drone in the target time slot and the task calculation delay corresponding to each service selection, the energy consumed by the GPU of the edge user, the edge server and the drone to complete the content generation task in the target time slot is calculated to construct an energy consumption model.

[0028] Furthermore, in the step S4, it includes:

[0029] Step S41, designing the task utility function of the edge user in the target time slot based on the service selection model;

[0030] Step S42, constructing an optimization objective function based on the task utility function and the energy consumption model;

[0031] Step S43, constructing optimization constraints based on the system model, the data transmission model and the service selection model;

[0032] Step S44: determining an overall optimization target based on the optimization objective function and the optimization constraints.

[0033] Further, in the step S5, it includes:

[0034] Step S51, determining the state space, action space and reward function of the Markov decision process based on the overall optimization objective and the reverse diffusion process of the diffusion model;

[0035] Step S52, constructing an actor network based on the state space, the action space and the reward function;

[0036] Step S53: construct a target dual critic network, design update gradients for the target dual critic network and the actor network respectively, and update parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

[0037] Furthermore, in step S1, the computational load required for the task is calculated according to the number of denoising steps required for the content generation task.

[0038] Furthermore, in the step S1, the service selection of the edge user includes offloading the content generation task to the drone, offloading the content generation task to the adjacent drone, and offloading the content generation task to the edge server.

[0039] Compared with the prior art, the beneficial effect of the present invention is that the present invention can design the system model and service selection according to the needs of edge users by constructing a system model and a data transmission model, so as to realize the simultaneous allocation of multiple tasks in one time slot and improve service efficiency. By constructing a service selection model and designing different priorities for content generation tasks, different system benefits brought by different tasks can be distinguished. Since the energy consumption required by the GPU to complete the content generation task is much higher than the energy consumed by signal transmission, by constructing an energy consumption model, only the energy consumption required for task completion is considered, which can reduce the amount of data processing and improve processing efficiency. By designing the overall optimization goal and clarifying the goal of the system, it is possible to ensure that the solution of task allocation and resource management is obtained while reducing the energy overhead of the system, so as to maximize the utility of edge user tasks. By constructing an actor network and a target dual critic network based on the overall optimization goal and the diffusion model, it is possible to generate strategies for dynamic resource allocation with the help of the powerful ability of the diffusion model to fit complex dynamic environments, reduce the estimation error of the Q value, and the optimal resource allocation strategy generated in this way can achieve efficient resource allocation with limited network resources.

[0040] Furthermore, the present invention designs a corresponding data transmission model according to the service selection of edge users, considers the additional interference of each channel type, and reasonably allocates computing resources, thereby saving network resources and achieving accurate and efficient use of limited network resources.

[0041] Furthermore, the present invention can reasonably allocate service selections according to computing resources and improve service capabilities by constructing a service selection sequence and service constraints.

[0042] Furthermore, the present invention designs a utility function according to content selection priority, which can ensure non-negative benefits brought by timely completion of high-priority tasks on the one hand, and prevent low-priority tasks from being ignored for a long time on the other hand.

[0043] Furthermore, the present invention converts the optimization problem into a Markov decision problem and introduces the reverse diffusion process into the decision-making process, which can reduce computational complexity, adapt to complex and changing environments, improve system performance, and achieve efficient and reasonable resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of a method for allocating network resources of a drone edge computing system based on a diffusion model according to an embodiment of the present invention;

[0045] Figure 2 This is a flow chart of step S1 of an embodiment of the present invention;

[0046] Figure 3 This is a flow chart of step S2 of an embodiment of the present invention;

[0047] Figure 4 This is a flow chart of step S4 of an embodiment of the present invention;

[0048] Figure 5 This is a flow chart of step S5 of an embodiment of the present invention;

[0049] Figure 6 A flow chart of a service selection model for providing content generation tasks in an embodiment of the present invention;

[0050] Figure 7 A schematic diagram of a process for generating an optimal resource allocation strategy according to an embodiment of the present invention;

[0051] Figure 8 A comparison chart of the average reward per task of the drone edge computing network resource allocation method based on the diffusion model in an embodiment of the present invention and the current mainstream reinforcement learning algorithms Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) under different user arrival strategies;

[0052] Fig. 9 This is a comparison chart of the average reward per training task under different user arrival strategies of the drone edge computing network resource allocation method based on the diffusion model in an embodiment of the present invention and the current mainstream reinforcement learning algorithms Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC). DETAILED DESCRIPTION

[0053] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0055] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0056] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0057] See also Figure 1 , Figure 6 , Figure 7 As shown, Figure 1 This is a flow chart of a method for allocating network resources of a drone edge computing system based on a diffusion model according to an embodiment of the present invention; Figure 6 A flow chart of a service selection model for providing content generation tasks in an embodiment of the present invention; Figure 7 The present invention provides a method for allocating network resources of a drone edge computing system based on a diffusion model, comprising:

[0058] Step S1, obtaining the content generation task information randomly generated by any edge user in the multi-UAV-assisted edge computing network in the target time slot and the service selection of the edge user, and constructing a system model and a data transmission model corresponding to each service selection;

[0059] The multi-drone assisted edge computing network includes a plurality of drones, a plurality of edge users, and a plurality of edge servers; the content generation task information includes the task data size, the computing load required for the task, the edge user priority, and the maximum service delay limit that the edge user can tolerate;

[0060] See also Figure 2 As shown, it is a flow chart of step S1 of an embodiment of the present invention; specifically, in step S1, it includes:

[0061] Step S11, constructing a system model based on the content generation task information;

[0062] Step S12, determining a corresponding channel type based on the service selection of the edge user, and obtaining additional interference corresponding to each channel type;

[0063] Step S13, determining a channel coefficient corresponding to each channel type according to the system model;

[0064] Step S14: constructing a data transmission model corresponding to each channel type according to each of the channel coefficients and the additional interference corresponding to each of the channel types.

[0065] Specifically, in step S1, the computational load required for the task is calculated according to the number of denoising steps required for the content generation task.

[0066] Specifically, in step S1, the service selection of the edge user includes offloading the content generation task to the drone, offloading the content generation task to the adjacent drone, and offloading the content generation task to the edge server.

[0067] In a specific embodiment, the multi-drone assisted edge computing network consists of M drones, N edge users, and 1 edge server. To facilitate system analysis, the time T is divided into several time slots of duration τ, and each drone is equipped with an omnidirectional antenna. In the target time slot t, each edge user n will randomly generate content generation tasks, and the task system model D of the edge user in the target time slot t can be constructed. n (t) = {d n (t),c n (t),μ n (t),p n (t)}, where d n (t) is the task data size, c n (t) is the computational load required for the content generation task, μ n (t) is the maximum service delay that edge user n can tolerate, p n (t) is the priority of edge user n. In this model, each UAV acts as an AIGC service provider and can also act as a relay to offload tasks within its service range W to other UAVs or edge servers.

[0068] It is understandable that the present invention ignores the downlink transmission of the task result, considering that the task result is usually much smaller than the task itself. Based on the role played by the drone, the service option for edge users is to offload the content generation task to the drone, and the corresponding channel type is the ground-to-air channel between the edge user and the drone; the service option for edge users is to offload the content generation task to the adjacent drone, and the corresponding channel type is the ground-to-air channel between the edge user and the drone plus the air-to-air channel between drones; the service option for edge users is to offload the content generation task to the edge server, and the corresponding channel type is the ground-to-air channel between the edge user and the drone plus the air-to-ground channel between the drone and the edge server;

[0069] For the ground-to-air channel, considering that each drone can simultaneously serve multiple edge users by using multi-user multiple-input multiple-output technology, the three-dimensional coordinates of drone m at the target time slot t can be expressed as follows: in, and represent the coordinates corresponding to the X, Y and Z axes respectively. The position of edge user n can be expressed as: n (t) = [x n (t),y n (t),0] T , the present invention considers both line-of-sight (LoS) channels and non-line-of-sight (NLoS) channels in the ground-to-air communication link. The probability of a line-of-sight channel between edge user n and drone m is It can be expressed as:

[0070]

[0071] Among them, α and are constants determined by the environment, d n,m (t) is the distance between edge user n and drone m at time slot t. The probability of non-line-of-sight channel can be obtained from formula (1): From this we can get the channel gain between the two:

[0072]

[0073] Where λ0 is the channel gain per unit distance, represents the path loss exponent, ζ∈(0,1) is the additional attenuation factor under the non-line-of-sight channel, so the average channel gain of the ground-to-air channel can be expressed as:

[0074]

[0075] Use g n,m(t) represents the channel coefficient between edge user n and drone m, which can be expressed as:

[0076]

[0077] in, is the small-scale fading coefficient, which obeys the expectation Therefore, the data transmission rate of edge user n and drone m at time slot t can be expressed as follows:

[0078]

[0079] Among them, B GA represents the channel bandwidth between the ground-to-air channels, P n (t) is the transmit power of edge user n in time slot t, and are the additional Gaussian noise and the interference caused by other edge users j within the coverage area of ​​UAV m;

[0080] For air-to-air channels, in order to make full use of the computing resources of all drones, the present invention assumes that drones can receive AIGC tasks from ground edge users and forward them to other drones. Considering the full-duplex communication capability of each drone, the receiving and forwarding tasks are considered to occur simultaneously. In order to simplify the design, the present invention assumes that the data transmission path loss between drones can be expressed by free space path loss (FSPL), because there are usually fewer obstructions between drones and they are mainly LoS channels. Based on the Friis transmission formula, the path loss between drone m and drone m' can be expressed as:

[0081] L m,m′ (t) = 20log 10 (d m,m′ (t))+20log 10 (f)+32.45, (6)

[0082] Among them, d m,m' (t) is the distance between UAV m and UAV m' at time slot t, measured in kilometers (km). f represents the signal frequency, measured in megahertz (MHz). Therefore, the data transmission rate of the air-to-air channel can be expressed as:

[0083]

[0084] in, represents Gaussian noise, B AA represents the channel bandwidth between empty-empty channels, P m(t) represents the signal power of UAV m at time slot t;

[0085] For the air-ground channel model, similar to the ground-air channel model, the data transmission rate between UAV m and the edge server can be expressed as:

[0086]

[0087] in, represents Gaussian noise, B AG represents the channel bandwidth between the air-ground channel, and the location of the edge server is represented by q0 = [x0, y0, 0] T .

[0088] The present invention designs a corresponding data transmission model according to the service selection of edge users, considers the additional interference of each channel type, and reasonably allocates computing resources, which can save network resources and achieve accurate and efficient utilization of limited network resources.

[0089] Step S2, calculating the corresponding task completion delay according to the system model and each of the data transmission models, and building a service selection model based on the completion delay of each task and the priority of the content generation task;

[0090] See also Figure 3 As shown, it is a flow chart of step S2 of an embodiment of the present invention; specifically, in step S2, it includes:

[0091] Step S21, constructing a service selection sequence based on the service selection of the edge user, and designing corresponding service constraints;

[0092] Step S22, calculating the task completion delay corresponding to each service selection based on the service selection sequence and the system model;

[0093] Specifically, in step S22, it includes:

[0094] The task calculation delay and task transmission delay corresponding to each service selection are determined based on the system model, the service selection sequence and the corresponding data transmission model, and the task completion delay corresponding to each service selection is calculated according to the task calculation delay and the task transmission delay.

[0095] Step S23: constructing a service selection model according to the completion delay of each task and the priority of the content generation task.

[0096] Specifically, in step S23, it includes:

[0097] The utility function corresponding to each task priority is determined according to the task completion delay and the content generation task priority, and a service selection model is constructed according to each utility function.

[0098] The present invention can reasonably allocate service selection according to computing resources and improve service capabilities by constructing a service selection sequence and service constraints.

[0099] In a specific embodiment, since each edge user has a certain computing capability, they can choose to complete the computation locally or upload the content generation task to the drone and the edge server with sufficient computing resources. The service selection sequence of edge user n in time slot t can be expressed as: in, Indicates that the user chooses to perform calculations locally. indicates that the task is executed on the edge server, and the rest indicate that the edge user offloads the task to the designated drone for calculation. The task of each edge user can only be calculated at one place in time slot t, so the service selection sequence of each edge user n follows the following constraints:

[0100]

[0101] Use F n (t) represents the task completion delay of edge user n in time slot t. When edge user n chooses to perform calculations locally, the task calculation delay can be expressed as follows:

[0102]

[0103] in, is the time required for each step of denoising when completing the content generation task at the edge user n. In this case, the task completion delay can be expressed as:

[0104]

[0105] When edge user n chooses to offload the content generation task to the drone, the task computation delay can be expressed as:

[0106]

[0107] in, k=1,...,M, the computing power of UAV k is expressed by the time required for each denoising step. Considering that each UAV has limited computing resources at time slot t, the following constraints need to be met:

[0108]

[0109] Among them, C krepresents the denoising capability of drone k. Assuming that edge user n is within the service range of drone m at time slot t, considering that edge user n can offload tasks to drone m or to other drones m' through drone m, and considering that all drones have full-duplex communication capabilities, the task transmission delay can be expressed as follows:

[0110]

[0111] Therefore, the task completion delay can be expressed as:

[0112]

[0113] If the computing resources required for the content generation task are too large, they can be transmitted to the edge server for calculation through drone relay. In this case, the task calculation delay can be expressed as:

[0114]

[0115] in, is the denoising capability of the edge server. Considering that the edge server has sufficient computing resources, the present invention does not consider the computing resource limitation of the edge server. At this time, the task transmission delay can be expressed as:

[0116]

[0117] Therefore, the task completion delay can be expressed as:

[0118]

[0119] Considering that tasks of different priorities bring different utilities, this embodiment defines a service priority for each task, and uses a to represent the boundary between high and low priorities. n When (t)≤a, it is a low priority task; otherwise, it is a high priority task. The utility function of the high priority task is defined as follows:

[0120]

[0121] in, is the utility coefficient of completing different high-priority tasks, μ max Is the highest priority; P H <0 is the penalty constant for not completing the high priority task in time. Different from other inventions, this embodiment uses the utility coefficient δ n (t), we can specifically distinguish the different system gains brought by the completion of different tasks in the high priority interval. At the same time, by using log2(1+μ n (t)-F nThe (t)) function can, on the one hand, ensure the non-negative benefits brought by the timely completion of high-priority tasks. On the other hand, considering its decreasing slope, it can effectively reflect the situation that the benefits brought by the completion of high-priority tasks within the tolerance delay range increase with the increase of advance completion time, but the growth rate slows down, which is in line with the actual long-tail effect.

[0122] The utility function of the low priority task is defined as follows:

[0123]

[0124] in, is the utility coefficient for completing different low-priority tasks.

[0125] Because low-priority tasks are usually less sensitive to time, this embodiment sets a fixed reward value R for completing low-priority tasks. L > 0, in order to avoid low priority tasks being ignored for a long time, the present invention sets the penalty value 0>P L >P H To punish low-priority tasks for not completing them on time.

[0126] The present invention designs a utility function according to content selection priority, which can ensure non-negative benefits brought by timely completion of high-priority tasks on the one hand, and prevent low-priority tasks from being ignored for a long time on the other hand.

[0127] Step S3, constructing an energy consumption model based on the energy consumed by the edge user, the edge server, and the GPU of the drone to complete the content generation task in the target time slot;

[0128] Specifically, step S3 includes:

[0129] Based on the GPU power of the edge user, the edge server and the drone in the target time slot and the task calculation delay corresponding to each service selection, the energy consumed by the GPU of the edge user, the edge server and the drone to complete the content generation task in the target time slot is calculated to construct an energy consumption model.

[0130] In a specific embodiment, considering that the energy consumption required by the GPU to complete the content generation task is much higher than the energy consumed by signal transmission, this embodiment only considers the energy consumption required to complete the task in order to simplify the model. The energy consumed by the system in time slot t can be expressed as:

[0131]

[0132] in, k ES and These are the GPU power of edge users, edge servers, and drones, respectively.

[0133] Step S4, determining an overall optimization target based on the system model, the service selection model, the data transmission model and the energy consumption model;

[0134] See also Figure 4 As shown, it is a flow chart of step S4 of an embodiment of the present invention; specifically, in step S4, it includes:

[0135] Step S41, designing the task utility function of the edge user in the target time slot based on the service selection model;

[0136] Step S42, constructing an optimization objective function based on the task utility function and the energy consumption model;

[0137] Step S43, constructing optimization constraints based on the system model, the data transmission model and the service selection model;

[0138] Step S44: determining an overall optimization target based on the optimization objective function and the optimization constraints.

[0139] In a specific embodiment, assuming that n (t) = {0,1} indicates whether the task of edge user n in time slot t is a high priority task. This can be determined by the task priority P n (t) and the priority boundary a. At this time, the task utility function of edge user n in time slot t can be expressed as follows:

[0140]

[0141] The main goal of this embodiment is to obtain a solution for task allocation and resource management while reducing system energy consumption, so as to maximize the utility of edge user tasks. The overall optimization goal can be expressed as follows:

[0142]

[0143] Among them, ψ1 and ψ2 represent the proportion of edge user task service utility and energy consumption in the final optimization target, respectively. d is a constant used to measure whether the user is within the service range of the drone. The C2 expression indicates that the total task completion delay of each edge user needs to be within the time slot length. The C3 expression indicates that the transmission power of all edge users and drones needs to be less than the upper limit P. n and P m ,The C4 expression indicates that the service selection of each edge user and the ,limitation of the computing capacity of each UAV follow Equation (9) and Equation (13), respectively.

[0144] Step S5, constructing an actor network and a target dual critic network based on the overall optimization goal and the diffusion model, and updating parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

[0145] See also Figure 5 As shown, it is a flow chart of step S5 of an embodiment of the present invention; specifically, in the step S5, it includes:

[0146] Step S51, determining the state space, action space and reward function of the Markov decision process based on the overall optimization objective and the reverse diffusion process of the diffusion model;

[0147] Step S52, constructing an actor network based on the state space, the action space and the reward function;

[0148] Step S53: construct a target dual critic network, design update gradients for the target dual critic network and the actor network respectively, and update parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

[0149] The present invention converts the optimization problem into a Markov decision problem and introduces the reverse diffusion process into the decision-making process, which can reduce the computational complexity, adapt to complex and changing environments, improve system performance, and achieve efficient and reasonable resource allocation.

[0150] In a specific embodiment, in order to solve the AIGC resource allocation problem in a multi-UAV assisted edge computing network, this embodiment converts the optimization problem into a Markov decision problem (MDP) and introduces the reverse diffusion process into the decision-making process. The main goal of MDP is to obtain long-term system rewards according to formula (23). MDP can be represented by a five-tuple <S, A, P, R, γ>, where S represents the state space of the system, A represents the action space of the resource allocation decision, P represents the transition function of the system state, R represents the system reward function, γ∈[0,1] represents the discount factor, and the specific composition of MDP is as follows:

[0151] State space: The state of the system at time slot t is S(t) = [S MU (t),S ASP (t)], the first component S MU (t) contains the task information of all edge users, considering that edge user n has a certain AIGC computing capacity g n (t), therefore, it needs to be added to S MUIn order to reduce the computational complexity, this embodiment converts the location of each edge user into the drone number i of the drone service area to which it belongs. n (t) and its distance u from the drone n (t), and add it to S MU (t); the second component S ASP (t) includes the AIGC computing capacity C of each drone m (t) and the AIGC computing capacity of the edge server C0(t);

[0152] Action space: Based on the aforementioned service selection model, this embodiment defines the action at time slot t as A(t)={S n (t)|n=1,2,3,...,N}, where S n (t) is the user service selection sequence defined above. However, considering the exponential growth of the action space that may be brought about by the increase in the number of users, this embodiment changes the user service selection sequence S n (t) is redefined as a scalar S with size [-1,1] n (t), and reduce the dimension of the action space by mapping the interval into M+1 subintervals corresponding to M+1 possible task choices;

[0153] Reward function: According to the optimization formula (23), the reward designed in this embodiment is defined as:

[0154]

[0155] Among them, P C <0 represents the fixed penalty term when the restriction in the optimization formula (23) is violated;

[0156] Reverse diffusion process: The diffusion model originated from the field of image generation. It mainly includes two processes: forward diffusion and reverse diffusion. The forward diffusion process continuously adds Gaussian noise to the image data distribution until the data finally becomes pure Gaussian noise, while the reverse diffusion process continuously removes the pure Gaussian noise to obtain the final required data distribution and complete the image generation process. However, in this embodiment, due to the dynamic complexity of the environment, there is no known optimal strategy that can be used as the input of the forward diffusion. Therefore, this embodiment only uses the reverse diffusion process in the diffusion model, and uses the powerful ability of the diffusion model to fit complex dynamic environments to generate a strategy for dynamic resource allocation. It sets the known {η k |k=0,1,...,K} represents the hyperparameter of Gaussian distribution in the K-step forward diffusion process, a K~N(0,I) is the result of K-step forward diffusion, which is also pure Gaussian noise. It obeys a Gaussian distribution with an expectation of 0 and a variance of I, where N represents the Gaussian distribution and I represents the unit matrix. The purpose of the reverse diffusion process is to convert a K Step by step, we can finally get a0 as the resource allocation strategy of the system. K In the case of a K Denoising to get a K-1 The probability distribution of is proved to obey the following Gaussian distribution with parameters θ:

[0157]

[0158] in,

[0159]

[0160] in, μ θ (a K ,K,s) is the expected function related to the parameter θ, s is the state of the system at this time. According to formula (25), it is still difficult to get the system state from the simple Gaussian noise a K Continuously denoising to obtain the required strategy a0, according to Bayes' theorem, we can get the following distribution:

[0161]

[0162] Among them, the expected function is:

[0163]

[0164] Based on the reparameterization technique, the above expectation function can be rewritten as follows:

[0165]

[0166] Among them, K ~N(0,I) represents Gaussian noise, distribution function P θ (a K-1 ∣a K ) can also be rewritten as in formula (29):

[0167]

[0168] In this way, we only need to use the neural network to θ (a K ,K,s) can obtain the distribution function P θ (a K-1 ∣a K ), so that we can use this distribution to get the expected function from the Gaussian noise a KThe final required action a0 is obtained by gradual denoising. According to the reparameterization technique, the target action a0 can be expressed as:

[0169]

[0170] Among them, ò~N(0,I), ⊙ represents the matrix Hadamard product operation;

[0171] Algorithm framework: using π θ represents the AIGC resource allocation strategy, which is composed of the actor network composed of the above inverse diffusion model, by predicting the noise ò θ , pure Gaussian noise a K ~N(0,I) obtains the required action at this time through reverse diffusion K steps. According to formula (23), the main goal of the algorithm can be obtained as:

[0172]

[0173] Among them, π * represents the optimal resource allocation strategy, represents expectation, a t ~π θ (·|s t ) indicates action a t According to the strategy π θ When the input is state s t , Y represents the weight of the entropy regularization term. Different from the traditional reinforcement learning algorithm, this embodiment also attempts to maximize the entropy value of the strategy action to enhance the exploratory nature of the strategy to adapt to the complex and changing environment. Under normal circumstances, if the action entropy is high, it means that the strategy has a uniform probability of selecting all actions in a certain state. On the contrary, if the strategy always selects a few or a certain action, the action entropy is low. In the reinforcement learning process, the strategy is improved by constantly interacting with the environment to explore. If the strategy is too certain and only selects a few actions, it is easy to fall into a local optimal solution and the generalization ability will be poor. The entropy regularization term is expressed as follows:

[0174]

[0175] In reinforcement learning, the Q-value function is used to evaluate the quality of an action. The Q-value function can be expressed as follows:

[0176]

[0177] Where D represents the experience replay buffer, s t+1 ~D represents s t+1 is stored from the experience replay buffer (s t ,a t ,r t ,s t+1) sequence, Φ represents the critic network parameter used for Q-value estimation. A single critic network to estimate Q-value will have the problem of value estimation bias, which will lead to an unstable strategy learning process. In order to reduce the estimation error of Q-value, this embodiment adopts the structure of dual critic network. In fact, it is taken from the smaller estimate of the Q value of the two critic networks. Φ1 and Φ2 are the parameters of the two critic networks respectively. Therefore, the state value function can be expressed as follows:

[0178]

[0179] In order to ensure the stability of training iterations, this embodiment uses a target critic network and a target action network, with parameters and It means that the update of the critic network follows the following gradient:

[0180]

[0181] Among them, Φ represents the gradient of Φ, f t represents the timing difference target and can be expressed as follows:

[0182]

[0183] γ represents the discount factor, and the update of the actor network follows the following gradient:

[0184]

[0185] According to the soft update mechanism, the parameters of the target critic network and the target action network are updated as follows:

[0186]

[0187] Among them, ω represents the soft update parameter, which is used to control the update speed of the target network;

[0188] The actor network based on the reverse diffusion process obtains the environmental state of the current time slot according to the system model, generates and executes actions after K steps of reverse diffusion, obtains the corresponding reward and the state of the next time slot, and stores them in the experience replay buffer. After looping the above process A times, a certain combination of <state, action, reward, next state> is sampled from the experience replay buffer, and the parameters of the critic network and the actor network are updated according to the gradients of formulas (36), (37) and (38), and then the parameters of the target critic network and the actor network are updated according to formula (39). Repeat all the above processes B times to obtain the actor network corresponding to the final required allocation strategy. The specific algorithm flow includes:

[0189] S001: Initialize the experience replay buffer D, the critic network parameters Φ1, Φ2, and the actor network parameters θ;

[0190] S002: Initialize target actor network parameters and target critic network parameters

[0191] S003: for training round e=1,…,E do;

[0192] S004: Reset the multi-unmanned assisted edge computing network environment;

[0193] S005: for training steps t=1,…,T do;

[0194] S006: Get system status s t And randomly generate a Gaussian noise a K ~N(0,I);

[0195] S007: for denoising steps i = K,…,1do;

[0196] S008: Using actor networks to predict noise ò θ ;

[0197] S009: Calculate the expectation of the reverse diffusion process according to formula (30);

[0198] S010: Calculate a according to formula (31) i-1 ;

[0199] S011: end for;

[0200] S012: Execute action a t =a i And get the state s of the next time slot t+1 and reward r t ;

[0201] S013: t ,a t ,r t ,s t+1 ) is stored in the experience playback buffer D;

[0202] S014: end for;

[0203] S015: Randomly sample from the experience replay buffer D with a size of B = {(s t ,a t ,r t ,s t+1 )} sequence;

[0204] S016: Update the actor and critic network parameters using B according to formulas (36), (37) and (38);

[0205] S017: Update the target network parameters according to formula (39);

[0206] S018: end for;

[0207] The present invention can design a system model and a service selection according to the needs of edge users by constructing a system model and a data transmission model, so as to realize the simultaneous allocation of multiple tasks in one time slot and improve service efficiency. By constructing a service selection model and designing different priorities for content generation tasks, different system benefits brought by different tasks can be distinguished. Since the energy consumption required by the GPU to complete the content generation task is much higher than the energy consumed by signal transmission, by constructing an energy consumption model, only the energy consumption required for task completion is considered, which can reduce the amount of data processing and improve processing efficiency. By designing an overall optimization goal and clarifying the goals of the system, it is possible to ensure that a solution for task allocation and resource management is obtained while reducing the energy overhead of the system, so as to maximize the utility of edge user tasks. By constructing an actor network and a target dual critic network based on the overall optimization goal and the diffusion model, it is possible to generate a strategy for dynamic resource allocation with the help of the powerful ability of the diffusion model to fit complex dynamic environments, reduce the estimation error of the Q value, and the optimal resource allocation strategy generated in this way can achieve efficient resource allocation using limited network resources.

[0208] See also Figure 8-Figure 9 , Figure 8 A comparison chart of the average reward per task of the drone edge computing network resource allocation method based on the diffusion model in an embodiment of the present invention and the current mainstream reinforcement learning algorithm proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies; Fig. 9 The figure is a comparison chart of the average reward per training task of the drone edge computing network resource allocation method based on the diffusion model in an embodiment of the present invention and the current mainstream reinforcement learning algorithms proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies; it can be seen that when the user arrival rate is low, the effects of all methods are similar, but when the user arrival rate increases, the complexity of the environment will increase significantly due to the increase in the number of users. Under this condition, the method proposed in the present invention has a strong fitting ability for complex environments due to the introduction of the reverse diffusion model, and the effect will be significantly better than the traditional method.

[0209] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A method for allocating network resources for UAV edge computing based on a diffusion model, characterized in that: include: Step S1, obtaining the content generation task information randomly generated by any edge user in the multi-UAV-assisted edge computing network in the target time slot and the service selection of the edge user, and constructing a system model and a data transmission model corresponding to each service selection; The multi-drone assisted edge computing network includes a plurality of drones, a plurality of edge users, and a plurality of edge servers; the content generation task information includes the task data size, the computing load required for the task, the edge user priority, and the maximum service delay limit that the edge user can tolerate; Step S2, calculating the corresponding task completion delay according to the system model and each of the data transmission models, and building a service selection model based on the completion delay of each task and the priority of the content generation task; Step S3, constructing an energy consumption model based on the energy consumed by the edge user, the edge server, and the GPU of the drone to complete the content generation task in the target time slot; Step S4, determining an overall optimization target based on the system model, the service selection model, the data transmission model and the energy consumption model; Step S5, constructing an actor network and a target dual critic network based on the overall optimization goal and the diffusion model, and updating parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

2. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 1 is characterized in that: In the step S1, it includes: Step S11, constructing a system model based on the content generation task information; Step S12, determining a corresponding channel type based on the service selection of the edge user, and obtaining additional interference corresponding to each channel type; Step S13, determining a channel coefficient corresponding to each channel type according to the system model; Step S14: constructing a data transmission model corresponding to each channel type according to each of the channel coefficients and the additional interference corresponding to each of the channel types.

3. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 2 is characterized in that: In the step S2, it includes: Step S21, constructing a service selection sequence based on the service selection of the edge user, and designing corresponding service constraints; Step S22, calculating the task completion delay corresponding to each service selection based on the service selection sequence and the system model; Step S23: constructing a service selection model according to the completion delay of each task and the priority of the content generation task.

4. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 3 is characterized in that: In the step S22, it includes: The task calculation delay and task transmission delay corresponding to each service selection are determined based on the system model, the service selection sequence and the corresponding data transmission model, and the task completion delay corresponding to each service selection is calculated according to the task calculation delay and the task transmission delay.

5. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 4 is characterized in that: In the step S23, it includes: The utility function corresponding to each task priority is determined according to the task completion delay and the content generation task priority, and a service selection model is constructed according to each utility function.

6. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 5 is characterized in that: In the step S3, it includes: Based on the GPU power of the edge user, the edge server and the drone in the target time slot and the task calculation delay corresponding to each service selection, the energy consumed by the GPU of the edge user, the edge server and the drone to complete the content generation task in the target time slot is calculated to construct an energy consumption model.

7. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 6 is characterized in that: In the step S4, it includes: Step S41, designing the task utility function of the edge user in the target time slot based on the service selection model; Step S42, constructing an optimization objective function based on the task utility function and the energy consumption model; Step S43, constructing optimization constraints based on the system model, the data transmission model and the service selection model; Step S44: determining an overall optimization target based on the optimization objective function and the optimization constraints.

8. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 7 is characterized in that: In the step S5, it includes: Step S51, determining the state space, action space and reward function of the Markov decision process based on the overall optimization objective and the reverse diffusion process of the diffusion model; Step S52, constructing an actor network based on the state space, the action space and the reward function; Step S53: construct a target dual critic network, design update gradients for the target dual critic network and the actor network respectively, and update parameters of the actor network and the target dual critic network to generate an optimal resource allocation strategy.

9. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 8 is characterized in that: In step S1, the computational load required for the task is calculated according to the number of denoising steps required for the content generation task.

10. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 9 is characterized in that: In the step S1, the service selection of the edge user includes offloading the content generation task to the drone, offloading the content generation task to the adjacent drone, and offloading the content generation task to the edge server.

Citation Information

Patent Citations

  • Diffusion model-based request allocation method in cloud edge collaborative system

    CN117938960A

  • Diffusion model-based request allocation method for cloud edge collaborative architecture

    CN118509491A

  • Large-capacity low-energy-consumption unmanned aerial vehicle covert communication method

    CN118678325A

  • AIGC edge-end collaborative reasoning method based on diffusion model

    CN119106723A