A method for allocating resources of a UAV edge computing network based on a diffusion model
By constructing a system model and a data transmission model, and combining the diffusion model with the reinforcement learning algorithm, an optimal resource allocation strategy is generated, which solves the challenge of resource allocation in the UAV-assisted edge computing network and achieves efficient support and resource utilization for different types of AIGC tasks.
Patent Information
- Application Number
- CN202411922092.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing technologies are unable to effectively utilize limited network resources to efficiently support different types of artificial intelligence generated content (AIGC) tasks, resulting in challenges in resource allocation in drone-assisted edge computing networks.
A resource allocation method for UAV edge computing network based on diffusion model is adopted. By constructing system model, data transmission model, service selection model and energy consumption model, and combining actor network and target dual critic network, the optimal resource allocation strategy is generated to optimize task allocation and resource management.
It achieves efficient support for different types of AIGC tasks under limited network resources, improves service efficiency and resource utilization, reduces energy consumption, adapts to complex and changing environments, ensures that high-priority tasks are completed in time, and avoids low-priority tasks from being ignored.
Smart Images

Figure CN119946723B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of edge computing, and in particular to a method for allocating resources of an unmanned aerial vehicle edge computing network based on a diffusion model. BACKGROUND
[0002] Edge computing network resource allocation refers to providing efficient allocation strategies for limited energy, computing, and communication resources in the network through linear programming, heuristic algorithms, game theory, reinforcement learning, and other methods to meet performance indicators such as energy optimization and computing resource utilization. Traditional network infrastructure is limited by flexibility and cost, and cannot meet the growing demand for edge terminal computing. Unmanned aerial vehicle-assisted edge computing utilizes the high flexibility and mobility of unmanned aerial vehicles to deploy computing resources on unmanned aerial vehicles to provide low-latency and efficient computing services for terminal edge devices.
[0003] With its own flexibility and low cost, unmanned aerial vehicle (UAV) assisted edge computing network, as an important support means for the integration of space, air, and ground networks in the sixth generation mobile communication system, is widely used in scenarios such as smart cities and smart industries. However, with the advent of the era of large models represented by ChatGPT, the surge in demand for artificial intelligence-generated content (AIGC) by terminal devices has further increased the network burden, and the large number and variety of terminal computing tasks have further challenged unmanned aerial vehicle-assisted edge computing networks.
[0004] Currently, there is little research on multi-unmanned aerial vehicle-assisted edge computing networks that provide AIGC services for terminal users, and most research focuses on single unmanned aerial vehicles as AIGC service providers without considering the quality of service requirements for different types of AIGC tasks. Therefore, how to use efficient resource allocation under limited network resources to support different types of AIGC tasks has become a technical problem that limits the development of unmanned aerial vehicle-assisted edge computing networks. SUMMARY
[0005] To this end, the present application provides a method for allocating resources of an unmanned aerial vehicle edge computing network based on a diffusion model to overcome the problem that existing technologies cannot use efficient resource allocation under limited network resources to support different types of AIGC tasks.
[0006] To achieve the above purpose, the present application provides a method for allocating resources of an unmanned aerial vehicle edge computing network based on a diffusion model, comprising:
[0007] In step S1, content generation task information randomly generated by any edge user in a target time slot of a multi-UAV assisted edge computing network and service selection of the edge user are acquired, a system model and a data transmission model corresponding to each service selection are constructed;
[0008] The multi-UAV assisted edge computing network includes a plurality of UAVs, a plurality of edge users and a plurality of edge servers; the content generation task information includes task data size, required computing load of the task, edge user priority and maximum service delay upper limit that can be tolerated by the edge user;
[0009] In step S2, a task completion delay corresponding to each service selection is calculated according to the system model and the data transmission model, and a service selection model is constructed based on the task completion delay and the content generation task priority;
[0010] In step S3, an energy consumption model is constructed based on energy consumed by GPUs of the edge user, the edge server and the UAV in completing the content generation task in the target time slot;
[0011] In step S4, an overall optimization target is determined based on the system model, the service selection model, the data transmission model and the energy consumption model;
[0012] In step S5, an actor network and a target double critic network are constructed based on the overall optimization target and a diffusion model, and the actor network and the target double critic network are parameter updated to generate an optimal resource allocation strategy.
[0013] Further, in the step S1, the following steps are included:
[0014] In step S11, a system model is constructed based on the content generation task information;
[0015] In step S12, a channel type corresponding to the service selection of the edge user is determined, and additional interference corresponding to each channel type is acquired;
[0016] In step S13, a channel coefficient corresponding to each channel type is determined according to the system model;
[0017] In step S14, a data transmission model corresponding to each channel type is constructed according to the channel coefficient and the additional interference corresponding to each channel type.
[0018] Further, in the step S2, the following steps are included:
[0019] In step S21, a service selection sequence is constructed based on the service selection of the edge user, and a corresponding service constraint is designed;
[0020] Step S22, calculating a task completion time delay corresponding to each service selection based on the service selection sequence and the system model;
[0021] Step S23, constructing a service selection model according to each task completion time delay and the content generation task priority.
[0022] Further, in the step S22, comprising:
[0023] Based on the system model, the service selection sequence and the corresponding data transmission model, determining a task calculation time delay and a task transmission time delay corresponding to each service selection, and calculating a task completion time delay corresponding to each service selection according to each task calculation time delay and each task transmission time delay.
[0024] Further, in the step S23, comprising:
[0025] According to each task completion time delay and the content generation task priority, determining a utility function corresponding to each task priority, and constructing a service selection model according to each utility function.
[0026] Further, in the step S3, comprising:
[0027] Based on the GPU power of the edge user, the edge server and the UAV in the target time slot and the task calculation time delay corresponding to each service selection, calculating the energy consumed by the edge user, the edge server and the GPU of the UAV in the target time slot to complete the content generation task, to construct an energy consumption model.
[0028] Further, in the step S4, comprising:
[0029] Step S41, designing a task utility function of the edge user in the target time slot based on the service selection model;
[0030] Step S42, constructing an optimization objective function based on the task utility function and the energy consumption model;
[0031] Step S43, constructing an optimization constraint based on the system model, the data transmission model and the service selection model;
[0032] Step S44, determining an overall optimization objective based on the optimization objective function and the optimization constraint.
[0033] Further, in the step S5, comprising;
[0034] Step S51, determining a state space, an action space and a reward function of a Markov decision process based on the overall optimization objective and a reverse diffusion process of a diffusion model;
[0035] Step S52, constructing an actor network based on the state space, the action space and the reward function;
[0036] Step S53, constructing a target double critic network, designing the update gradient of the target double critic network and the actor network respectively, and performing parameter update on the actor network and the target double critic network to generate an optimal resource allocation strategy.
[0037] Further, in the step S1, the computing load required by the content generation task is calculated according to the number of denoising steps required by the content generation task.
[0038] Further, in the step S1, the service selection of the edge user includes offloading the content generation task to a UAV, offloading the content generation task to a neighboring UAV, and offloading the content generation task to an edge server.
[0039] Compared with the prior art, the beneficial effects of the present application are that the present application can construct a system model and a data transmission model, design a system model and a service selection according to the needs of an edge user, allocate multiple tasks in one time slot, and improve service efficiency. By constructing a service selection model, different priorities are designed for content generation tasks, and different system benefits brought by different tasks can be distinguished. Since the energy consumption required by GPU to complete the content generation task is much higher than the energy consumption of signal transmission, by constructing an energy consumption model, only the energy consumption required for task completion is considered, the data processing amount is reduced, and the processing efficiency is improved. By designing an overall optimization target, the target of the system is clear, and the scheme of task allocation and resource management can be obtained under the condition of reducing system energy consumption, so as to maximize the task utility of the edge user. By constructing an actor network and a target double critic network based on the overall optimization target and the diffusion model, the powerful ability of the diffusion model to fit complex dynamic environment can be used to generate a strategy for dynamic resource allocation, reduce the estimation error of Q value, and generate an optimal resource allocation strategy that can realize efficient resource allocation with limited network resources.
[0040] Further, the present application designs a corresponding data transmission model according to the service selection of the edge user, considers the additional interference of each channel type, reasonably allocates computing resources, saves network resources, and realizes precise and efficient use of limited network resources.
[0041] Further, the present application can reasonably allocate service selection according to computing resources by constructing a service selection sequence and service constraints, and improve service capability.
[0042] Further, the present application designs the utility function according to the content selection priority, which can ensure the non-negative profit brought by the timely completion of high-priority tasks, and avoid long-term neglect of low-priority tasks.
[0043] Further, the present application converts the optimization problem into a Markov decision problem, and introduces a reverse diffusion process into the decision-making process, which can reduce the computational complexity, adapt to complex and variable environments, improve the system performance, and realize efficient and reasonable resource allocation. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The flowchart of the UAV edge computing network resource allocation method based on the diffusion model of the embodiment of the present application is shown in the figure.
[0045] Figure 2 The flowchart of step S1 of the embodiment of the present application is shown in the figure.
[0046] Figure 3 The flowchart of step S2 of the embodiment of the present application is shown in the figure.
[0047] Figure 4 The flowchart of step S4 of the embodiment of the present application is shown in the figure.
[0048] Figure 5 The flowchart of step S5 of the embodiment of the present application is shown in the figure.
[0049] Figure 6 The flowchart of the service selection model for providing content generation tasks of the embodiment of the present application is shown in the figure.
[0050] Figure 7 The flowchart of generating the optimal resource allocation strategy of the embodiment of the present application is shown in the figure.
[0051] Figure 8 The comparison chart of the average reward per task of the UAV edge computing network resource allocation method based on the diffusion model of the embodiment of the present application and the current mainstream reinforcement learning algorithm proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies is shown in the figure.
[0052] Figure 9 The comparison chart of the average reward per training task of the UAV edge computing network resource allocation method based on the diffusion model of the embodiment of the present application and the current mainstream reinforcement learning algorithm proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies is shown in the figure. DETAILED DESCRIPTION
[0053] In order to make the purpose and advantages of the present application clearer and more apparent, the present application will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.
[0054] The preferred embodiments of the present application will be described below with reference to the drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present application and not to limit the protection scope of the present application.
[0055] It should be noted that, in the description of the present application, the terms of direction or position relationship such as "upper", "lower", "left", "right", "inner", "outer" and the like are based on the direction or position relationship shown in the drawings, which is only for the convenience of description and does not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0056] In addition, it should also be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connection" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0057] Please refer to Figure 1 、 Figure 6 、 Figure 7 , Figure 1 is a flow chart of the method for allocating resources of a UAV edge computing network based on a diffusion model according to an embodiment of the present application; Figure 6 is a flow chart of a service selection model for providing a content generation task according to an embodiment of the present application; Figure 7 is a flow chart of generating an optimal resource allocation strategy according to an embodiment of the present application. The present application provides a method for allocating resources of a UAV edge computing network based on a diffusion model, which comprises:
[0058] Step S1, obtaining content generation task information randomly generated by any edge user in a target time slot multi-UAV assisted edge computing network and service selection of the edge user, constructing a system model and a data transmission model corresponding to each service selection;
[0059] The multi-UAV assisted edge computing network comprises a plurality of UAVs, a plurality of edge users and a plurality of edge servers; the content generation task information comprises task data size, required computing load of the task, edge user priority and maximum service delay upper limit that can be tolerated by the edge user;
[0060] Please refer to Figure 2 , which is a flow chart of step S1 according to an embodiment of the present application; specifically, in the step S1, it comprises:
[0061] Step S11, constructing a system model based on the content generation task information;
[0062] Step S12, determining corresponding channel types based on the service selection of the edge user, and obtaining additional interference corresponding to each channel type;
[0063] Step S13, determining channel coefficients corresponding to each channel type according to the system model;
[0064] Step S14, constructing a data transmission model corresponding to each channel type according to each channel coefficient and the additional interference corresponding to each channel type.
[0065] Specifically, in the step S1, the computing load required by the content generation task is calculated according to the number of denoising steps required by the content generation task.
[0066] Specifically, in the step S1, the service selection of the edge user includes offloading the content generation task to a UAV, offloading the content generation task to a neighboring UAV, and offloading the content generation task to an edge server.
[0067] In one specific embodiment, a multi-UAV assisted edge computing network is composed of M UAVs, N edge users and 1 edge server. In order to facilitate system analysis, time T is evenly divided into time slots with duration τ, and each UAV is equipped with an omnidirectional antenna. In the target time slot t, each edge user n will randomly generate a content generation task, and a task system model D n (t) of the edge user in the target time slot t can be constructed n (t), c n (t), μ n (t), p n (t)}, where d n (t) is the task data size, c n (t) is the computing load required by the content generation task, μ n (t) is the maximum service delay upper limit that the edge user n can tolerate, p n (t) is the priority of the edge user n. Each UAV acts as an AIGC service provider in this model, and can also act as a relay to offload tasks within its service range W to other UAVs or edge servers.
[0068] It can be understood that, considering that the task result is usually much smaller in data volume than the task itself, the present application ignores the downlink backhaul of the task result. Starting from the role played by the UAV, the service selection of the edge user is to unload the content generation task to the UAV, and the corresponding channel type is the air-ground channel between the edge user and the UAV; the service selection of the edge user is to unload the content generation task to the adjacent UAV, and the corresponding channel type is the air-ground channel between the edge user and the UAV plus the air-air channel between the UAVs; the service selection of the edge user is to unload the content generation task to the edge server, and the corresponding channel type is the air-ground channel between the edge user and the UAV plus the air-ground channel between the UAV and the edge server;
[0069] For the air-ground channel, considering that each UAV can serve multiple edge users at the same time by using multi-user multi-input multi-output technology, the three-dimensional coordinates of the UAV m at the target time slot t can be represented as follows: wherein, and represent the coordinates of the X, Y and Z axes respectively, and the position of the edge user n can be represented as: q n (t)=[x n (t),y n (t),0] T The present application considers both the line-of-sight (LoS) channel and the non-line-of-sight (NLoS) channel in the air-ground communication link. The probability of the line-of-sight channel between the edge user n and the UAV m can be represented as:
[0070]
[0071] wherein, both a and are constants determined by the environment, d n,m (t) is the distance between the edge user n and the UAV m at the time slot t, and the probability of the non-line-of-sight channel can be obtained from formula (1) Thus, the channel gain between the two can be obtained as:
[0072]
[0073] wherein, λ0 is the channel gain per unit distance, represents the path loss exponent, and ζ∈(0, 1) is an additional attenuation factor under the non-line-of-sight channel, therefore, the average channel gain of the air-ground channel can be represented as:
[0074]
[0075] g n,m (t) represents the channel coefficient between the edge user n and the UAV m, which can be expressed as:
[0076]
[0077] wherein, is a small-scale fading coefficient, which is subject to the expectation , thus, the data transmission rate of the edge user n and the UAV m at time slot t can be expressed as follows:
[0078]
[0079] wherein, B GA represents the channel bandwidth between the ground-air channels, P n (t) is the transmission power of the edge user n at time slot t, and are an additional Gaussian noise and the interference caused by other edge users j within the coverage range of the UAV m, respectively;
[0080] For the air-air channel, in order to fully utilize the computing resources of all UAVs, the present application sets that the UAV can receive the AIGC task of the ground edge user and forward it to other UAVs, considering the full-duplex communication capability of each UAV, the receiving and forwarding task is considered to occur at the same time, in order to simplify the design, the present application assumes that the data transmission path loss between the UAVs can be represented by the free space path loss (FSPL), because the UAVs usually have less obstruction and are mainly LoS channels. Based on the Friis transmission formula, the path loss between the UAV m and the UAV m' can be expressed as:
[0081] L m,m′ (t) = 20log 10 (d m,m′ (t)) + 20log 10 (f) + 32.45, (6)
[0082] wherein, d m,m' (t) is the distance between the UAV m and the UAV m' at time slot t, measured in kilometer (km) as the unit. f represents the signal frequency, measured in megahertz (MHz) as the unit, thus, the data transmission rate of the air-air channel can be expressed as:
[0083]
[0084] wherein, represents the Gaussian noise, B AA represents the channel bandwidth between the air-air channels, P m(t) represents the signal power of the unmanned aerial vehicle m at the time slot t;
[0085] For the air-ground channel model, similar to the air-ground channel model, the data transmission rate between the unmanned aerial vehicle m and the edge server can be represented as:
[0086]
[0087] wherein, represents the Gaussian noise, B AG represents the channel bandwidth between the air-ground channel, and the position of the edge server is represented as q0=[x0,y0,0] T .
[0088] The application can save network resources and realize precise and efficient use of limited network resources by designing corresponding data transmission models according to service selection of edge users, considering additional interference of each channel type, and reasonably allocating computing resources.
[0089] Step S2, calculating corresponding task completion time delay according to the system model and each data transmission model, and constructing a service selection model based on each task completion time delay and content generation task priority;
[0090] Please refer to Figure 3 , which is a flowchart of step S2 of the embodiment of the application. Specifically, in the step S2, it includes:
[0091] Step S21, constructing a service selection sequence based on the service selection of the edge user, and designing a corresponding service constraint;
[0092] Step S22, calculating the task completion time delay corresponding to each service selection based on the service selection sequence and the system model;
[0093] Specifically, in the step S22, it includes:
[0094] Based on the system model, the service selection sequence and the corresponding data transmission model, the task calculation time delay and the task transmission time delay corresponding to each service selection are determined, and the task completion time delay corresponding to each service selection is calculated according to each task calculation time delay and each task transmission time delay.
[0095] Step S23, constructing a service selection model according to each task completion time delay and the content generation task priority.
[0096] Specifically, in the step S23, it includes:
[0097] According to the task completion delay and the content generation task priority, a utility function corresponding to each task priority is determined, and a service selection model is constructed according to each utility function.
[0098] The present application can reasonably allocate service selection according to computing resources by constructing a service selection sequence and service constraints, and improve service capability.
[0099] In a specific embodiment, because each edge user has certain computing capability, the edge user can choose to perform the content generation task locally or upload the content generation task to the unmanned aerial vehicle and the edge server with sufficient computing resources to complete the calculation. The service selection sequence of the edge user n at time slot t can be represented as: Wherein, represents that the user chooses to perform the calculation locally, represents that the task is executed at the edge server, and the rest represents that the edge user offloads the task to the designated unmanned aerial vehicle for calculation. The task of each edge user can only be calculated at one place at time slot t, so the service selection sequence of each edge user n follows the following constraints:
[0100]
[0101] F n (t) represents the task completion delay of the edge user n at time slot t. When the edge user n chooses to perform the calculation locally, the task calculation delay can be represented as follows:
[0102]
[0103] Wherein, is the time required for each step of denoising when completing the content generation task at the edge user n. In this case, the task completion delay can be represented as:
[0104]
[0105] When the edge user n chooses to offload the content generation task to the unmanned aerial vehicle, the task calculation delay can be represented as:
[0106]
[0107] Wherein, k=1,...,M, the calculation capability of the unmanned aerial vehicle k is represented by using the time required for each step of denoising. Considering that the computing resources of each unmanned aerial vehicle are limited at time slot t, the following constraints need to be met:
[0108]
[0109] Wherein, C kTo represent the de-noising capability of the UAV k, it is assumed that the edge user n is in the service range of the UAV m at the time slot t, and the edge user n can either offload the task to the UAV m or offload it to other UAVs m' through the UAV m. In consideration of the full-duplex communication capability of all the UAVs, the task transmission delay can be represented as follows:
[0110]
[0111] Therefore, the task completion delay can be represented as:
[0112]
[0113] If the computing resource required by the content generation task is too large, it can be transmitted to the edge server for calculation in the manner of UAV relay. At this time, the task calculation delay can be represented as:
[0114]
[0115] wherein, is the de-noising capability of the edge server, and the computing resource limitation of the edge server is not considered in the present application because the edge server has sufficient computing resources. At this time, the task transmission delay can be represented as:
[0116]
[0117] Therefore, the task completion delay can be represented as:
[0118]
[0119] In consideration of the different utilities brought by different priority tasks, the present embodiment defines the service priority of each task, and uses a to represent the dividing line between high and low priority. When the task priority P n (t) of the edge user n is less than a, it is a low priority task; otherwise, it is a high priority task. The utility function of the high priority task is defined as follows:
[0120]
[0121] wherein, is the utility coefficient of completing different high priority tasks, and μ max is the highest priority; P H <0 is the penalty constant term for not completing the high priority task in time. Different from other applications, the present embodiment can specifically distinguish the different system gains brought by the completion of different tasks in the high priority interval by using the utility coefficient δ n (t). Meanwhile, by using log2(1+μ n (t)-F nThe utility function of the high-priority task can ensure non-negative benefits brought by timely completion of the high-priority task, and can effectively reflect the benefits brought by completion of the high-priority task within a tolerable time delay range, which increases with an increase in the time of completion in advance, but the growth rate slows down, which conforms to the actual long-tail effect.
[0122] The utility function of the low-priority task is defined as follows:
[0123]
[0124] wherein, is a utility coefficient of completion of different low-priority tasks.
[0125] Because the low-priority task is usually less sensitive to time, the embodiment sets a fixed reward value R L > 0, so as to avoid long-term neglect of the low-priority task. L > P H to punish the low-priority task that is not completed in time.
[0126] The utility function is designed according to the content selection priority, which can ensure non-negative benefits brought by timely completion of the high-priority task, and can avoid long-term neglect of the low-priority task.
[0127] In step S3, the energy consumption model is constructed based on the energy consumed by the edge user, the edge server and the GPU of the UAV in completing the content generation task in the target time slot.
[0128] Specifically, in the step S3, the following steps are included.
[0129] The energy consumed by the edge user, the edge server and the GPU of the UAV in completing the content generation task in the target time slot is calculated based on the GPU power of the edge user, the edge server and the UAV in the target time slot and the task computation time delay corresponding to each service selection, so as to construct the energy consumption model.
[0130] In a specific embodiment, considering that the energy consumed by the GPU to complete the content generation task is much higher than the energy consumed by signal transmission, the embodiment only considers the energy consumed by the task completion in order to simplify the model, and the energy consumed by the system in the time slot t can be represented as:
[0131]
[0132] wherein, k ES and GPU power of edge users, edge servers and UAVs respectively.
[0133] Step S4, determining an overall optimization target based on the system model, the service selection model, the data transmission model and the energy consumption model;
[0134] Please refer to Figure 4 The flow chart of step S4 of the embodiment of the application is shown in the figure, and specifically, in the step S4, the following steps are included:
[0135] Step S41, designing a task utility function of the edge user in a target time slot based on the service selection model;
[0136] Step S42, constructing an optimization objective function based on the task utility function and the energy consumption model;
[0137] Step S43, constructing an optimization constraint based on the system model, the data transmission model and the service selection model;
[0138] Step S44, determining an overall optimization target based on the optimization objective function and the optimization constraint.
[0139] In one specific embodiment, it is assumed that υ n (t)={0,1} represents whether the task of the edge user n in the time slot t is a high priority task, which can be determined by the task priority P n (t) and the priority boundary a, at this time, the task utility function of the edge user n in the time slot t can be represented as follows:
[0140]
[0141] The main goal of the embodiment is to obtain a task allocation and resource management scheme in the case of reducing system energy overhead, so as to maximize the edge user task utility, and the overall optimization target can be represented as follows:
[0142]
[0143] Wherein, ψ1 and ψ2 represent the proportion of the edge user task service utility and the energy consumption overhead in the final optimization target respectively, d is a constant used to measure whether the user is within the service range of the UAV, C2 expression represents that the total task completion delay of each edge user needs to be within the time slot length, C3 expression represents that the transmission power of all edge users and UAVs needs to be less than the upper limit P n and P m , C4 expression represents that the service selection of each edge user and the computing capacity of each UAV follow formula (9) and formula (13) respectively.
[0144] Step S5, based on the overall optimization target and the diffusion model, constructing an actor network and a target double critic network, and updating parameters of the actor network and the target double critic network to generate an optimal resource allocation strategy.
[0145] Referring to Figure 5 The figure is a flow chart of step S5 of the embodiment of the application; specifically, in the step S5, it includes;
[0146] Step S51, determining a state space, an action space and a reward function of a Markov decision process based on the overall optimization target and a reverse diffusion process of the diffusion model;
[0147] Step S52, constructing an actor network based on the state space, the action space and the reward function;
[0148] Step S53, constructing a target double critic network, designing an update gradient of the target double critic network and the actor network respectively, and updating parameters of the actor network and the target double critic network to generate an optimal resource allocation strategy.
[0149] The application converts the optimization problem into a Markov decision problem, and introduces a reverse diffusion process into the decision-making process, which can reduce the computational complexity, adapt to complex and variable environments, improve the system performance, and realize efficient and reasonable resource allocation.
[0150] In one specific embodiment, in order to solve the AIGC resource allocation problem in the multi-UAV assisted edge computing network, the embodiment converts the optimization problem into a Markov decision problem (MDP) and introduces a reverse diffusion process into the decision-making process. The main goal of MDP is to get long-term system rewards according to formula (23), and MDP can be represented by five tuples <S, A, P, R, γ>, S represents the state space of the system, A represents the action space of the resource allocation decision, P represents the state transition function of the system, R represents the system reward function, and γ∈[0, 1] represents the discount factor. The specific composition of MDP is as follows:
[0151] State space: the state of the system at time slot t is S(t)=[S MU (t),S ASP (t)], the first component S MU (t) contains the task information of all edge users, considering that the edge user n has a certain AIGC computing capability g n (t), therefore, it needs to be added to S MUIn (t), in order to reduce the computational complexity, the embodiment converts the position of each edge user into the UAV number i of the UAV service area to which it belongs n (t) and the distance u of the edge user from the UAV n (t) and adds it to S MU (t) in (t); the second component S ASP (t) contains the AIGC computing capacity C of each UAV m (t) and the AIGC computing capacity C0(t) of the edge server;
[0152] Action space: based on the aforementioned service selection model, the embodiment defines the action at time slot t as A(t) = {S n (t) | n = 1, 2, 3,..., N}, where S n (t) is the user service selection sequence defined as described above, however, considering that the exponential growth of the action space may be caused by the increase in the number of users, the embodiment redefines the user service selection sequence S n (t) as a scalar S n (t) with a size of [-1, 1], and by mapping this interval to M+1 sub-intervals, the dimension of the action space is reduced, corresponding to M+1 possible task selections;
[0153] Reward function: according to the optimization formula (23), the reward designed by the embodiment is defined as:
[0154]
[0155] where P C <0 represents a fixed penalty term when a violation of the restriction in the optimization formula (23) occurs;
[0156] Inverse diffusion process: the diffusion model originates from the field of picture generation, which mainly includes two processes, forward diffusion and inverse diffusion. The forward diffusion process adds Gaussian noise to the picture data distribution until the last data becomes pure Gaussian noise, while the inverse diffusion process removes the pure Gaussian noise to obtain the final data distribution needed to complete the picture generation process. However, in the embodiment, due to the dynamic complexity of the environment, there is no known optimal strategy that can be used as the input of the forward diffusion, therefore, the embodiment only uses the inverse diffusion process in the diffusion model, and by virtue of the powerful ability of the diffusion model to fit complex dynamic environments, generates strategies for dynamic resource allocation, sets known {η k | k = 0, 1,..., K} represent the hyperparameters of the Gaussian distribution in the K-step forward diffusion process, a K~N(0,I) is the result of K-step forward diffusion, which is also a pure Gaussian noise. It obeys a Gaussian distribution with an expectation of 0 and a variance of I, where N represents the Gaussian distribution and I represents the unit matrix. The purpose of the reverse diffusion process is to convert a K Gradually remove the noise, and finally get a0 as the resource allocation strategy of the system at this time. K In the case of a K Denoising to get a K-1 The probability distribution of is proved to obey the following Gaussian distribution with respect to the parameter θ:
[0157]
[0158] in,
[0159]
[0160] in, μ θ (a K ,K,s) is the expected function related to the parameter θ, s is the state of the system at this time, according to formula (25) it is still difficult to get the system from the simple Gaussian noise a K Continuously denoising to obtain the required strategy a0, according to Bayes' theorem, we can get the following distribution:
[0161]
[0162] Among them, the expected function is:
[0163]
[0164] Based on the reparameterization technique, the above expectation function can be rewritten as follows:
[0165]
[0166] Among them, K ~N(0,I) represents Gaussian noise, distribution function P θ (a K-1 ∣a K ) can also be rewritten as in formula (29):
[0167]
[0168] In this way, we only need to use the neural network to θ (a K ,K,s) can obtain the distribution function P θ (a K-1 ∣a K ), so that we can get the expected function from Gaussian noise a according to this distribution. KThe final required action a0 is obtained by step-by-step denoising, according to the reparameterization trick, the target action a0 can be expressed as:
[0169]
[0170] Wherein, ò ~ N(0, I), ⊙ represents the matrix Hadamard product operation;
[0171] Algorithm framework: use π θ To represent the AIGC resource allocation strategy, which is composed of the above inverse diffusion model composed of the actor network, through the prediction of noiseò θ , the pure Gaussian noise a K ~ N(0, I) is obtained by inverse diffusion K steps to obtain the required action at this time, according to formula (23), the main goal of the algorithm can be obtained as:
[0172]
[0173] Wherein, π * Represents the optimal resource allocation strategy, Represents the expectation, a t ~ π θ (·|s t ) indicates that the action a t Is obtained according to the strategy π θ Under the input state s t , Y represents the weight of the entropy regularization term, unlike traditional reinforcement learning algorithms, this embodiment simultaneously attempts to maximize the entropy value of the strategy action to enhance the exploratory nature of the strategy to adapt to complex and variable environments. Generally, if the action entropy is high, it means that the strategy selects all actions uniformly in a certain state, on the contrary, if the strategy always selects a few or a certain action, the action entropy will be low. In the process of reinforcement learning, the strategy improves by constantly interacting with the environment to explore. If the strategy is too certain and only selects a few actions, it is easy to fall into a local optimal solution, and the generalization ability will also be poor. The entropy regularization term is expressed as follows:
[0174]
[0175] In reinforcement learning, the Q value function is used to evaluate the goodness of the action, and the Q value function can be expressed as follows:
[0176]
[0177] Wherein, D represents the experience replay buffer, s t+1 ~ D indicates that s t+1 Is stored in the experience replay buffer (s t , a t , r t , s t+1) sequence, and represents critic network parameters for Q value estimation. A single critic network for estimating Q values can have value estimation bias problems, which can lead to an unstable policy learning process. To reduce the estimation error of Q values, the embodiment uses a double critic network structure, In fact, a smaller one of the two critic networks is taken for Q value estimation, and and are parameters of the two critic networks, respectively. Therefore, the state value function can be expressed as follows:
[0178]
[0179] To ensure the stability of the training iteration, the embodiment uses a target critic network and a target action network, which are represented by parameters and respectively. The update of the critic network follows the gradient:
[0180]
[0181] where represents the gradient with respect to, and represents the time difference target, which can be expressed as follows: Φ t
[0182]
[0183] represents a discount factor, and the update of the actor network follows the gradient:
[0184]
[0185] According to the soft update mechanism, the parameter updates of the target critic network and the target action network are as follows:
[0186]
[0187] where represents a soft update parameter used to control the update speed of the target network.
[0188] The actor network based on the reverse diffusion process obtains the environment state of the current time slot according to the system model, generates an action after K steps of reverse diffusion and executes it, obtains the corresponding reward and the state of the next time slot, and stores them into an experience replay buffer. After the above process A is repeated A times, a certain combination of <state, action, reward, next state> is sampled from the experience replay buffer, and the parameters of the critic network and the actor network are updated according to the gradients of formulas (36), (37) and (38), respectively. Then, the parameters of the target critic network and the target actor network are updated according to formula (39). The above processes are repeated B times to obtain the final actor network corresponding to the allocation policy. The specific algorithm flow includes:
[0189] S001: Initialize experience replay buffer D, critic network parameters Φ1, Φ2, actor network parameters θ;
[0190] S002: Initialize target actor network parameters and target critic network parameters
[0191] S003: for training episode e = 1,..., E do;
[0192] S004: Reset multi-UAV assisted edge computing network environment;
[0193] S005: for training step t = 1,..., T do;
[0194] S006: Get system state s t and randomly generate a Gaussian noise a K ~ N(0, I);
[0195] S007: for denoising step i = K,..., 1 do;
[0196] S008: Use actor network to predict noise θ ;
[0197] S009: Calculate expectation of inverse diffusion process according to formula (30);
[0198] S010: Calculate a i-1 according to formula (31);
[0199] S011: end for;
[0200] S012: Perform action a t = a i and get next time slot state s t+1 and reward r t ;
[0201] S013: Store (s t , a t , r t , s t+1 ) into experience replay buffer D;
[0202] S014: end for;
[0203] S015: Randomly sample a sequence of size B = {(s t , a t , r t , s t+1 )} from experience replay buffer D;
[0204] S016: Update the actor and critic network parameters using B according to formulas (36), (37) and (38);
[0205] S017: Update the target network parameters according to formula (39);
[0206] S018: end for;
[0207] The application can design a system model and service selection according to the needs of edge users by constructing a system model and a data transmission model, realize the allocation of multiple tasks at a time slot, and improve service efficiency. By constructing a service selection model, different priorities can be designed for content generation tasks, and different system benefits brought by different tasks can be distinguished. Since the energy consumption required by GPU to complete the content generation task is much higher than the energy consumption required by signal transmission, by constructing an energy consumption model, only the energy consumption required for task completion is considered, the data processing amount can be reduced, and the processing efficiency can be improved. By designing the overall optimization target, the target of the system is clear, which can ensure that the task allocation and resource management scheme is obtained under the condition of reducing the energy consumption of the system, so as to maximize the task utility of the edge user. By constructing the actor network and the target double critic network based on the overall optimization target and the diffusion model, the powerful fitting ability of the diffusion model to the complex dynamic environment can be used to generate a strategy for dynamic resource allocation, reduce the estimation error of the Q value, and generate an optimal resource allocation strategy that can realize efficient resource allocation with limited network resources.
[0208] Please refer to Figures 8-9 , Figure 8 The comparison chart of the average reward per task of the diffusion model based UAV edge computing network resource allocation method of the embodiment of the application and the current mainstream reinforcement learning algorithms proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies; Figure 9 The comparison chart of the average reward per training task of the diffusion model based UAV edge computing network resource allocation method of the embodiment of the application and the current mainstream reinforcement learning algorithms proximal policy optimization (PPO) and soft actor-critic (SAC) under different user arrival strategies; it can be seen that under the condition of low user arrival rate, the effects of all methods are similar, but when the user arrival rate increases, the complexity of the environment will increase significantly due to the increase in the number of users, under this condition, the method proposed in the application has strong fitting ability to complex environment because of the introduction of the reverse diffusion model, and the effect is significantly better than that of the traditional method.
[0209] The technical scheme of the present application has been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical schemes after the changes or replacements will all fall within the protection scope of the present application.
Claims
1. A method for allocating network resources for UAV edge computing based on a diffusion model, characterized in that: include: Step S1: Obtain the content generation task information randomly generated by any edge user in the multi-UAV-assisted edge computing network in the target time slot and the service selection of the edge user, and build a system model and a data transmission model corresponding to each service selection; The multi-UAV-assisted edge computing network includes several UAVs, several edge users, and several edge servers; the content generation task information includes the task data size, the computing load required for the task, the edge user priority, and the maximum service delay that the edge user can tolerate; Step S2, calculating the corresponding task completion delay according to the system model and each of the data transmission models, and building a service selection model based on the completion delay of each task and the priority of the content generation task; Step S3, building an energy consumption model based on the energy consumed by the edge user, the edge server, and the GPU of the drone to complete the content generation task in the target time slot; Step S4, determining an overall optimization target based on the system model, the service selection model, the data transmission model, and the energy consumption model; Step S5: constructing an actor network and a target dual-critic network based on the overall optimization goal and the diffusion model, and updating parameters of the actor network and the target dual-critic network to generate an optimal resource allocation strategy.
2. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 1 is characterized in that: In the step S1, it includes: Step S11, constructing a system model based on the content generation task information; Step S12: determining a corresponding channel type based on the service selection of the edge user, and obtaining additional interference corresponding to each channel type; Step S13, determining a channel coefficient corresponding to each channel type according to the system model; Step S14: constructing a data transmission model corresponding to each channel type according to each of the channel coefficients and the additional interference corresponding to each of the channel types.
3. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 2 is characterized in that: In the step S2, it includes: Step S21: constructing a service selection sequence based on the service selection of the edge user and designing corresponding service constraints; Step S22, calculating the task completion delay corresponding to each service selection based on the service selection sequence and the system model; Step S23: constructing a service selection model according to the completion delay of each task and the priority of the content generation task.
4. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 3 is characterized in that: In the step S22, it includes: The task calculation delay and task transmission delay corresponding to each service selection are determined based on the system model, the service selection sequence and the corresponding data transmission model, and the task completion delay corresponding to each service selection is calculated according to the task calculation delay and the task transmission delay.
5. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 4 is characterized in that: In the step S23, it includes: A utility function corresponding to each task priority is determined according to the task completion delay and the content generation task priority, and a service selection model is constructed according to each utility function.
6. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 5 is characterized in that: In the step S3, it includes: Based on the GPU power of the edge user, the edge server and the drone in the target time slot and the task calculation delay corresponding to each service selection, the energy consumed by the GPU of the edge user, the edge server and the drone to complete the content generation task in the target time slot is calculated to construct an energy consumption model.
7. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 6 is characterized in that: In the step S4, it includes: Step S41, designing a task utility function of the edge user in the target time slot based on the service selection model; Step S42: constructing an optimization objective function based on the task utility function and the energy consumption model; Step S43: constructing optimization constraints based on the system model, the data transmission model, and the service selection model; Step S44: determining an overall optimization objective based on the optimization objective function and the optimization constraints.
8. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 7 is characterized in that: In the step S5, it includes: Step S51, determining the state space, action space and reward function of the Markov decision process based on the overall optimization objective and the reverse diffusion process of the diffusion model; Step S52, constructing an actor network based on the state space, the action space, and the reward function; Step S53: construct a target dual-critic network, design update gradients for the target dual-critic network and the actor network respectively, and update parameters of the actor network and the target dual-critic network to generate an optimal resource allocation strategy.
9. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 8, characterized in that: In step S1 , the computational load required for the content generation task is calculated according to the number of denoising steps required for the task.
10. The method for allocating network resources of UAV edge computing based on diffusion model according to claim 9 is characterized in that: In step S1, the service selection of the edge user includes offloading the content generation task to a drone, offloading the content generation task to an adjacent drone, and offloading the content generation task to an edge server.