Unmanned aerial vehicle assisted mobile edge calculation optimization method and device, and related product

By using drones to send artificial interference signals and employing a multi-agent optimization scheme, the security and energy efficiency issues of drone-assisted mobile edge computing systems were resolved. This enabled efficient data transmission and system optimization, thereby improving the security and energy efficiency of drone-assisted mobile edge computing systems.

CN121908277APending Publication Date: 2026-04-21GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU
Filing Date
2025-12-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In the traditional cloud computing paradigm, terminal devices have limited computing power and scarce battery resources, making it difficult to meet the data processing requirements of high latency and low reliability. Drone-assisted mobile edge computing systems face challenges in security and energy efficiency optimization in dynamic network environments, especially when wireless communication broadcast characteristics are vulnerable to eavesdropping attacks and drone energy is limited, making it difficult to achieve a fine balance. System optimization involves non-convex mixed integer nonlinear programming problems with multi-dimensional coupled variables, and traditional optimization algorithms have high computational complexity and poor real-time adaptability.

Method used

The design involves sending artificial interference signals to eavesdroppers during the hovering phase of a drone. By combining physical layer security technology with the randomness of wireless channels to ensure transmission security, an optimization model is constructed to maximize the long-term security computational efficiency of the secure communication system. Through a multi-agent, dual-delay, deep deterministic policy gradient optimization scheme, the secure communication system is modeled as a Markov decision process, achieving multi-dimensional optimization.

Benefits of technology

It significantly reduces the possibility of data theft, enhances the ability of communication systems to resist eavesdropping attacks, achieves efficient energy utilization, improves overall system performance and optimization efficiency, ensures the confidentiality and stability of communication, and extends system uptime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908277A_ABST
    Figure CN121908277A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle assisted mobile edge calculation optimization method and device and a related product, and relates to the technical field of edge calculation. According to the method, an unmanned aerial vehicle is designed to send an artificial interference signal to an eavesdropper in a hovering stage to directly interfere the eavesdropper to receive the signal, so that the eavesdropper is difficult to obtain effective data, and the possibility that the data is stolen is remarkably reduced; meanwhile, the transmission security is guaranteed by combining a physical layer security technology and utilizing the randomness of a wireless channel, so that the eavesdropping attack resistance of the security communication system is enhanced from multiple aspects, and the security of communication is ensured; moreover, by combining the optimization variables, the constraint conditions and the constructed optimization model and comprehensively considering the total energy consumption of the security communication system, fine balance among trajectory planning, resource allocation and an anti-eavesdropping strategy can be realized, the overall energy efficiency of the security communication system is effectively improved, the operation time of the security communication system is prolonged, and the energy waste is reduced. And efficient utilization of energy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of edge computing technology, and in particular to a method, apparatus and related products for optimizing mobile edge computing with drone assistance. Background Technology

[0002] In today's era of rapid digital development, 6G communication technology is penetrating deeply into many key areas such as the Industrial Internet of Things, smart city traffic management, drone swarm collaboration, and autonomous driving environmental perception with unprecedented momentum. These areas have extremely high requirements for real-time data processing. With the integration of 6G technology, massive amounts of heterogeneous data are emerging like a tidal wave, and the demand for real-time processing is growing exponentially.

[0003] Traditional cloud computing has played a crucial role in the development of information technology. However, it has revealed numerous insurmountable problems in the face of current new demands. Terminal devices, such as IoT devices widely used in various fields, have inherent limitations. On the one hand, their computing power is relatively limited, making them inadequate when dealing with massive data processing tasks. On the other hand, these devices typically rely on battery power, and battery resources are scarce. This makes it difficult for traditional cloud computing to meet the stringent requirements of low latency and high reliability in high-load application scenarios.

[0004] To address this challenge, Mobile Edge Computing (MEC) emerged and quickly became one of the core supporting technologies for 6G. It innovatively pushes computing resources down to the network edge; this architectural change effectively reduces data transmission latency and significantly alleviates network congestion. However, even with its significant advantages, MEC still faces two major challenges in dynamically changing network environments: security and energy efficiency optimization.

[0005] From a security perspective, the broadcast nature of wireless communication makes drone-assisted MEC systems highly vulnerable to eavesdropping attacks. Traditional security solutions are mostly based on high-level encryption technologies, which rely on algorithmic complexity to ensure data transmission security. However, as attackers' computing power continues to improve, this encryption method, which depends on algorithmic complexity, is susceptible to brute-force key cracking, leading to data leakage and posing serious security risks to the system.

[0006] Regarding energy efficiency optimization, while physical layer security technologies can leverage the randomness of wireless channels to ensure data transmission security, the mobility of drone platforms means that these channels are constantly changing. Simultaneously, the drone's own energy is limited. This necessitates a precise balance in trajectory planning, resource allocation, and anti-eavesdropping strategies. Otherwise, either excessive focus on security leads to excessive energy consumption, or excessive energy conservation compromises data transmission security.

[0007] Furthermore, optimizing the entire system is an extremely complex problem, involving coupled variables from multiple dimensions such as UAV trajectory, user task offloading, and communication and computing resource allocation. These variables are intertwined, forming a non-convex mixed-integer nonlinear programming problem. Traditional optimization algorithms, such as alternating optimization and convex approximation, exhibit high computational complexity and poor real-time adaptability when dealing with such complex problems. With the rapid development of 6G networks, the network environment exhibits high dynamism and non-stationarity, making these traditional algorithms clearly insufficient to meet the demands of practical applications. Summary of the Invention

[0008] In view of the above problems, this application is made to provide a method, apparatus, and related products for optimizing UAV-assisted mobile edge computing to overcome or at least partially solve the above problems. The technical solution is as follows: Firstly, a drone-assisted mobile edge computing optimization method is provided, the method comprising: A secure communication system for drone-assisted mobile edge computing was established, comprising one full-duplex drone, K One single-antenna ground user equipment and one passive eavesdropper with a random location; With a fixed drone flight altitude, the time slot length The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. The design optimization objective is to maximize the long-term secure computing efficiency of the secure communication system, i.e., the maximum amount of data processed per unit of energy consumption, also known as the secure data volume processed per unit of energy consumption. Combining the optimization variables and constraints, the optimization model is constructed as follows:

[0009] The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; To address the non-convexity and dynamic nature of the constructed optimization model, an optimization scheme based on multi-agent dual-delay deep deterministic policy gradient is proposed. This scheme models the secure communication system as a Markov decision process and achieves multi-dimensional optimization through multi-agent and shared evaluation networks.

[0010] In one possible implementation, the method further includes: Position and channel modeling: UAVs in the 19th century n The horizontal coordinate of the time slot is User equipment k In the n The coordinates of the time slot are The location of the passive eavesdropper is within the estimated range. The ground-to-air channel uses a probabilistic line-of-sight model, with the line-of-sight probability as follows:

[0011] in, For user equipment k In the n Time slots and the drone's elevation angle, , For coefficients; The path loss, which combines free space loss, line-of-sight additional loss, and non-line-of-sight additional loss, is expressed by the following formula:

[0012] in, For user equipment k In the n Free space loss of time slots, For user equipment k In the n Additional line-of-sight loss in time slots, For user equipment k In the nNon-line-of-sight additional loss in time slots.

[0013] In one possible implementation, the method further includes: Establish the capability for drones to handle unloading tasks Limited by CPU frequency limit Constraints, i.e. , For user equipment k In the n The amount of data unloaded from the time slot.

[0014] In one possible implementation, the method further includes: Establish user equipment k In the n The formula for the local computational energy consumption of a time slot is:

[0015] in, The hardware capacitance coefficient of the user equipment; Establish drones in n The formula for calculating the energy consumption of a time slot is:

[0016] in, The hardware capacitance coefficient of the drone. For drones in the n CPU frequency of the time slot.

[0017] In one possible implementation, the method further includes: Total energy consumption for establishing a secure communication system The formula is as follows:

[0018] in, For user equipment k In the n Energy consumption of time slots, including user equipment k In the n Local computing power consumption of time slots and user equipment k In the n Energy consumption of offloaded data in time slots ; Calculate the weighting coefficients for the energy consumption of drones. The weighting factor for the propulsion energy consumption of drones. For drones in the n Energy consumption for time slot calculation For drones in the n Energy consumption for time slot propulsion For drones in the n Interference energy consumption in time slots.

[0019] In one possible implementation, the four-tuple of a Markov decision process is defined as follows: State space S: contains the UAV in the state space S. n Current position of the time slot With all user equipment in the n Data queue length of time slot ,Right now It comprehensively reflects the resource and task load status of the secure communication system; Action space A is divided into two types of agent actions: drone trajectory actions. , yes Vectorization controls flight speed and direction; Resource allocation actions Among them, discrete unloading decision Mapped to the continuous interval [0,1], supporting gradient optimization; State transition: UAV position is updated deterministically according to the movement model. , yes Vectorization, yes Vectorization; user equipment k In the n Data queue length of time slots according to Random updates The number of tasks arriving follows a bounded third-order distribution; Design a reward function that balances global energy efficiency and local queue stability for dual-component rewards: ; in, To normalize the computational efficiency for safety; As a reward for queue stability, , and These are the weighting coefficients.

[0020] Secondly, a drone-assisted mobile edge computing optimization device is provided, the device comprising: The deployment unit is used to build a secure communication system for UAV-assisted mobile edge computing. The secure communication system includes one full-duplex UAV. K One single-antenna ground user equipment and one passive eavesdropper with a random location; The operation unit is used to fix the drone's flight altitude within the time slot. The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. The construction unit is used to design an optimization model with the objective of maximizing the long-term secure computing efficiency of the secure communication system, i.e., the maximum amount of data processed per unit of energy consumption, also known as the amount of secure data processed per unit of energy consumption. Combining the optimization variables and constraints, the constructed optimization model is as follows:

[0021] The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; The optimization unit is used to address the non-convexity and dynamics of the constructed optimization model. Based on the multi-agent dual-delay deep deterministic policy gradient optimization scheme, it models the secure communication system as a Markov decision process and achieves multi-dimensional optimization through multi-agent and shared evaluation networks.

[0022] Thirdly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the UAV-assisted mobile edge computing optimization method described in any of the preceding claims.

[0023] Fourthly, a storage medium is provided that stores a computer program, wherein the computer program is configured to execute the UAV-assisted mobile edge computing optimization method described in any of the preceding claims at runtime.

[0024] Fifthly, a computer program product is provided, including a computer program configured to execute the UAV-assisted mobile edge computing optimization method described in any of the preceding claims at runtime.

[0025] By utilizing the above technical solutions, the UAV-assisted mobile edge computing optimization method, apparatus, and related products provided in this application significantly reduce the possibility of data theft by designing the UAV to send artificial interference signals to eavesdroppers during the hovering phase, directly interfering with the eavesdroppers' signal reception and making it difficult for them to obtain valid data. Simultaneously, by combining physical layer security technology with the randomness of wireless channels to ensure transmission security, the method enhances the ability of the secure communication system to resist eavesdropping attacks from multiple levels, ensuring the confidentiality of communication. Furthermore, the optimization model constructed by this application, combining optimization variables and constraints, comprehensively considers the total energy consumption of the secure communication system, achieving a fine balance between trajectory planning, resource allocation, and anti-eavesdropping strategies, effectively improving the overall energy efficiency of the secure communication system. This approach improves efficiency, extends the operating time of secure communication systems, reduces energy waste, and achieves efficient energy utilization. Furthermore, addressing the non-convex mixed-integer nonlinear programming problem involving multi-dimensional coupled variables in system optimization, this application proposes an optimization scheme based on a multi-agent, dual-delay, deep deterministic policy gradient. The secure communication system is modeled as a Markov decision process, and multi-dimensional optimization is achieved through multiple agents and a shared evaluation network. Each agent is responsible for UAV trajectory planning and resource allocation, enabling parallel processing of optimization tasks across different dimensions. Compared to traditional optimization algorithms (such as alternating optimization and convex approximation) which suffer from high computational complexity and poor real-time adaptability, this approach significantly improves optimization efficiency, enabling rapid processing of complex optimization problems, finding better solutions, and enhancing overall system performance. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0027] Figure 1 A flowchart of the UAV-assisted mobile edge computing optimization method provided in an embodiment of this application is shown; Figure 2 A schematic diagram of a secure communication system provided in an embodiment of this application is shown; Figure 3 A structural diagram of the UAV-assisted mobile edge computing optimization device provided in an embodiment of this application is shown. Detailed Implementation

[0028] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."

[0030] To address the aforementioned technical problems, embodiments of this application provide a drone-assisted mobile edge computing optimization method, such as... Figure 1 As shown, the UAV-assisted mobile edge computing optimization method may include the following steps S101 to S104: Step S101: Establish a secure communication system for UAV-assisted mobile edge computing, wherein the secure communication system includes one full-duplex UAV, K A single-antenna ground user equipment and one randomly located passive eavesdropper, such as Figure 2 As shown, in Figure 2 In this context, smart terminals carried by people, vehicles, and other similar devices are considered user equipment. Step S102: Fix the drone's flight altitude within the time slot. The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. Step S103: The design optimization objective is to maximize the long-term secure computing efficiency of the secure communication system, that is, the maximum amount of data processed per unit of energy consumption, also known as the amount of secure data processed per unit of energy consumption. Combining the optimization variables and constraints, the optimization model is constructed as follows:

[0031] The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n The CPU (Central Processing Unit) frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; Step S104: In view of the non-convexity and dynamics of the constructed optimization model, the secure communication system is modeled as a Markov decision process based on the optimization scheme of multi-agent dual-delay deep deterministic policy gradient, and multi-dimensional optimization is achieved through multi-agent and shared evaluation network.

[0032] This application's embodiments involve designing a drone to send artificial interference signals to eavesdroppers during the hovering phase, directly interfering with the eavesdropper's signal reception and making it difficult for them to obtain valid data, thus significantly reducing the possibility of data theft. Simultaneously, by combining physical layer security technology with the randomness of wireless channels to ensure transmission security, the application enhances the secure communication system's ability to resist eavesdropping attacks from multiple levels, ensuring communication confidentiality. Furthermore, the optimization model constructed by this application, combining optimization variables and constraints, comprehensively considers the total energy consumption of the secure communication system, achieving a fine balance between trajectory planning, resource allocation, and anti-eavesdropping strategies, effectively improving the overall energy efficiency of the secure communication system, extending its operating time, and reducing energy waste. This application addresses the issue of non-convex mixed-integer nonlinear programming problems involving multi-dimensional coupled variables in system optimization. It proposes an optimization scheme based on a multi-agent, dual-delay, deep deterministic policy gradient, modeling the secure communication system as a Markov decision process. Multi-agent optimization is achieved through a shared evaluation network, with each agent responsible for UAV trajectory planning and resource allocation. This allows for parallel processing of optimization tasks across different dimensions. Compared to traditional optimization algorithms (such as alternating optimization and convex approximation) which suffer from high computational complexity and poor real-time adaptability, this scheme significantly improves optimization efficiency, enabling rapid processing of complex optimization problems, finding better solutions, and enhancing overall system performance.

[0033] This application provides one possible implementation method, which may further include the following steps: Position and channel modeling: UAVs in the 19th century n The horizontal coordinate of the time slot is User equipmentk In the n The coordinates of the time slot are The location of the passive eavesdropper is within the estimated range. The ground-to-air channel uses a probabilistic line-of-sight model, with the line-of-sight probability as follows:

[0034] in, For user equipment k In the n Time slots and the drone's elevation angle, , For coefficients; The path loss, which combines free space loss, line-of-sight additional loss, and non-line-of-sight additional loss, is expressed by the following formula:

[0035] in, For user equipment k In the n Free space loss of time slots, For user equipment k In the n Additional line-of-sight loss in time slots, For user equipment k In the n Non-line-of-sight additional loss in time slots.

[0036] Considering the dynamic changes in the channel caused by the mobility of the UAV platform, the embodiments of this application accurately describe the characteristics of the channel changes with the location of the UAV and the user equipment when constructing the location and channel model. By acquiring the location information of the UAV and the user equipment in real time, and based on the established channel model, the secure communication system can track the dynamic changes of the channel in real time and adjust the secure transmission strategy in a timely manner, such as adjusting the signal power and encoding method, so as to ensure the security and stability of the communication link in complex dynamic environments and greatly improve the reliability of data transmission.

[0037] This application provides one possible implementation method, which may further include the following steps: Establish the capability for drones to handle unloading tasks Limited by CPU frequency limit Constraints, i.e. , For user equipment k In the n The amount of data unloaded from the time slot.

[0038] This application provides one possible implementation method, which may further include the following steps: Establish user equipment k In then The formula for the local computational energy consumption of a time slot is:

[0039] in, The hardware capacitance coefficient of the user equipment; Establish drones in n The formula for calculating the energy consumption of a time slot is:

[0040] in, The hardware capacitance coefficient of the drone. For drones in the n CPU frequency of the time slot.

[0041] This application provides one possible implementation method, which may further include the following steps: Total energy consumption for establishing a secure communication system The formula is as follows:

[0042] in, For user equipment k In the n Energy consumption of time slots, including user equipment k In the n Local computing power consumption of time slots and user equipment k In the n Energy consumption of offloaded data in time slots ; Calculate the weighting coefficients for the energy consumption of drones. The weighting factor for the propulsion energy consumption of drones. For drones in the n Energy consumption for time slot calculation For drones in the n Energy consumption for time slot propulsion For drones in the n Interference energy consumption in time slots.

[0043] This embodiment comprehensively considers the total energy consumption of a secure communication system, encompassing the energy consumption of local computation and offloading transmission for user equipment, as well as the energy consumption of UAV computation, propulsion, and interference. By establishing a detailed energy consumption model, combined with the task offloading and computation model, a fine balance can be achieved between trajectory planning, resource allocation, and anti-eavesdropping strategies. For example, in a UAV swarm collaboration scenario, the established model can accurately calculate the energy consumption of each link based on the energy status and task requirements of each UAV, rationally plan trajectories and allocate resources, prevent some UAVs from prematurely exiting the task due to energy depletion, effectively improve the overall energy efficiency of the system, extend system operating time, reduce energy waste, and achieve efficient energy utilization.

[0044] This application provides a possible implementation method, and the four-tuple definition of a Markov decision process is as follows: State space S: contains the UAV in the state space S. n Current position of the time slot With all user equipment in the n Data queue length of time slot ,Right now It comprehensively reflects the resource and task load status of the secure communication system; Action space A is divided into two types of agent actions: drone trajectory actions. , yes Vectorization controls flight speed and direction; Resource allocation actions Among them, discrete unloading decision Mapped to the continuous interval [0,1], supporting gradient optimization; State transition: UAV position is updated deterministically according to the movement model. , yes Vectorization, yes Vectorization; user equipment k In the n Data queue length of time slots according to Random updates The number of tasks arriving follows a bounded third-order distribution; Design a reward function that balances global energy efficiency and local queue stability for dual-component rewards: ; in, To normalize the computational efficiency for safety; As a reward for queue stability, , and These are the weighting coefficients.

[0045] This embodiment models the secure communication system as a Markov decision process, and achieves multi-dimensional optimization through multiple agents and a shared evaluation network. Each agent is responsible for UAV trajectory planning and resource allocation, and can process optimization tasks of different dimensions in parallel.

[0046] The above introduces Figure 1 The embodiments shown have various implementation methods for each stage. The following will further explain the UAV-assisted mobile edge computing optimization method of this application through specific embodiments.

[0047] Currently, the following technical problems are faced when processing massive amounts of data: 1) Limitations of the Traditional Cloud Computing Paradigm: Under the traditional cloud computing paradigm, terminal devices such as IoT devices have limited computing power. Faced with the exponentially growing volume of heterogeneous data, they cannot quickly process and analyze the data, resulting in low processing efficiency. At the same time, these devices have scarce battery resources; continuous high-load tasks will quickly deplete their power, making it impossible to guarantee stable operation over long periods. Especially in the Industrial Internet of Things (IIoT), a large number of sensors need to collect and process data in real time. The traditional cloud computing model cannot meet the requirements for low latency and high reliability, affecting the precise control and efficient operation of industrial production. In smart city traffic management, traffic data changes rapidly, and traditional models cannot process and analyze it in a timely manner, easily leading to problems such as exacerbating traffic congestion.

[0048] 2) Security vulnerabilities of Mobile Edge Computing (MEC): In dynamic network environments, while MEC offers the advantage of decentralized computing resources, its security faces challenges. The broadcast nature of wireless communication makes drone-assisted MEC systems vulnerable to eavesdropping attacks. Traditional security solutions based on high-level encryption rely on algorithm complexity; as attackers' computing power increases, keys become susceptible to brute-force attacks, increasing the likelihood of data leakage. For example, if autonomous driving environmental perception data is stolen and tampered with, it will seriously threaten driving safety.

[0049] 3) MEC Energy Efficiency Optimization Challenges: While physical layer security technologies can leverage the randomness of wireless channels to ensure transmission security, the mobility of UAV platforms causes dynamic changes in the channel, making it difficult to maintain a stable and secure transmission state. Furthermore, UAVs have limited energy, making it difficult to achieve a precise balance between trajectory planning, resource allocation, and anti-eavesdropping strategies. In UAV swarm collaboration scenarios, unreasonable planning and allocation can cause some UAVs to run out of energy and prematurely exit the mission, affecting the overall collaboration effect.

[0050] 4) Challenges in system optimization algorithms: System optimization involves multi-dimensional coupled variables, forming a non-convex mixed integer nonlinear programming problem. Traditional optimization algorithms, such as alternating optimization and convex approximation, have high computational complexity and poor real-time adaptability in the dynamic and non-stationary environment of 6G networks. They cannot make optimal decisions quickly according to network changes, resulting in resource waste or task processing delays, making it difficult to meet the high-efficiency operation requirements of 6G communication technology application scenarios.

[0051] This specific embodiment provides a framework for model building, problem construction, and optimization based on Multi-Agent Deep Reinforcement Learning (MADRL), which will be described in detail below.

[0052] (I) Model Construction A secure communication system for drone-assisted mobile edge computing was established, comprising one full-duplex drone, K A single-antenna ground user equipment and one randomly located passive eavesdropper, such as Figure 2 As shown.

[0053] With a fixed drone flight altitude, the time slot length The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues local computation.

[0054] Position and channel modeling: UAVs in the 19th century n The horizontal coordinate of the time slot is User equipment k In the n The coordinates of the time slot are The location of the passive eavesdropper is within the estimated range. The ground-to-air channel uses a probabilistic line-of-sight model, with the line-of-sight probability as follows:

[0055] in, For user equipment k In the n Time slots and the drone's elevation angle, , For coefficients; The path loss, which combines free space loss, line-of-sight additional loss, and non-line-of-sight additional loss, is expressed by the following formula:

[0056] in, For user equipment k In the n Free space loss of time slots, For user equipment k In the n Additional line-of-sight loss in time slots, For user equipment k In the n Non-line-of-sight additional loss in time slots.

[0057] Task offloading and computation model: A binary offloading strategy is adopted. 0 indicates local computation, and 1 indicates offloading to the drone.

[0058] User equipment k In the n The energy consumption of the offloaded data in the time slot is , ; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n The time ratio of time slot transmission signal; This represents the time slot duration.

[0059] Similarly, drones in the n The interference energy consumption of the time slot is , ; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in a time slot.

[0060] User equipment k In the n The local computing power of the time slot is , ; For user equipment k In the n CPU frequency of the time slot; The period is calculated per unit bit.

[0061] Establish the capability for drones to handle unloading tasks Limited by CPU frequency limit Constraints, i.e. , For user equipment k In the n The amount of data offloaded from a time slot; For communication overhead.

[0062] Establish user equipment k In the n The formula for the local computational energy consumption of a time slot is:

[0063] in, This refers to the hardware capacitance coefficient of the user equipment.

[0064] Establish drones in n The formula for calculating the energy consumption of a time slot is:

[0065] in, The hardware capacitance coefficient of the drone. For drones in the n CPU frequency of the time slot.

[0066] The formula for the propulsion power of a drone is as follows: The basis of drone propulsion energy consumption is propulsion power, which needs to consider three parts: blade profile power (to overcome propeller airfoil drag), induced power (to overcome gravity and generate lift), and aerodynamic drag power (to overcome air resistance). The total propulsion power formula is as follows:

[0067] in, , , , These are total propulsion power, blade rotation power, induced power, and aerodynamic drag power, respectively. For drones in the n The horizontal flight speed of the time slot.

[0068] Safe rate and energy consumption model: For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers.

[0069] The total energy consumption of a secure communication system includes user equipment energy consumption (local computing and offloading transmission) and UAV energy consumption (computing, propulsion, and jamming). This establishes the total energy consumption of the secure communication system. The formula is as follows:

[0070] in, For user equipment k In the n Energy consumption of time slots, including user equipment k In the n Local computing power consumption of time slots and user equipment k In the n Energy consumption of offloaded data in time slots ; Calculate the weighting coefficients for the energy consumption of drones. The weighting factor for the propulsion energy consumption of drones. For drones in the n Energy consumption for time slot calculation For drones in the n Energy consumption for time slot propulsion For drones in the n Interference energy consumption in time slots.

[0071] (II) Problem Construction The optimization objective of this specific embodiment is to maximize the long-term secure computing efficiency of the secure communication system, that is, the maximum amount of data processed per unit of energy consumption, also known as the amount of secure data processed per unit of energy consumption. Combining the optimization variables and constraints, the optimization model is constructed as follows:

[0072] The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the nThe horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , .

[0073] (III) Optimization Framework Based on Multi-Agent Deep Reinforcement Learning (MADRL) To address the non-convexity and dynamic nature of the constructed optimization model, an optimization scheme based on multi-agent dual-delay deep deterministic policy gradient is proposed. This scheme models the secure communication system as a Markov Decision Process (MDP) and achieves multi-dimensional optimization through multi-agent and shared evaluation networks.

[0074] (1) The definition of a four-tuple in a Markov Decision Process (MDP) is as follows: State space S: contains the UAV in the state space S. n Current position of the time slot With all user equipment in the n Data queue length of time slot ,Right now It comprehensively reflects the resource and task load status of the secure communication system; Action space A is divided into two types of agent actions: drone trajectory actions. , yes Vectorization controls flight speed and direction; Resource allocation actions Among them, discrete unloading decision Mapped to the continuous interval [0,1], supporting gradient optimization; State transition: UAV position is updated deterministically according to the movement model. , yes Vectorization, yes Vectorization; user equipment k In the n Data queue length of time slots according to Random updates The number of tasks arriving follows a bounded third-order distribution; Design a reward function that balances global energy efficiency and local queue stability for dual-component rewards: ; in, To normalize the computational efficiency for safety; As a reward for queue stability, , and These are the weighting coefficients.

[0075] (2) MADRL core architecture The multi-agent system is a distributed agent system: Two independent agents are set up, responsible for UAV trajectory planning (Actor1) and resource allocation (Actor2) respectively. Each agent outputs an action based on its current state, and the joint action-driven system state transition formula is as follows:

[0076] in, , For policy functions; Output an action for the current state; , For agent policy network parameters, achieve multi-dimensional optimization and decoupling; This is for the state transition of the action joint drive system.

[0077] Shared Critic Network: Employs a twin critic (value evaluation network), learning a value function based on global state and joint actions.

[0078] in, E[] represents the network evaluation parameters; E[] represents the expectation operator. The discount factor is a number between 0 and 1 used to measure the current value of future rewards. For the future t Each time step provides instant feedback and rewards.

[0079] This network addresses the nonstationarity problem of multi-agent systems by providing a unified optimization objective for all agents, thus avoiding local optima.

[0080] Perturbed Actors: Add clipped Gaussian noise when generating the target action, as shown in the formula: ;in, For the target policy network; for The next state output action; These are the parameters of the target policy network; The standard deviation of noise is used to reduce Q Value overestimation bias, enhance exploration capabilities and environmental robustness.

[0081] Network training and update process: Initialization: Construct the Actor1 (trajectory) and Actor2 (resource) policy networks, the twin critic network, and the corresponding target network; initialize the experience replay buffer. .

[0082] Experience collection: The agent adds Ornstein-Uhlenbeck noise (a stochastic process with time correlation and mean regression properties, mainly used to explore continuous action spaces) to generate exploratory actions, interact with the environment, and store the experience. Balance exploration and utilization.

[0083] Parameter update: when When there is sufficient data, sample mini-batch data (a method of dividing a large dataset into smaller subsets for training, aiming to balance computational efficiency and model convergence stability) to calculate the target action and target. Q Values; update the evaluation network by minimizing the critic loss; update the target network using a soft update strategy.

[0084] The specific embodiment can achieve the following technical effects: Traditional security schemes based on high-level encryption are vulnerable to attacks due to increased computing power, posing a risk of key brute-force cracking. This specific embodiment addresses this by designing the drone to send artificial interference signals to eavesdroppers during its hovering phase. This technique directly interferes with the eavesdropper's signal reception, making it difficult for them to obtain valid data and significantly reducing the likelihood of data theft. This provides more reliable security for critical information such as environmental perception data for autonomous driving. Simultaneously, physical layer security technology leverages the randomness of wireless channels to ensure transmission security, enhancing the system's ability to resist eavesdropping attacks from multiple levels and ensuring communication confidentiality. Considering the dynamic channel changes caused by the drone platform's mobility, this embodiment accurately describes the channel characteristics as the drone and user positions change when constructing the position and channel model. By acquiring the drone and user's position information in real time and based on the established channel model, the system can track dynamic channel changes in real time and adjust security transmission strategies promptly, such as adjusting signal power and encoding methods. This ensures the security and stability of the communication link in complex dynamic environments, significantly improving the reliability of data transmission.

[0085] This specific embodiment comprehensively considers the total system energy consumption, covering the energy consumption of user equipment local computing and offloading transmission, as well as the energy consumption of UAV computing, propulsion, and interference. By establishing a detailed energy consumption model, combined with the task offloading and computing model, a fine balance can be achieved between trajectory planning, resource allocation, and anti-eavesdropping strategies. For example, in a UAV swarm collaboration scenario, the established model can accurately calculate the energy consumption of each link based on the energy status and task requirements of each UAV, rationally plan trajectories and allocate resources, prevent some UAVs from prematurely exiting the task due to energy depletion, effectively improve the overall system energy efficiency, extend system operating time, reduce energy waste, and achieve efficient energy utilization.

[0086] This specific embodiment addresses the non-convex mixed-integer nonlinear programming problem involving multi-dimensional coupled variables in system optimization, proposing an optimization framework based on Multi-Agent Deep Reinforcement Learning (MADRL). This framework models the system as a Markov Decision Process (MDP), utilizing distributed agents and a shared evaluation network to decouple multi-dimensional optimization. The distributed agents are responsible for UAV trajectory planning and resource allocation, respectively, enabling parallel processing of optimization tasks across different dimensions. Compared to traditional optimization algorithms (such as alternating optimization and convex approximation) which suffer from high computational complexity and poor real-time adaptability, this framework significantly improves optimization efficiency, enabling rapid handling of complex optimization problems, finding better solutions, and enhancing overall system performance. In the dynamic and non-stationary environment of 6G networks, the optimization framework of this specific embodiment exhibits good real-time adaptability. The agents continuously interact with the environment to collect experience, learning and adjusting policies based on state transitions and reward functions. For example, when faced with situations such as random changes in the number of tasks arriving or dynamic changes in channel status, the intelligent agent can adjust its actions in a timely manner based on environmental feedback, such as adjusting the drone trajectory and reallocating resources, so that the system always maintains a near-optimal operating state, effectively copes with the dynamism and non-stationarity of 6G networks, and improves the robustness and reliability of the system.

[0087] It should be noted that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In practical applications, all the above possible implementation methods can be arbitrarily combined in a combined manner to form possible embodiments of this application, which will not be described in detail here.

[0088] Based on the UAV-assisted mobile edge computing optimization methods provided in the above embodiments, and based on the same inventive concept, this application also provides a UAV-assisted mobile edge computing optimization device.

[0089] Figure 3 This is a structural diagram of the UAV-assisted mobile edge computing optimization device provided in an embodiment of this application. Figure 3 As shown, the UAV-assisted mobile edge computing optimization device may specifically include a construction unit 310, an operation unit 320, a construction unit 330, and an optimization unit 340.

[0090] Unit 310 is used to build a secure communication system for UAV-assisted mobile edge computing. The secure communication system includes one full-duplex UAV. K One single-antenna ground user equipment and one passive eavesdropper with a random location; Operation unit 320 is used to fix the flight altitude of the UAV within the time slot. The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. Element 330 is constructed to design an optimization model that maximizes the long-term secure computing efficiency of the secure communication system, i.e., the maximum amount of data processed per unit of energy consumption, also known as the amount of secure data processed per unit of energy consumption. Combining the optimization variables and constraints, the constructed optimization model is as follows:

[0091] The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; The optimization unit 340 is used to optimize the non-convexity and dynamics of the constructed optimization model. Based on the optimization scheme of multi-agent dual-delay deep deterministic policy gradient, the secure communication system is modeled as a Markov decision process, and multi-dimensional optimization is achieved through multi-agent and shared evaluation network.

[0092] This application embodiment provides a possible implementation, wherein the construction unit 330 is further configured to: Position and channel modeling: UAVs in the 19th century n The horizontal coordinate of the time slot is User equipment k In the n The coordinates of the time slot are The location of the passive eavesdropper is within the estimated range. The ground-to-air channel uses a probabilistic line-of-sight model, with the line-of-sight probability as follows:

[0093] in, For user equipment k In the n Time slots and the drone's elevation angle, , For coefficients; The path loss, which combines free space loss, line-of-sight additional loss, and non-line-of-sight additional loss, is expressed by the following formula:

[0094] in, For user equipment k In the n Free space loss of time slots, For user equipment k In the n Additional line-of-sight loss in time slots, For user equipment k In the n Non-line-of-sight additional loss in time slots.

[0095] This application embodiment provides a possible implementation, wherein the construction unit 330 is further configured to: Establish the capability for drones to handle unloading tasks Limited by CPU frequency limit Constraints, i.e. , For user equipment k In the n The amount of data unloaded from the time slot.

[0096] This application embodiment provides a possible implementation, wherein the construction unit 330 is further configured to: Establish user equipment k In the n The formula for the local computational energy consumption of a time slot is:

[0097] in, The hardware capacitance coefficient of the user equipment; Establish drones in n The formula for calculating the energy consumption of a time slot is:

[0098] in, The hardware capacitance coefficient of the drone. For drones in the n CPU frequency of the time slot.

[0099] This application embodiment provides a possible implementation, wherein the construction unit 330 is further configured to: Total energy consumption for establishing a secure communication system The formula is as follows:

[0100] in, For user equipment k In the n Energy consumption of time slots, including user equipment k In the n Local computing power consumption of time slots and user equipment k In the n Energy consumption of offloaded data in time slots ; Calculate the weighting coefficients for the energy consumption of drones. The weighting factor for the propulsion energy consumption of drones. For drones in the n Energy consumption for time slot calculation For drones in the n Energy consumption for time slot propulsion For drones in the n Interference energy consumption in time slots.

[0101] This application provides a possible implementation method, and the four-tuple definition of a Markov decision process is as follows: State space S: contains the UAV in the state space S. n Current position of the time slot With all user equipment in the n Data queue length of time slot ,Right now It comprehensively reflects the resource and task load status of the secure communication system; Action space A is divided into two types of agent actions: drone trajectory actions. , yes Vectorization controls flight speed and direction; Resource allocation actions Among them, discrete unloading decision Mapped to the continuous interval [0,1], supporting gradient optimization; State transition: UAV position is updated deterministically according to the movement model. , yes Vectorization, yes Vectorization; user equipment k In the n Data queue length of time slots according to Random updates The number of tasks arriving follows a bounded third-order distribution; Design a reward function that balances global energy efficiency and local queue stability for dual-component rewards: ; in, To normalize the computational efficiency for safety; As a reward for queue stability, , and These are the weighting coefficients.

[0102] Based on the same inventive concept, this application also provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the UAV-assisted mobile edge computing optimization method of any of the above embodiments.

[0103] Based on the same inventive concept, this application also provides a storage medium storing a computer program, wherein the computer program is configured to execute the UAV-assisted mobile edge computing optimization method of any of the above embodiments at runtime.

[0104] Based on the same inventive concept, this application also provides a computer program product, including a computer program configured to execute the UAV-assisted mobile edge computing optimization method of any of the above embodiments at runtime.

[0105] Those skilled in the art will understand that the technical solution of this application, or all or part of it, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several program instructions to cause an electronic device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of this application when running the program instructions. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0106] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as electronic devices like personal computers, servers, or network devices) associated with program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the embodiments of this application.

[0107] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of this application, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to leave the protection scope of this application.

Claims

1. A method for optimizing mobile edge computing assisted by unmanned aerial vehicles (UAVs), characterized in that, The method includes: A secure communication system for drone-assisted mobile edge computing was established, comprising one full-duplex drone, K One single-antenna ground user equipment and one passive eavesdropper with a random location; With a fixed drone flight altitude, the time slot length The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. The design optimization objective is to maximize the long-term secure computing efficiency of the secure communication system, i.e., the maximum amount of data processed per unit of energy consumption, also known as the secure data volume processed per unit of energy consumption. Combining the optimization variables and constraints, the optimization model is constructed as follows: The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; To address the non-convexity and dynamic nature of the constructed optimization model, an optimization scheme based on multi-agent dual-delay deep deterministic policy gradient is proposed. This scheme models the secure communication system as a Markov decision process and achieves multi-dimensional optimization through multi-agent and shared evaluation networks.

2. The method according to claim 1, characterized in that, The method further includes: Position and channel modeling: UAVs in the 19th century n The horizontal coordinate of the time slot is User equipment k In the n The coordinates of the time slot are The location of the passive eavesdropper is within the estimated range. The ground-to-air channel uses a probabilistic line-of-sight model, with the line-of-sight probability as follows: in, For user equipment k In the n Time slots and the drone's elevation angle, , For coefficients; The path loss, which combines free space loss, line-of-sight additional loss, and non-line-of-sight additional loss, is expressed by the following formula: in, For user equipment k In the n Free space loss of time slots, For user equipment k In the n Additional line-of-sight loss in time slots, For user equipment k In the n Non-line-of-sight additional loss in time slots.

3. The method according to claim 2, characterized in that, The method further includes: Establish the capability for drones to handle unloading tasks Limited by CPU frequency limit Constraints, i.e. , For user equipment k In the n The amount of data unloaded from the time slot.

4. The method according to claim 3, characterized in that, The method further includes: Establish user equipment k In the n The formula for the local computational energy consumption of a time slot is: in, The hardware capacitance coefficient of the user equipment; Establish drones in n The formula for calculating the energy consumption of a time slot is: in, The hardware capacitance coefficient of the drone. For drones in the n CPU frequency of the time slot.

5. The method according to claim 4, characterized in that, The method further includes: Total energy consumption for establishing a secure communication system The formula is as follows: in, For user equipment k In the n Energy consumption of time slots, including user equipment k In the n Local computing power consumption of time slots and user equipment k In the n Energy consumption of offloaded data in time slots ; Calculate the weighting coefficients for the energy consumption of drones. The weighting factor for the propulsion energy consumption of drones. For drones in the n Energy consumption for time slot calculation For drones in the n Energy consumption for time slot propulsion For drones in the n Interference energy consumption in time slots.

6. The method according to claim 5, characterized in that, The definition of a four-tuple in a Markov decision process is as follows: State space S: contains the UAV in the state space S. n Current position of the time slot With all user equipment in the n Data queue length of time slot ,Right now It comprehensively reflects the resource and task load status of the secure communication system; Action space A is divided into two types of agent actions: drone trajectory actions. , yes Vectorization controls flight speed and direction; Resource allocation actions Among them, discrete unloading decision Mapped to the continuous interval [0,1], supporting gradient optimization; State transition: UAV position is updated deterministically according to the movement model. , yes Vectorization, yes Vectorization; user equipment k In the n Data queue length of time slots according to Random updates The number of tasks arriving follows a bounded third-order distribution; Design a reward function that balances global energy efficiency and local queue stability for dual-component rewards: ; in, To normalize the computational efficiency for safety; As a reward for queue stability, , and These are the weighting coefficients.

7. A drone-assisted mobile edge computing optimization device, characterized in that, The device includes: The deployment unit is used to build a secure communication system for UAV-assisted mobile edge computing. The secure communication system includes one full-duplex UAV. K One single-antenna ground user equipment and one passive eavesdropper with a random location; The operation unit is used to fix the drone's flight altitude within the time slot. The system completes two phases of operation: hovering and moving. During the hovering phase, the user equipment transmits unloading mission data to the drone, while the drone simultaneously sends artificial interference signals to the passive eavesdropper. During the moving phase, the drone updates its position and processes the received mission data, while the user equipment continues to perform local calculations. The construction unit is used to design an optimization model with the objective of maximizing the long-term secure computing efficiency of the secure communication system, i.e., the maximum amount of data processed per unit of energy consumption, also known as the amount of secure data processed per unit of energy consumption. Combining the optimization variables and constraints, the constructed optimization model is as follows: The constraints are as follows: ; ; ; ; ; ; ; ; The optimization variables include: For user equipment k In the n The time ratio of time slot transmission signal; For user equipment k In the n Task offloading decision for time slots: 0 indicates local computation, 1 indicates offloading to the drone; For drones in the n The horizontal coordinate of the time slot; For user equipment k In the n Transmit power of the time slot; For user equipment k In the n CPU frequency of the time slot; For drones in the n The interference power of sending artificial interference signals to passive eavesdroppers in time slots; For user equipment k In the n The amount of data processed in a time slot; For the first n The time slot includes the total energy consumption of the secure communication system, encompassing the energy consumption of user equipment and drones. N The number of time slots; For drones in the n The horizontal coordinate of the +1 time slot; For drones in the n Horizontal flight speed in time slots; This refers to the maximum horizontal flight speed of the drone; , For user equipment k In the n The local computing power of time slots , The calculation period is calculated per unit bit; For user equipment k In the n Secure transmission rate of time slots For user equipment k In the n Time slots and channel capacity for UAVs For user equipment k In the n Time slots and channel capacity for passive eavesdroppers; For communication overhead; For user equipment k In the n Data queue length of the time slot This represents the maximum queue capacity. User equipment energy consumption includes the energy consumption of user equipment offloading data. k In the n The energy consumption of the offloaded data in the time slot is , ; Drone energy consumption includes the interference energy consumption of the drone, and the drone in the first... n The interference energy consumption of the time slot is , ; The optimization unit is used to address the non-convexity and dynamics of the constructed optimization model. Based on the multi-agent dual-delay deep deterministic policy gradient optimization scheme, it models the secure communication system as a Markov decision process and achieves multi-dimensional optimization through multi-agent and shared evaluation networks.

8. An electronic device, characterized in that, The device includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the UAV-assisted mobile edge computing optimization method according to any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the UAV-assisted mobile edge computing optimization method according to any one of claims 1 to 6 at runtime.

10. A computer program product, comprising a computer program, characterized in that, The computer program is configured to execute the UAV-assisted mobile edge computing optimization method as described in any one of claims 1 to 6 at runtime.