Multi-objective evolution depth deterministic strategy gradient method for mobile edge network optimization

Through the multi-objective evolution deep deterministic strategy gradient method, the multi-objective problem of the drone-assisted mobile edge computing network is optimized, and the problem of difficult to balance energy consumption, delay and task completion in the existing technology is solved, and the comprehensive improvement of network performance and user service quality is achieved.

CN120050710APending Publication Date: 2025-05-27CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235864.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the prior art optimizes the drone-assisted mobile edge computing network, it is difficult to effectively balance multiple optimization goals such as energy consumption, delay and task completion, resulting in difficulty in comprehensively optimizing network performance and user service quality.

Method used

The multi-objective evolution deep deterministic strategy gradient method is adopted to optimize the user delay, system energy consumption and drone task completions of the drone assisted mobile edge computing network by designing a multi-objective optimization model, a deep deterministic strategy gradient algorithm and a multi-objective evolution algorithm framework.

Benefits of technology

Significantly reduce user delay and system energy consumption, and increase the number of tasks completed by drones, improving the comprehensive performance and user service quality of mobile edge computing networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050710A_ABST
    Figure CN120050710A_ABST
Patent Text Reader

Abstract

The invention provides a multi-objective evolution depth deterministic strategy gradient method for mobile edge network optimization, and the method comprises the following steps: 1, constructing an integrated multi-objective optimization model which aims at minimizing the time delay and energy consumption of a mobile edge computing network assisted by an unmanned aerial vehicle, maximizing the task completion number of the unmanned aerial vehicle, and reducing the time delay and energy consumption of the mobile edge computing network; the overall performance of the network is ensured; and step 2, providing a multi-objective evolution depth deterministic strategy gradient method which combines a multi-objective evolution algorithm and a depth deterministic strategy gradient algorithm to carry out optimization solution on the model established in the step 1. According to the invention, the efficient communication service is provided for the user through the mobile edge computing network assisted by the unmanned aerial vehicle, and the comprehensive performance of the mobile edge computing network is remarkably improved. The method not only provides a new thought and technical support for the development of the mobile edge computing network, but also opens up a wide prospect for the application of the mobile equipment such as the unmanned aerial vehicle in the mobile edge computing network in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) control, and particularly to a multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization. Background Art

[0002] With the rapid development of Internet of Things (IoT) technology, the application scenarios of IoT terminals have been continuously expanding, and the demand for computationally intensive tasks has also increased sharply. The UAV-assisted mobile edge computing network can significantly reduce the task computing delay and improve the resource utilization efficiency by offloading computing tasks to edge nodes for processing, thus providing an efficient computing solution for IoT terminals and attracting much attention. However, in practical applications, the efficient operation of the mobile edge computing network faces challenges of multiple optimization objectives such as high energy consumption and large delay. Existing optimization algorithms mainly formulate it as a single-objective or multi-objective optimization problem when dealing with the optimization challenges of the UAV-assisted mobile edge computing network. However, the single-objective optimization method only focuses on a single performance index such as energy consumption or delay, and it is difficult to achieve the overall optimal performance of the mobile edge computing network. At the same time, although the multi-objective optimization method attempts to consider multiple aspects of performance, it often simplifies it to a single-objective problem through a weighting strategy and often fails to achieve an effective balance between various objectives. Therefore, this patent constructs a multi-objective optimization model for a UAV-assisted mobile edge computing network, aiming to simultaneously minimize the user delay and system energy consumption, and maximize the number of UAV task completions, thereby effectively ensuring the network performance and user service quality. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization, which greatly enhances the overall efficiency of the UAV-assisted mobile edge network and lays a solid technical foundation for future mobile network optimization work.

[0004] To solve the above technical problem, the technical solution adopted by the present invention is as follows:

[0005] The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization includes the following steps: Step 1, establish a multi-objective optimization model for a UAV-assisted mobile edge computing network, including a user delay model, an energy consumption model, and a total number of UAV task completion models;

[0006] Step 2, design a multi-objective evolutionary deep deterministic policy gradient method, including: the design of state, action, and reward functions, the deep deterministic policy gradient algorithm framework, and the multi-objective evolutionary algorithm framework; optimize and solve the multi-objective optimization model in Step 1 through the designed algorithm, greatly reducing the user delay, system energy consumption, and increasing the number of UAV task completions.

[0007] In the above Step 1, the steps for establishing the user delay model include:

[0008] Step 1.1. Establish the user delay model

[0009] Since the computing resources of local users are limited, the local user data is offloaded to the drone for processing; therefore, the user delay includes offloading delay and local computing delay;

[0010] The offloading delay is the transmission time when the user offloads data to the drone after establishing communication with the drone, and the specific definition is as follows:

[0011]

[0012] Among them, β is the amount of data of the user, O r is the offloading ratio, V t is the information transmission rate between the user and the drone, and its expression is as follows:

[0013]

[0014] Among them, B is the channel bandwidth, P s is the transmission power of the drone, PL is the path loss of the signal channel, σ n is the noise power;

[0015] The local computing delay is the time required for the user to compute data locally, and the specific definition is as follows:

[0016]

[0017] Among them, β is the amount of data of the user, O r is the offloading ratio, f l is the local computing frequency;

[0018] Finally, the user delay model is established as:

[0019]

[0020] Among them, N is the number of users served by the drone within the time period.

[0021] In the above Step 1, the steps for establishing the energy consumption model include:

[0022] Step 1.2. Establish the system energy consumption model

[0023] The system energy consumption includes: offloading energy consumption, local computing energy consumption, and drone flight energy consumption;

[0024] The offloading energy consumption is the energy consumption when the user offloads data to the drone, and the specific definition is as follows:

[0025] E o = L o ·P s ; (5)

[0026] Among them, L o is the unloading delay, and P s is the transmission power of the UAV;

[0027] The local computing energy consumption is the energy consumption generated by the user's local computing, and the specific definition is as follows:

[0028] E l = k·β·(f l ) 2 L l ; (6)

[0029] Among them, k is the effective capacitance coefficient of the user's chip, β is the amount of data of the user, f l is the local computing frequency, and L t is the local computing delay;

[0030] The UAV flight energy consumption is:

[0031]

[0032] Among them, M is the number of flight time slots of the UAV, v t represents the flight speed of the UAV, P a and U tip are the sectional power and tip speed of the rotor blade respectively, P b and v 0 represent the induced power and the average induced speed of the rotor during hovering respectively, d z is the fuselage drag coefficient, ρ is the air density, r z is the rotor natural ratio, and μ is the rotor area;

[0033] Finally, the established system energy consumption model is:

[0034]

[0035] In the above Step1, the steps for establishing the UAV total task completion model include:

[0036] Step1.3. Establish the UAV total task completion model

[0037] The total number of user tasks completed by the UAV is the number of tasks unloaded by the user to the UAV, and the specific definition is:

[0038]

[0039] Among them, β is the amount of data of the user, O rLet $\lambda$ be the offloading ratio and $M$ be the number of flight time slots of the UAV.

[0040] In the above Step 1, the steps of establishing the multi-objective optimization model of the UAV-assisted mobile edge computing network include:

[0041] Step1.4. Establish the multi-objective optimization model of the UAV-assisted mobile edge computing network:

[0042] min ($L$ all , $E$ all , $C - T$ all ); (10)

[0043] where $C$ represents a constant; to be consistent with the problem of minimizing energy consumption and latency, the problem of maximizing the number of tasks is transformed into a minimization problem.

[0044] In the above Step 2, the steps of state design include:

[0045] State design; the state is defined as:

[0046]

[0047] where and respectively represent the horizontal and vertical coordinates of the current position of the UAV, represents the number of tasks collected by the UAV in time slot $t$.

[0048] In the above Step 2, the steps of action design include:

[0049] Action design; the action is defined as:

[0050]

[0051] where is the flight angle, $d$ t is the flight distance, $O$ t is $O$ t is the offloading decision parameter.

[0052] In the above Step 2, the steps of reward function design include:

[0053] Reward function design; the reward function is used to guide the agent to make the optimal decision based on the current state, and it is designed as:

[0054]

[0055] where $L$ t is the latency of the UAV in time slot $t$, $E$ t is the energy consumption of the UAV in time slot $t$, $T$ tD Let \(N_t\) be the number of tasks completed by the UAV in time slot \(t\), \(j\) be the penalty factor, and \(\Theta\) denote that the UAV is within the specified flight area.

[0056] In the above Step 2, the steps of designing the deep deterministic policy gradient algorithm framework include the design of the agent, the calculation of the target Q value, the value network, the policy network, the target value network, and the update of the target policy network:

[0057] The design of the agent is \((\pi θ ξ , Q θ ζ , \pi θ ξ' , Q θ ζ' ), where the policy network \(\pi θ ξ is used to select actions; the value network \(Q θ ζ is used to evaluate the value of the state-action pair; the target policy network \(\pi θ ξ' and the target value network \(Q θ ζ' are used to calculate the target Q value; the policy network interacts with the environment to obtain the sample pair \((s t , a t , r t , s t+1 ) and stores it in the experience replay pool. When the experience replay pool reaches the set capacity, the above four networks are updated by sampling a small batch of sample pairs;

[0058] Calculate the target Q value according to the sampled sample pair, and its definition is:

[0059] Q tar = r t + \gamma Q θ ζ' (s t+1 , \pi θ ξ' (s t+1 )); (14)

[0060] where \(\gamma\) is the discount factor; \(Q θ ζ' (s t+1 , \pi θ ξ' (s t+1 )) represents the Q value output by the target value network; \(\pi θ ξ' (s t+1 ) represents the action output by the target policy network;

[0061] The value network adopts multi-objective state-action value function update, which is defined as:

[0062]

[0063] Among them, and are respectively the weighted sum of the target Q value, the estimated Q value and the weight vector w;

[0064] The update of the policy network is adjusted according to the Q value output by the value network, and is updated by multiplying the gradient by the learning rate. The specific update method is as follows:

[0065]

[0066] Among them, π θ ξ (s t ) is the deterministic action output by the policy network; is the gradient of the value network for this deterministic action;

[0067] The soft update strategy is adopted to update the target value network. The specific update methods are respectively:

[0068] θ ξ' = εθ ξ +(1 - ε)θ ξ' ; (17)

[0069] Among them, ε is the soft update ratio, which is used to ensure that the parameters of the target policy network and the target value network can approach the parameters of the policy network and the value network smoothly and gradually during each update;

[0070] The soft update strategy is adopted to update the target policy network. The specific update methods are respectively:

[0071] π θ ζ' = επ θ ζ'+(1 - ε)π θ ζ'; (18)

[0072] Among them, ε is the soft update ratio.

[0073] In the above Step2, the design steps of the multi-objective evolutionary algorithm framework include the definition of individuals, the design of the two-way selection strategy, the preparation stage and the evolutionary stage of the multi-objective evolutionary method:

[0074] Definition of individuals:

[0075] An individual is γ = (w, agent), where w is the weight vector and agent is the agent in the deep deterministic policy gradient algorithm;

[0076] Design of the two-way selection strategy:

[0077] The two-way selection strategy uses the Chebyshev distance and the vertical distance to match the weight vector with the individual one by one:

[0078] First, in the process of matching the weight vector with the individual, the Chebyshev distance is used to evaluate the closeness between each weight vector and the individual, and the individual with the smallest distance is selected as the optimal match for the current weight vector; the definition of the Chebyshev distance is as follows:

[0079]

[0080] In the formula, m is the number of optimization objectives, which is 3 in this patent; is the ideal point; F(π θ ξ )=(f 1 (π θ ξ ),...,f m (π θ ξ ) is the target vector of the policy network;

[0081] Secondly, a vertical distance from the individual to the weight vector is designed to endow the individual with the right to select the weight vector, and the weight vector with the smallest vertical distance from the individual is selected for matching to solve the problem of reducing population diversity caused by multiple weight vectors simultaneously selecting the same individual; the vertical distance is defined as follows:

[0082]

[0083] In the formula, is the normalized target vector, and w is the weight vector;

[0084] Preparation stage of the multi-objective evolutionary method:

[0085] Randomly initialize n agents. The performance of these agents is relatively poor in the initial stage and needs to be improved through multiple interactions with the environment for learning; therefore, n weight vectors are randomly combined with these agents to obtain n initial individuals, and the depth deterministic policy gradient algorithm is used to iteratively update Φ times to obtain n·Φ new individuals, and all the new individuals are added to the initial population;

[0086] Evolution stage of the multi-objective evolutionary method:

[0087] Use the two-way selection strategy to match each individual in the initial population with the weight vector one by one to determine the evolution direction of each individual; the offspring individuals after matching will continue to participate in the evolution process of the next generation. At the same time, the deep deterministic policy gradient algorithm is still used to continuously update the agent; this update process will continue in each generation of evolution until the preset maximum number of evolution generations is reached; finally, the Pareto solution set is output, which includes the optimal solutions corresponding to different weight vectors, and these optimal solutions are the optimized user delay, energy consumption, and the number of tasks completed by the UAV.

[0088] A multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization provided by the present invention combines a multi-objective evolutionary algorithm and a deep deterministic policy gradient algorithm to optimize and solve the model established in step 1. The present invention provides efficient communication services for users through a UAV-assisted mobile edge computing network, significantly improving the comprehensive performance of the mobile edge computing network. This invention not only provides new ideas and technical support for the development of mobile edge computing networks, but also opens up broad prospects for the application of future mobile devices such as UAVs in mobile edge computing networks. Brief Description of the Drawings

[0089] The following further illustrates the present invention in conjunction with the drawings and embodiments:

[0090] Figure 1 It is a schematic structural diagram of the multi-objective evolutionary deep deterministic policy gradient method for the mobile edge computing network of the present invention;

[0091] Figure 2 It is a schematic diagram of the UAV trajectory obtained by the multi-objective evolutionary deep deterministic policy gradient method under different weight vectors in the embodiment;

[0092] Figure 3 It is a schematic diagram of the comparison results of the Pareto solutions obtained by the multi-objective evolutionary deep deterministic policy gradient method and other methods in terms of delay, energy consumption, and the number of tasks completed in the embodiment. Detailed Embodiment

[0093] The technical solution of the present invention is described in detail below in conjunction with the drawings and embodiments.

[0094] Embodiment 1:

[0095] As Figures 1-3As shown in the figure, a multi-objective evolutionary deep deterministic policy gradient method for a mobile edge computing network. First, a multi-objective optimization model integrating the task delay, system energy consumption, and the number of tasks completed by the unmanned aerial vehicle (UAV) in the mobile edge computing network is established. Secondly, a multi-objective evolutionary strategy and a deep deterministic policy gradient algorithm for optimizing and solving the above model are proposed. By organically integrating the multi-objective evolutionary algorithm and the deep deterministic policy gradient algorithm, the overall performance of the UAV-assisted mobile edge computing network can be optimized in real time, thereby providing users with better communication service quality. The method includes the following steps:

[0096] Step 1: Establish a multi-objective optimization model for the UAV-assisted mobile edge computing network, specifically including a delay model, an energy consumption model, and a task completion number model.

[0097] Step 1.1: Establish a user delay model.

[0098] Since the computing resources of local users are limited, it is considered to offload local user data to the UAV for processing. Therefore, the user delay mainly includes the offloading delay and the local computing delay.

[0099] The offloading delay is the transmission time when the user offloads data to the UAV after establishing communication with the UAV, and is specifically defined as follows:

[0100]

[0101] Among them, β is the amount of data of the user, O r is the offloading ratio, and V t is the information transmission rate between the user and the UAV, and its expression is as follows:

[0102]

[0103] Among them, B is the channel bandwidth, P s is the transmission power of the UAV, PL is the path loss of the signal channel, and σ n is the noise power.

[0104] The local computing delay is the time required for the user to compute data locally, and is specifically defined as follows:

[0105]

[0106] Among them, β is the amount of data of the user, O r is the offloading ratio, and f l is the local computing frequency.

[0107] Finally, the delay model of the user is established as:

[0108]

[0109] Among them, N is the number of users served by the UAV within the time period.

[0110] Step 1.2: Establish a system energy consumption model.

[0111] The system energy consumption mainly includes: offloading energy consumption, local computing energy consumption, and UAV flight energy consumption

[0112] The offloading energy consumption is the energy consumption when users offload data to the UAV, and the specific definition is as follows:

[0113] E o = L o · P s ; (5)

[0114] Among them, L o is the offloading delay, and P s is the transmission power of the UAV.

[0115] The local computing energy consumption is the energy consumption generated by local computing of users, and the specific definition is as follows:

[0116] E l = k · β · (f l ) 2 L l ; (6)

[0117] Among them, k is the effective capacitance coefficient of the user chip, β is the amount of data of the user, f l is the local computing frequency, and L t is the local computing delay.

[0118] The UAV flight energy consumption is:

[0119]

[0120] Among them, M is the number of flight time slots of the UAV, v t represents the flight speed of the UAV, P a and U tip are the sectional power and tip speed of the rotor blade respectively, P b and v 0 represent the induced power and the average induced speed of the rotor during hovering respectively, d z is the fuselage drag coefficient, ρ is the air density, r z is the rotor natural ratio, and μ is the rotor area.

[0121] Finally, the established system energy consumption model is:

[0122]

[0123] Step 1.3: Establish a model for the total number of tasks completed by the UAV.

[0124] The total number of user tasks completed by the UAV is the number of tasks unloaded by the user to the UAV, and the specific definition is:

[0125]

[0126] where β is the amount of data of the user, O r is the offloading ratio, and M is the number of flight time slots of the UAV.

[0127] Step 1.4: Establish a multi-objective optimization model for the UAV-assisted mobile edge network as:

[0128] min(L all ,E all ,C - T all ); (10)

[0129] where C represents a very large constant. To be consistent with the problem of minimizing energy consumption and latency, the problem of maximizing the number of tasks is transformed into a minimization problem.

[0130] (Step 2 According to Step 1, design a multi-objective evolutionary deep deterministic policy gradient method)

[0131] Step 2: Design a multi-objective evolutionary deep deterministic policy gradient method, which specifically includes: the design of state, action and reward functions, the framework of the deep deterministic policy gradient algorithm, and the framework of the multi-objective evolutionary algorithm.

[0132] Step 2.1: Design of state, action and reward functions.

[0133] Step 2.1.1: State design. In the UAV-assisted mobile edge computing network, the UAV acts as an agent to interact with the environment and makes optimal decisions based on the acquired state information. The state information includes the position information and task information of the UAV, so that the UAV can reasonably plan the flight path and optimize the offloading strategy. Therefore, the state is defined as:

[0134]

[0135] where and respectively represent the horizontal and vertical coordinates of the current position of the UAV, represents the number of tasks collected by the UAV in the t time slot.

[0136] Step 2.1.2: Action design. The flight action decision of the UAV needs to comprehensively consider the flight angle and flight distance to ensure efficient trajectory planning and task execution. At the same time, the task processing method involves the optimization of the task offloading ratio, so as to effectively balance the UAV and local computing resources. Therefore, the action is defined as:

[0137]

[0138] Step 2.1.3: Reward function design. The reward function is mainly used to guide the agent to make optimal decisions based on the current state, and it is designed as:

[0139]

[0140] where L t is the delay of the UAV at time slot t, E t is the energy consumption of the UAV at time slot t, is the number of tasks completed by the UAV at time slot t, j is the penalty factor, and Θ represents that the UAV is within the specified flight area.

[0141] Step 2.2: Design of the deep deterministic policy gradient algorithm framework. Specifically, it includes: the design of the agent, the calculation of the target Q value, the value network, the policy network, the target value network, and the update of the target policy network.

[0142] Step 2.2.1: The agent is set as (π θ ξ , Q θ ζ , π θ ξ' , Q θ ζ' ), where the policy network π θ ξ is used to select actions; the value network Q θ ζ is used to evaluate the value of the state-action pair; the target policy network π θ ξ' and the target value network Q θ ζ' are used to calculate the target Q value. The policy network interacts with the environment to obtain the sample pair (s t , a t , r t , s t+1 ) and stores it in the experience replay pool. When the experience replay pool reaches a certain capacity, the above four networks are updated by sampling a small batch of sample pairs.

[0143] Step 2.2.2: Calculate the target Q value according to the sampled sample pair, and its definition is:

[0144] Q tar = r t + γQ θ ζ' (s t+1 , π θ ξ' (st+1 )); (14)

[0145] where γ is the discount factor; Q θ ζ' (s t+1 , π θ ξ' (s t+1 )) represents the Q-value output by the target value network; π θ ξ' (s t+1 ) represents the action output by the target policy network.

[0146] Step 2.2.3: The value network is updated using a multi-objective state-action value function, which is defined as:

[0147]

[0148] where and are respectively the weighted sums of the target Q-value, the estimated Q-value, and the weight vector w.

[0149] Step 2.2.4: The update of the policy network is adjusted according to the Q-value output by the value network, and is updated by multiplying the gradient by the learning rate. The specific update method is as follows:

[0150]

[0151] where π θ ξ (s t ) is the deterministic action output by the policy network; is the gradient of the value network for this deterministic action.

[0152] Step 2.2.5: The target value network is updated using a soft update strategy. The specific update methods are respectively:

[0153] θ ξ' = εθ ξ + (1 - ε)θ ξ' ; (17)

[0154] where ε is the soft update ratio, which is used to ensure that the parameters of the target policy network and the target value network can approach the parameters of the policy network and the value network smoothly and gradually during each update.

[0155] Step 2.2.6: The target policy network is updated using a soft update strategy. The specific update methods are respectively:

[0156] π θ ζ' = επ θ ζ' + (1 - ε)π θζ'; (18)

[0157] where ε is the soft update ratio.

[0158] Step 2.3: Design of the multi-objective evolutionary algorithm. Specifically, it includes: definition of individuals, design of the two-way selection strategy, preparation stage and evolutionary stage of the multi-objective evolutionary method.

[0159] Step 2.3.1: Definition of individuals.

[0160] An individual is γ = (w, agent), where w is the weight vector and agent is the agent in the deep deterministic policy gradient algorithm.

[0161] Step 2.3.2: The two-way selection strategy uses the Chebyshev distance and the perpendicular distance to match the weight vector with the individual one by one.

[0162] First, in the process of matching the weight vector with the individual, the Chebyshev distance is used to evaluate the closeness between each weight vector and the individual, and the individual with the smallest distance is selected as the optimal match for the current weight vector. The definition of the Chebyshev distance is as follows:

[0163]

[0164] In the formula, m is the number of optimization objectives, which takes the value of 3 in this patent; is the ideal point; F(πθ ξ ) = (f 1 (π θ ξ ),..., f m (π θ ξ ) is the objective vector of the policy network.

[0165] Secondly, a perpendicular distance from the individual to the weight vector is designed to endow the individual with the right to select the weight vector, and the weight vector with the smallest perpendicular distance from the individual is selected for matching to solve the problem of reduced population diversity caused by multiple weight vectors simultaneously selecting the same individual. The definition of the perpendicular distance is as follows:

[0166]

[0167] In the formula, is the normalized objective vector and w is the weight vector.

[0168] Step 2.3.3: Preparation stage of the multi-objective evolutionary method.

[0169] Randomly initialize n agents. The performance of these agents is relatively poor in the initial stage and needs to be improved by interacting with the environment multiple times for learning. Therefore, randomly combine n weight vectors with these agents to obtain n initialized individuals, and iteratively update them Φ times through the Deep Deterministic Policy Gradient algorithm to get n·Φ new individuals, and add all the new individuals to the initial population.

[0170] Step 2.3.4: The evolution stage of the multi-objective evolution method.

[0171] Use the two-way selection strategy to match the individuals in the initial population with the weight vectors one by one to determine the evolution direction of each individual. The offspring individuals after matching will continue to participate in the evolution process of the next generation. At the same time, the Deep Deterministic Policy Gradient algorithm is still used to continuously update the agents. This update process will continue in each generation of evolution until the preset maximum number of evolution generations is reached; finally, output the Pareto solution set, which includes the optimal solutions corresponding to different weight vectors, and these optimal solutions are the optimized user delay, energy consumption, and the number of tasks completed by the UAV.

[0172] The present invention provides a multi-objective evolution Deep Deterministic Policy Gradient method for a mobile edge computing network. First, an integrated multi-objective optimization model is constructed, which aims to comprehensively improve the network performance by minimizing the delay and energy consumption of the UAV-assisted mobile edge computing network and simultaneously maximizing the number of task executions of the UAV. Secondly, in view of the challenge that the traditional Deep Deterministic Policy Gradient algorithm faces in effectively balancing the weight distribution among various objectives when dealing with multi-objective optimization tasks, a two-way selection strategy is innovatively proposed for the precise matching of weight vectors and individuals, thereby significantly enhancing the diversity of the population. Finally, on the basis of deeply integrating the multi-objective evolution algorithm and the Deep Deterministic Policy Gradient algorithm, a novel multi-objective evolution Deep Deterministic Policy Gradient algorithm is developed. This algorithm can dynamically and real-time optimize the overall performance of the UAV-assisted mobile edge computing network. By implementing the method proposed by the present invention, the overall efficiency of the UAV-assisted mobile edge network will be greatly enhanced, laying a solid technical foundation for future mobile network optimization work.

Claims

1. A multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization, characterized by: The steps include: Step 1: Establish a multi-objective optimization model for the UAV-assisted mobile edge computing network, including a user latency model, an energy consumption model, and a model for the total number of tasks completed by the UAV; Step 2. Design a multi-objective evolutionary deep deterministic policy gradient method, including: the design of state, action and reward functions, the deep deterministic policy gradient algorithm framework, and the multi-objective evolutionary algorithm framework; optimize and solve the multi-objective optimization model in Step 1 through the designed algorithm, thereby reducing user latency, system energy consumption and increasing the number of tasks completed by the drone.

2. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the above-mentioned Step 1, the steps of establishing the user delay model include: Step 1.

1. Establish user delay model Since local users have limited computing resources, local user data is offloaded to the drone for processing; therefore, user latency includes offloading latency and local computing latency; Unloading latency is the transmission time when the user unloads data to the drone after the user establishes communication with the drone. The specific definition is as follows: Among them, β is the amount of data of the user, O r is the unloading ratio, V t is the information transmission rate between the user and the drone, and its expression is as follows: Where B is the channel bandwidth, P s is the transmission power of the UAV, PL is the channel path loss, σ n is the noise power; Local computing latency is the time required for users to calculate data locally. The specific definition is as follows: Among them, β is the amount of data of the user, O r is the unloading ratio, f l Calculate frequency for local area; Finally, the user's delay model is established as: Where N is the number of users served by the drone within a time period.

3. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the above-mentioned Step 1, the steps of establishing the energy consumption model include: Step 1.2: Establish system energy consumption model System energy consumption includes: unloading energy consumption, local computing energy consumption and UAV flight energy consumption; Unloading energy consumption is the energy consumption of users unloading data to drones, which is defined as follows: AND o =L o ·P s ; (5) Among them, L o is the unloading delay, P s is the transmitting power of the UAV; Local computing energy consumption refers to the energy consumption generated by local computing of users. The specific definition is as follows: E l =k·β·(f l ) 2 L l (6) Where k is the effective capacitance coefficient of the user chip, β is the amount of user data, and f l is the local calculation frequency, L t Calculate latency locally; The energy consumption of drone flight is: Where M is the number of flight time slots of the drone, v t Indicates the flight speed of the drone, P a and U tip are the sectional power and tip speed of the rotor blade, P b and v0 represent the induced power and the average induced speed of the rotor in hovering, respectively. z is the fuselage drag coefficient, ρ is the air density, r z is the rotor inherent ratio, μ is the rotor area; Finally, the established system energy consumption model is:

4. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the above-mentioned Step 1, the steps of establishing the model of the total number of tasks completed by the drone include: Step 1.

3. Establish a model for the total number of tasks completed by drones The total number of user tasks completed by the drone is the number of tasks unloaded by the user to the drone, which is specifically defined as: Among them, β is the amount of data of the user, O r is the unloading ratio, and M is the number of flight time slots of the UAV.

5. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the above-mentioned Step 1, the steps of establishing a multi-objective optimization model for a drone-assisted mobile edge computing network include: Step 1.

4. Establish a multi-objective optimization model for drone-assisted mobile edge computing network: min(L all ,E all ,CT all );(10)Where C represents a constant; In order to be consistent with the problem of minimizing energy consumption and delay, the problem of maximizing the number of tasks is transformed into a minimization problem.

6. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the aforementioned Step 2, the state design steps include: State design; states are defined as: in, and Respectively represent the horizontal and vertical coordinates of the current position of the drone, Represents the number of tasks collected by the drone in time slot t.

7. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the aforementioned Step 2, the steps of action design include: Action design; actions are defined as: in, is the flight angle, d t is the flight distance, O t is the uninstall decision parameter.

8. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In Step 2, the reward function design steps include: Reward function design: The reward function is used to guide the agent to make the best decision based on the current state. It is designed as follows: Among them, L t is the delay of the UAV in time slot t, E t is the energy consumption of the UAV in time slot t, is the number of tasks completed by the UAV in time slot t, j is the penalty factor, and Θ indicates that the UAV is within the specified flight area.

9. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the above Step 2, the steps of designing the deep deterministic policy gradient algorithm framework include the design of the agent, the calculation of the target Q value, the value network, the policy network, the target value network and the target policy network update: The design of the intelligent agent is Among them, the policy network Used to select actions; value network Used to evaluate the value of state-action pairs; target policy network and target value network Used to calculate the target Q value; the policy network interacts with the environment to obtain (s t ,a t ,r t ,s t+1 ) After the sample pairs are collected, they are stored in the experience replay pool. When the experience replay pool reaches the set capacity, the above four networks are updated by sampling small batches of sample pairs; The target Q value is calculated based on the sample pairs obtained by sampling, which is defined as: Q tar =r t +γQ θ ζ' (s t+1 ,πθ ξ' (s t+1 )); (14) Among them, γ is the discount factor; Represents the Q value output by the target value network; represents the action output by the target policy network; The value network is updated using a multi-objective state-action value function, which is defined as: in, and They are the weighted sum of the target Q value, the estimated Q value and the weight vector w; The update of the policy network is adjusted according to the Q value output by the value network, and is updated by multiplying the gradient by the learning rate. The specific update method is as follows: in, is the deterministic action output by the policy network; is the gradient of the value network for this deterministic action; A soft update strategy is used to update the target value network. The specific update methods are: i ξ '=θ ξ +(1-e)θ ξ '; (17) Among them, ε is the soft update ratio, which is used to ensure that the parameters of the target policy network and the target value network can smoothly and gradually approach the parameters of the policy network and the value network at each update; A soft update strategy is used to update the target strategy network. The specific update methods are: Among them, ε is the soft update ratio.

10. The multi-objective evolutionary deep deterministic policy gradient method for mobile edge network optimization according to claim 1, characterized in that: In the aforementioned Step 2, the multi-objective evolutionary algorithm framework design steps include the definition of individuals, the design of a two-way selection strategy, the preparation phase and the evolution phase of the multi-objective evolutionary method: Definition of Individual: The individual is γ = (w, agent), where w is the weight vector and agent is the agent in the deep deterministic policy gradient algorithm; Two-way selection strategy design: The two-way selection strategy uses Chebyshev distance and vertical distance to match weight vectors with individuals one by one: First, in the process of matching the weight vector with the individual, the Chebyshev distance is used to evaluate the closeness between each weight vector and the individual, and the individual with the smallest distance is selected as the optimal match for the current weight vector; the definition of the Chebyshev distance is as follows: Where m is the number of optimization targets, and the value is 3 in this patent; is the ideal point; is the target vector of the policy network; Secondly, a vertical distance from an individual to a weight vector is designed to give the individual the right to choose a weight vector, and the weight vector with the smallest vertical distance to the individual is selected for matching, solving the problem of reduced population diversity caused by multiple weight vectors selecting the same individual at the same time; the vertical distance is defined as follows: In the formula, is the normalized target vector, w is the weight vector; Preparation stage of multi-objective evolutionary method: Randomly initialize n agents. The performance of these agents is relatively poor in the initial stage, and they need to improve their quality through multiple interactive learning with the environment; therefore, n weight vectors are randomly combined with these agents to obtain n initialized individuals, and the deep deterministic policy gradient algorithm is iteratively updated Φ times to obtain n·Φ new individuals, and all new individuals are added to the initial population; Evolutionary stages of multi-objective evolutionary methods: A two-way selection strategy is used to match individuals in the initial population with weight vectors one by one to determine the evolutionary direction of each individual; the matched offspring individuals will continue to participate in the evolutionary process of the next generation. At the same time, the deep deterministic policy gradient algorithm is still used to continuously update the intelligent agent; this update process will continue in each generation of evolution until the preset maximum number of evolution generations is reached; finally, the Pareto solution set is output, which includes the optimal solutions corresponding to different weight vectors. These optimal solutions are the optimized user latency, energy consumption and number of tasks completed by the drone.