Mobile edge computing unloading strategy and resource optimization method based on physical layer security
By building a task offload model and a secure transmission channel model, combining deep reinforcement learning methods, dynamically adjusting the task offload ratio and artificial noise power, the security and resource efficiency problems of wireless offloading in mobile edge computing are solved, and an efficient and secure mobile edge computing system is realized.
Patent Information
- Application Number
- CN202411468412.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the industrial Internet of Things, the wireless offload mechanism of mobile edge computing increases the risk of user data being eavesdropped, and the existing technology is difficult to solve the problem of improving task resource efficiency and reliability, especially in terms of taking into account secure transmission and resource optimization.
Using a mobile edge computing offload strategy based on physical layer security, we use the task offload model and a secure transmission channel model to build a deep reinforcement learning method, dynamically adjust the task offload ratio and artificial noise power, optimize the delay and energy consumption performance, and ensure the safe transmission of the task.
It realizes efficient resource allocation and secure information transmission in the industrial Internet of Things environment, significantly reduces system delay and energy consumption, has the ability to fight eavesdropping and attacks, and provides an efficient and secure mobile edge computing system.
Smart Images

Figure CN120263789A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial Internet of things applications, and particularly relates to a mobile edge computing offloading strategy and resource optimization method based on physical layer security. Background Technique
[0002] In the industrial Internet of things, the application of mobile edge computing offloads computationally intensive tasks from resource-constrained terminals to edge servers, solving the problems of insufficient computing resources and excessive latency. However, the wireless offloading mechanism of mobile edge computing increases the risk of user data being eavesdropped.
[0003] With the rapid development of the industrial Internet of things (IIoT), the number of mobile devices has increased rapidly, and users' demands for diverse emerging services have also been growing continuously. The huge amount of data and service demands generated by user terminals have continuously increased the requirements for the computing power of the network. Relying solely on terminal devices to complete computationally intensive tasks and the like will face significant challenges. Computational offloading provides an effective solution for mobile users, enabling mobile users to offload their computationally intensive tasks to devices with sufficient computing resources, thereby effectively alleviating the limitations of terminal device computing resources. Compared with traditional cloud computing, mobile edge computing (MEC) allows resource-limited devices to offload some or all tasks to a mobile edge computing server (MES, MEC Server) near the base station for processing. It not only effectively alleviates the computing pressure of terminal devices but also avoids the network congestion problems that may be caused by traditional centralized computing and the high latency caused by long-distance data transmission, meeting the actual needs of latency-sensitive tasks.
[0004] Although the task offloading mechanism in mobile edge computing has significantly improved system efficiency, the subsequent wireless transmission security challenges cannot be ignored. Given its inherent broadcast nature, data faces a high risk of being intercepted by unauthorized third parties during the transmission process between user devices and MES. How to ensure the security of wireless transmission between terminal devices and MES has become an urgent problem to be solved. In existing research, no offloading strategy solution has been proposed for heterogeneous task requirements in IIoT that takes into account both secure transmission and resource joint optimization, making it difficult to solve the problems of improving task resource efficiency and reliability in this scenario. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a mobile edge computing offloading strategy and resource optimization method based on physical layer security. According to the heterogeneous characteristics of tasks and the diversity changes of the environment, the task offloading ratio and artificial noise power are dynamically adjusted, aiming to optimize the delay and energy consumption performance simultaneously and ensure the secure transmission of tasks. The present invention models the optimization problem of the offloading strategy as a Markov decision process and uses the proximal policy optimization algorithm to solve the problem. Through numerical simulation, it is shown that this strategy can effectively cope with the dual challenges of efficient resource allocation and information security transmission in the IIoT environment, providing a new solution for realizing an efficient and secure mobile edge computing system.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: A mobile edge computing offloading strategy and resource optimization method based on physical layer security, and the specific steps are as follows:
[0007] Step 1. Modeling the secure offloading of mobile edge computing tasks;
[0008] Step 11). Construct a task offloading model: In the system, there are N user devices, K small base stations, and L eavesdroppers. Each small base station is equipped with an edge server. The set of user devices is represented as U = {1, 2, 3,..., N}, the set of edge servers is represented as M = {1, 2, 3,…, K}, and the set of eavesdroppers is represented as E = {1, 2,…, l,…, L}. Assume that user device u n generates one or more computationally intensive and delay-sensitive tasks Q n =(d n , c n , t n ) in each time slot t ∈ T. Among them, d n represents the number of generated computing tasks, c n represents the number of cycles for processing each bit of data on user u n , and t n represents the maximum tolerable delay for completing this task. Determine an offloading ratio λ n , and offload the task to the edge server for calculation. (1 - λ n )d n of the data is calculated locally, and the remaining λ n d n of the data is offloaded to the MES for calculation.
[0009] Step 12). Secure transmission channel model: The channel gain between the user device and the MES is represented as The channel gain between the user device and the eavesdropper is represented as The channel gain between the MES and the eavesdropper is represented as Among them, M rDenote the receiving antenna of the MES as M t Denote the transmitting antenna of the MES as M e Denote the receiving antenna of the eavesdropper. When the MES sends an interference signal, interference from the output to the input will be generated through the self-interference channel, denoted as Introduce a parameter ρ ∈ [0, 1] to represent the self-interference effect of the server full-duplex module. The smaller ρ is, the better the full-duplex system is. ρ = 0 represents the full-duplex system under ideal conditions. This system model takes into account large-scale fading and small-scale fading and there is Among them, the small-scale channel fading is generated by an independent complex Gaussian distribution with unit variance and captured by H. The large-scale channel fading is represented by denoted as and represent the path loss constant and the path loss exponent respectively. d UM represents the Euclidean distance between the user equipment and the MES, d UE represents the Euclidean distance between the user equipment and the eavesdropper, d ME represents the Euclidean distance between the MES and the eavesdropper;
[0010] When the user equipment offloads tasks, it sends a signal χ with a transmit power of p t The signal and the transmit power satisfy constraint. When the MES exchanges data with the user equipment, it sends an interference signal n J , the interference signal n J is obtained from a complex Gaussian distribution, denoted as Among them, the covariance matrix Q of the interference signal is a Hermitian matrix and Q ≥ 0. For a given Q, the power of the corresponding interference signal is p J , that is, tr(Q) = p J , assuming that the mobile device has only one antenna and the edge server adopts the maximum ratio combining (MRC) receiver given by r M = h UM / ||h UM ||, then the signal received by the edge server is:
[0011]
[0012] Among them, is the channel noise, is the noise power of the antenna. Assuming that the eavesdropper also uses the maximum ratio combining receiver r E = h UE / ||h UE ||, then the signal received by the eavesdropper is:
[0013]
[0014] Among them, is the corresponding channel noise, is the noise power;
[0015] Secrecy rate is expressed as:
[0016]
[0017] Among them, [v] + = max{v, 0}, represents the transmission rate of the task unloaded from the user equipment to the edge server, represents the task transmission rate between the user equipment and the eavesdropper, and is usually expressed as:
[0018]
[0019] Among them, B represents the uplink bandwidth of the user equipment and the MES, and the secrecy rate is designed by adjusting the power of the AN, H M (Q) and H E (Q) are defined as:
[0020]
[0021] t ∈ [0, τ] represents the time elapsed after making the offloading decision, ρ ∈ [0, 1] represents the self-interference effect, h UM represents the channel gain between the user equipment and the MES, h UE represents the channel gain between the user and the eavesdropper, h ME represents the channel gain between the MES and the eavesdropper;
[0022] Step 13): Calculate the task offloading model;
[0023] Step 131): Calculate the delay model;
[0024] The task ratio λ n ∈ [0, 1] is used to determine how much task data should be offloaded to the MES, and the generated tasks are divided into two parts: one part is executed locally, and the other part is offloaded to the MES for execution;
[0025] a. Local execution delay: When the mobile device processes the computing task locally, the processing time depends on the computing resources of the mobile device, and the local execution delay is expressed as:
[0026]
[0027] Among them, c n represents the number of CPU cycles required to process one data, CPU frequency processed by the local device;
[0028] b. Task offloading delay: When the tasks of the terminal are placed in the MES for computing offloading, there will be a transmission delay for the tasks to be transmitted from the user equipment to the MES, an execution delay for the MES to process the tasks, and a transmission delay for the processing results to be transmitted to the terminal. Without considering the feedback transmission time of the downlink, the total delay of task offloading consists of the data transmission delay Execution delay of the MES to process the tasks and is expressed as:
[0029]
[0030] From the local computing delay, the execution delay of the MES to process the tasks is:
[0031]
[0032] where is the CPU frequency of the MEC in the t-th time slot;
[0033] Step 132). Energy consumption model;
[0034] The total energy consumption is divided into two parts: the energy consumption required for local data processing and the transmission energy consumption during task offloading.
[0035] a. Local execution energy consumption; is the power consumption of the CPU chip, with a coefficient of k, depending on the CPU architecture. The local execution energy consumption is expressed as:
[0036]
[0037] b. Offloading transmission energy consumption
[0038] For the task offloading part, energy will be consumed during the data transmission process. During the offloading process, the offloading transmission energy consumption is expressed as:
[0039]
[0040] where p t is the transmission power of the user equipment when offloading tasks;
[0041] Step 14). Optimization objective establishment;
[0042] According to equations (1) to (10), the total delay of the system in t time slots is expressed as:
[0043]
[0044] The total energy consumption of the system in t time slots is expressed as:
[0045]
[0046] Optimize the time delay and energy consumption simultaneously to minimize the cost of task offloading calculation, and establish the optimal decision-making model as follows:
[0047]
[0048] Among them, β is the weight factor, and λ in C1 n is the task offloading ratio, C2 is the time delay constraint to ensure that the total time delay for task completion does not exceed the maximum tolerable time delay τ, C3 indicates that the computing resources allocated by the MES for the tasks offloaded to this server do not exceed the computing power of this edge server, C4 indicates the constraint on the transmission power p AN when the MES exchanges data with the Internet of Things device, and C5 indicates the constraint on limiting the transmission power of the user device;
[0049] Step 2: Solve the problem of deep reinforcement learning based on multi-objective optimization;
[0050] The DRL method is used to solve the above-mentioned minimization problem of time delay and energy consumption, which is modeled as a Markov decision process (MDP), and the optimal policy is learned from the training environment through the DRL method to solve the problem of minimizing the time delay and energy consumption of the system;
[0051] Step 21): MDP modeling;
[0052] Perceive the environmental state to obtain conditional information. Considering the random server situation and the dynamic task changes during data transmission, model the problem of task offloading and resource allocation of computing tasks as an MDP, and learn the optimal policy from the training environment through the DRL method to solve the problem of minimizing the weighted sum of delay and energy consumption. Define this process as a triple M=(S, A, R), where S represents the state space, A represents the action space, and R represents the reward function. In the MDP, the agent continuously interacts with the dynamic environment to optimize its own policy. The agent adjusts its own policy by observing the state s t+1 and the reward r t to maximize the cumulative reward and try to find the optimal computing offloading policy to minimize the problem of the weighted sum of time delay and energy consumption of the system. The basic three elements of the MDP are defined as follows:
[0053] Element 1: State space State;
[0054] To comprehensively consider the characteristics between device tasks and MES resources in the IIoT, the agent needs to observe the system state at the beginning of each time slot. Its state includes the tasks Q n =(d n , c n , t n) and the task offloading ratio λ n , represent the state space as:
[0055] s t ={d n ,c n ,t n ,λ n} (16)
[0056] where d n represents the amount of computational task data generated, c n represents the number of cycles to process each bit of data on user u n , t n represents the maximum latency that the task can tolerate;
[0057] Element 2. Action space Action;
[0058] The agent finds the optimal offloading strategy by allocating and offloading tasks for each time slot, and outputs the action a t according to the state space. At time slot t, the action a t ∈A t is defined as A t ={α,f,Q}, where α={x1(t),x2(t),…,x k (t)}, x n (t)∈{0,1,…,K}. When x n (t)=0, it means that the task offloading is executed locally. When x n (t)=n, it means that the task is offloaded to the nth server for execution, and the resource allocation decision represents the computing frequency allocated by the edge server to task Q n at time slot t;
[0059] Element 3. Reward Reward;
[0060] Define the reward as r(t)=-Cost, where Cost is the weighted sum of latency and energy consumption defined in formula (15);
[0061] Step 22), DRL training framework based on PPO;
[0062] Adopt the PPO algorithm to achieve secure task transmission and optimize task offloading decisions and resource allocation decisions. Use the PPO algorithm to solve this MDP problem, select appropriate offloading strategies under different states to optimize the system performance. The choice of the PPO algorithm is based on its excellent performance and efficient learning ability, which is especially suitable for dealing with complex problems with high-dimensional state spaces and action spaces.
[0063] The Actor network is an agent used to directly control the offloading scheme strategy. Its output actions determine whether a task is offloaded and the computing resources allocated to each task by the server. The agent observes the task offloading ratio and the tasks to be executed in the current time slot as the state, and then outputs the probability distribution of actions. The agent selects its next action by sampling from this probability distribution. The Critic network is used to approximate the value function, which is defined as:
[0064]
[0065] where γ is the discount factor, denotes the expected value, r(·) represents the reward function with respect to the state and actions. The role of this function is to evaluate the value of the state or the state-action pair, measure the quality of different strategies, estimate the expected return of different strategies in a certain state through the value function, and select the strategy that can improve the return, so as to improve the strategy to achieve the goal. In other words, Actor-Critic is a combination of policy optimization and value optimization. The Actor decides what action to take, while the Critic evaluates the action and indicates how the Actor should adjust. During the learning process, PPO stores the interactions between the agent and the environment in memory, and this memory will be used to update the Actor-Critic network. Different from the experience replay of DQN, the PPO algorithm uses all the memory to update the neural network. The updated neural network is used to collect new experiences and refill the memory. The gradient estimate is given by the following formula:
[0066]
[0067] where π θ are the parameters of the policy neural network, is the estimated value of the advantage function, is defined as:
[0068]
[0069] When performing multiple optimization steps on the policy, a large number of updates are required. To solve this problem and limit the new policy from deviating too far from the old policy, the ratio of probabilities is used to control the update amplitude of the new policy. Let ρ t (θ) represent the probability ratio between the new policy π θ and the old policy :
[0070]
[0071] If the difference between the new and old policies is significant and the advantage function is large, the update amplitude is appropriately increased: If the ρ t ratio is closer to 1, it indicates that the difference between the new and old policies is smaller. The objective function is expressed as:
[0072]
[0073] Among them, ε c is a hyperparameter between 0 and 1, clip is a truncation parameter, and the purpose of this function is to control the amplitude of the policy update, restricting the range of the policy update ratio within [1 - ε c , 1 + ε c ;
[0074] To provide a more stable estimate of the advantage function, improve the training stability and convergence speed of the algorithm, the Generalized Advantage Estimation (GAE) method is adopted, and the parameter λ is adjusted to find a balance between bias and variance, thereby improving the overall performance of the algorithm. It is:
[0075]
[0076] Among them, T is a given length, and λ is a hyperparameter. δ(t) is defined as:
[0077] δ(t) = r(t) + γV(s(t + 1)) - V(s(t)) (23)
[0078] Calculate the gradient of the target network according to Equation (21), and update the parameters θ and ξ through gradient descent to complete one round of iteration.
[0079] The present invention proposes a mobile edge computing task offloading strategy considering secure transmission in the IIoT environment, which can dynamically adjust the task offloading ratio and artificial noise power according to the heterogeneous characteristics of tasks and the diversity changes of the environment, optimize the system task processing delay, and reduce the energy consumption of the system. Especially in the face of the potential risk of eavesdropping and attacks during the transmission of computational task offloading, the proposed strategy will ensure secure transmission during the task offloading process. In the specific implementation process, the optimization problem of the offloading strategy is modeled as a Markov decision process, and the proximal policy optimization algorithm is used to solve the problem. The present invention demonstrates through simulation that this strategy can effectively address the dual challenges of efficient resource allocation and information secure transmission in the IIoT environment, providing a new solution for realizing an efficient and secure mobile edge computing system. The intelligent body trained by the present invention can generate resource allocation and policy optimization with relatively low complexity, has a relatively fast optimization convergence speed, and has good real-time processing capabilities. The proposed solution of the present invention can significantly reduce the system delay and energy consumption, and the present invention can effectively improve the security problem during the computational task offloading process and has the ability to resist eavesdropping and attacks. Description of the Drawings
[0080] Figure 1 It is a task offloading model diagram for mobile edge computing.
[0081] Figure 2 It is a diagram of the DRL training framework based on PPO.
[0082] Figure 3 It is a schematic diagram of the convergence of the algorithm under different learning rates.
[0083] Figure 4 It is a schematic diagram of the convergence of the algorithm under different mini-batch samples.
[0084] Figure 5 It is a schematic diagram of the change of the system average cost under different delay constraints. Figure 5 (a) is the system average delay under different τ, and 5(b) is the schematic diagram of the system average energy consumption under different τ. Figure 5 (c) is the schematic diagram of the system average cost under different τ.
[0085] Figure 6 It is a schematic diagram of the relationship between different weight parameters and the system average cost.
[0086] Figure 7 It is a schematic diagram of the convergence of the system reward under different strategies. Specific implementation manners
[0087] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0088] The mobile edge computing offloading strategy and resource optimization method based on physical layer security are as follows:
[0089] Step 1: Model the secure offloading of mobile edge computing tasks;
[0090] Step 11): Construct a task offloading model;
[0091] As Figure 1 shown, construct a task offloading model for mobile edge computing for IIoT. In this model, user devices and edge servers are randomly distributed in the same area. Users can choose a single edge server to access and perform computing offloading. Each user device has a certain degree of local computing ability, and the user device offloads some or all of the tasks to the edge server for collaborative computing. Considering the existence of potential eavesdroppers, their eavesdropping behavior will put user information at risk of leakage. To solve this problem, the edge server emits interference signals (artificial noise) to actively prevent eavesdropping and help users achieve secure offloading. It is assumed that the edge server has a full-duplex module, which can effectively reduce the self-interference of the interference signal.
[0092] In the system, there are N user devices, K small base stations, and L eavesdroppers. Each small base station is equipped with an edge server. The set of user devices is denoted as U = {1, 2, 3,..., N}, the set of edge servers is denoted as M = {1, 2, 3,…, K}, and the set of eavesdroppers is denoted as E = {1, 2,…, l,…, L}. Assume that user device u n generates one or more computation-intensive and latency-sensitive tasks Q n =(d n , c n , t n ) in each time slot t ∈ T, where d n represents the number of generated computation tasks (in bits), c n represents the number of cycles to process each bit of data on user u n , and t n represents the maximum tolerable latency to complete the task. In the designed model, the partial offloading mode is studied. To save time and energy, a suitable offloading ratio λ n is determined, and the task is offloaded to the edge server for computation. (1 - λ n )d n of the data is computed locally, and the remaining λ n d n of the data is offloaded to the MES for computation.
[0093] Step 12), secure transmission channel model;
[0094] The channel gains between the user device and the MES, between the user device and the eavesdropper, and between the MES and the eavesdropper are respectively denoted as where M r and M t respectively represent the receiving antenna and the transmitting antenna of the MES, and M e represents the receiving antenna of the eavesdropper. When the MES sends an interference signal, interference from the output to the input will be generated through the self-interference channel, denoted as The parameter ρ ∈ [0, 1] is introduced to represent the self-interference effect of the server's full-duplex module. The smaller ρ is, the better the full-duplex system. ρ = 0 represents the ideal full-duplex system. This system model considers large-scale fading and small-scale fading, and there is where the small-scale channel fading is generated by independent complex Gaussian distributions with unit variance and H, and the large-scale channel fading is represented by . and respectively represent the path loss constant and the path loss exponent, d UM , d UE , d MErespectively represent the Euclidean distances between the user equipment and the MES, between the user equipment and the eavesdropper, and between the MES and the eavesdropper.
[0095] When the user equipment unloads tasks, it sends a signal χ with a transmission power of p t The signal χ and the transmission power satisfy the constraint. To prevent data eavesdropping, the MES sends an interference signal n J when exchanging data with the user equipment. The interference signal is obtained from the complex Gaussian distribution, denoted as where the covariance matrix Q of the interference signal is a Hermitian matrix and Q ≥ 0 (Q is a positive semi - definite matrix to ensure that the problem has a global minimum). For a given Q, the power of the corresponding interference signal is p J , that is, tr(Q) = p J . Assume that the mobile device has only one antenna, and the edge server adopts the maximum ratio combining (MRC) receiver given by r M = h UM / ||h UM ||. Then the signal received by the edge server is:
[0096]
[0097] where, is the channel noise, is the noise power of the antenna. Assume that the eavesdropper also uses the maximum ratio combining receiver r E = h UE / ||h UE ||. Then the signal received by the eavesdropper is:
[0098]
[0099] where is the corresponding channel noise, is the noise power.
[0100] Since there is eavesdropping during the task offloading process, the secrecy rate is an important indicator to measure the physical layer security performance. Therefore, the secrecy rate is usually used to evaluate the secrecy performance of the wireless communication system. It refers to the maximum transmission rate without information leakage during the offloading process. The secrecy rate can be expressed as:
[0101]
[0102] where, [v] + = max{v, 0}, represents the transmission rate of the task from the user equipment to the edge server, Denotes the task transmission rate between the user equipment and the eavesdropper, usually expressed as:
[0103]
[0104] Where B represents the uplink bandwidth of the user equipment and the MES, and the secrecy rate can be designed by adjusting the power of the AN, H M (Q) and H E (Q) are defined as:
[0105]
[0106] t ∈ [0, τ] represents the time elapsed after making the offloading decision, ρ ∈ [0, 1] represents the self-interference impact, h UM 、h UE 、h ME Represent the channel gains between the user equipment and the MES, the user and the eavesdropper, and the MES and the eavesdropper, respectively.
[0107] Step 13), calculate the task offloading model;
[0108] Step 131), calculate the delay model: The locally generated computing tasks support the partial offloading mode. According to the actual conditions of the network environment and computing resources, select an appropriate offloading ratio. The task ratio λ n ∈ [0, 1] is used to determine how much task data should be offloaded to the MES. Therefore, the generated tasks are divided into two parts: one part is executed locally, and the other part is offloaded to the MES for execution.
[0109] a. Local execution delay: When the mobile device processes the computing task locally, the processing time depends on the computing resources of the mobile device. The local execution delay is expressed as:
[0110]
[0111] n Where c represents the number of CPU cycles required to process one data,
[0112] is the CPU frequency of the local device for processing (local computing power).
[0112] b. Task offloading delay: Placing the task of the terminal in the MES for computing offloading will generate the transmission delay of the task from the user equipment to the MES, the execution delay of the MES processing the task, and the transmission delay of the processing result to the terminal. In actual calculation, the processing result after the task is processed by the MES is much smaller than the input data volume, and the offloading data rate is lower than the upload data rate. Therefore, the feedback transmission time of the downlink is not considered. So, the total delay of task offloading consists of the data transmission delay The execution delay of the MES processing the task Composition. The offloading data transmission delay is expressed as:
[0113]
[0114] From the local computing delay, the execution delay of the MES processing task is:
[0115]
[0116] Among them, is the CPU frequency of the MEC in the t-th time slot.
[0117] Step 132), Energy consumption model: Since the energy of the user equipment is limited, in addition to paying attention to the system delay, the energy consumption problem is also an aspect that the edge computing system needs to focus on. The total energy consumption is divided into two parts: the energy consumption required for local data processing and the transmission energy consumption during the task offloading process.
[0118] a. Local execution energy consumption
[0119] For the partial tasks executed locally, the local energy consumption is related to the local power consumption and the local computing time. According to circuit theory, is the power consumption of the CPU chip, and the coefficient is k, which depends on the CPU architecture. Therefore, the local execution energy consumption can be expressed as:
[0120]
[0121] b. Offloading transmission energy consumption
[0122] For the task offloading part, the data transmission process will consume energy. During the offloading process, the offloading transmission energy consumption can be expressed as:
[0123]
[0124] Among them, p t is the transmission power of the user equipment when offloading tasks.
[0125] Step 14), Establishment of optimization objective
[0126] Since the amount of result data generated by processing tasks in the edge server is usually very small, compared with the delay generated during the uplink transmission process, the transmission delay of the calculation results can be ignored. Therefore, the feedback delay and energy consumption are ignored to simplify the problem. According to equations (1) to (10), the total delay and total energy consumption of the system in t time slots are respectively expressed as:
[0127]
[0128] In the MEC system, the goal is to achieve a balance between reducing task latency and energy consumption under secure wireless communication conditions. Therefore, to simultaneously optimize latency and energy consumption and minimize its (the cost of task offloading and computing), the optimal decision-making model is established as follows:
[0129]
[0130] where β is the weight factor, β ∈ [0, 1], and the parameter β can be adjusted to determine the more important factor between latency and energy cost. λ in C1 n is the task offloading ratio; C2 is the latency constraint, ensuring that the total latency for task completion does not exceed the maximum tolerable latency τ; C3 indicates that the computing resources allocated by the MES for the tasks offloaded to this server do not exceed the computing capacity of this edge server; C4 represents the constraint on the transmission power p when the MES exchanges data with the Internet of Things device AN ; C5 represents the constraint on limiting the transmission power of the user equipment.
[0131] Based on Equation (15), it can be seen that the proposed optimization problem is a mixed-integer non-linear programming problem, a complex multi-objective optimization problem with many constraints in a heterogeneous scenario. Even if the offloading model and problem of mobile edge computing in this industrial Internet of Things can be described, the solution process will be quite difficult. Therefore, for the mobile edge computing task offloading scenario in the IIoT environment, a computing task offloading strategy based on physical layer security is planned. By introducing deep reinforcement learning, the limitations of traditional computing methods in dealing with high-dimensional and non-linear problems are solved. Through intelligent learning and decision-making, the solution to complex offloading problems is achieved, significantly enhancing the flexibility and practicality of problem-solving.
[0132] Step 2: Solving the deep reinforcement learning problem based on multi-objective optimization;
[0133] To achieve efficient decision-making, the DRL method is used to solve the above-mentioned minimization problems of latency and energy consumption, which is modeled as a Markov decision process (MDP), and the optimal policy is learned from the training environment through the DRL method to solve the system's latency and energy consumption minimization problem.
[0134] Step 21): MDP modeling
[0135] In the proposed scenario, the environmental state is perceived to obtain condition information. Considering the random server situation and the dynamic task changes during data transmission, the problem of computing task offloading and resource allocation is modeled as an MDP, and an optimal policy is learned from the training environment through the DRL method to solve the problem of minimizing the weighted sum of latency and energy consumption. This process is defined as a triple M=(S, A, R), where S represents the state space, A represents the action space, and R represents the reward function. In the MDP, the agent continuously interacts with the dynamic environment to optimize its own policy. The agent adjusts its own policy by observing the state s t+1 and the reward r t to maximize the cumulative reward and tries to find an optimal computing offloading policy to minimize the weighted sum of system latency and energy consumption. The basic three elements of the MDP are defined as follows:
[0136] Step 211), State Space State
[0137] To comprehensively consider the characteristics between device tasks and MES resources in the IIoT, the agent needs to observe the system state at the beginning of each time slot. The state includes the task Q n =(d n , c n , t n ) to be executed in each time slot and the task offloading ratio λ n . Therefore, the state space is represented as:
[0138] s t ={d n , c n , t n , λ n} (39)
[0139] where d n represents the amount of computing task data generated, c n represents the number of cycles to process each bit of data on user u n , and t n represents the maximum latency that can be tolerated to complete the task.
[0140] Step 212), Action Space Action
[0141] The agent finds the optimal offloading policy by allocating and offloading the tasks in each time slot. According to the state space, the action a t is output. At time slot t, the action a t ∈A t can be defined as A t ={α, f, Q}, where α={x1(t), x2(t),…, x k (t)}, x n(t) ∈ {0, 1, …, K}, when x n (t) = 0, it means that the task is executed locally for offloading, x n (t) = n means that the task is offloaded to the nth server for execution, and the resource allocation decision represents the computing frequency allocated by the edge server to task Q at time slot t n of.
[0142] Step 213), Reward Reward
[0143] The agent executes actions based on the observed state and obtains rewards from the environment. The goal of the agent is to select actions that can obtain the highest return. The reward function is usually related to the objective function, and the goal is to minimize the weighted sum of latency and energy consumption Cost, that is, to achieve the minimum total cost Cost of the user. Therefore, the reward must be negatively correlated with the value of the total cost. The reward is defined as r(t) = -Cost, where Cost is the weighted sum of latency and energy consumption defined in formula (13).
[0144] In the MEC environment, as the number of devices increases, the state space and action space of the MDP grow exponentially, and it is difficult to solve the optimization problem in polynomial time. Utilizing the ability of deep reinforcement learning to handle high-dimensional state spaces and the powerful learning ability of reinforcement learning, a PPO algorithm with an Actor-Critic framework is adopted to solve the problem of minimizing the weighted sum of latency and energy consumption of the system.
[0145] Step 22), DRL training framework based on PPO
[0146] The PPO algorithm is adopted to achieve secure task transmission and optimize task offloading decisions and resource allocation decisions. The PPO algorithm is used to solve this MDP problem, and appropriate offloading strategies are selected under different states to achieve the optimization of system performance. The choice of the PPO algorithm is based on its excellent performance and efficient learning ability, which is especially suitable for dealing with complex problems with high-dimensional state spaces and action spaces. The DRL training framework based on PPO is as Figure 2 shown.
[0147] The Actor network is an agent used to directly control the offloading scheme strategy. Its output actions determine whether the task is offloaded and the computing resources allocated by the server to each task. The agent observes the task offloading ratio and tasks to be executed at the current time slot as the state, and then outputs the probability distribution of actions. The agent selects its next action by sampling from this probability distribution. The Critic network is used to approximate the value function, defined as:
[0148]
[0149] Among them, γ is the discount factor, represents the expected value, and r(·) represents the reward function for states and actions. The function is used to evaluate the value of a state or a state-action pair, can measure the quality of different policies, estimate the expected return of different policies in a certain state through the value function, and can select a policy that can improve the return, so as to improve the policy to achieve the goal.
[0150] Actor-Critic is a combination of policy optimization and value optimization. The Actor decides what action to take, and the Critic evaluates the action and indicates how the Actor should adjust. During the learning process, PPO stores the interactions between the agent and the environment in memory, and this memory will be used to update the Actor-Critic network. Different from the experience replay of DQN, the PPO algorithm uses all the memory to update the neural network, and the updated neural network is used to collect new experiences and refill the memory. The gradient estimate is given by the following formula:
[0151]
[0152] Among them, π θ is the parameter of the policy neural network, is the estimated value of the advantage function, is defined as:
[0153]
[0154] When performing multiple optimization steps on the policy, a large number of updates are required. To solve this problem and limit the new policy from deviating too far from the old policy, the ratio of probabilities is used to control the update amplitude of the new policy. Let ρ t (θ) represent the probability ratio between the new policy π θ and the old policy :
[0155]
[0156] If the difference between the new and old policies is significant and the advantage function is large, the update amplitude is appropriately increased: If the ρ t ratio is closer to 1, it indicates that the difference between the new and old policies is smaller. The objective function is expressed as:
[0157]
[0158] Among them, ε c is a hyperparameter between 0 and 1, and clip is a truncation parameter. The purpose of this function is to control the update amplitude of the policy and limit the range of the policy update ratio to [1 - ∈ c , 1 + ∈ c .
[0159] To provide more stable estimates of the advantage function, improve the training stability and convergence speed of the algorithm, the Generalized Advantage Estimation (GAE) method is adopted. By adjusting the parameter λ, a balance is found between bias and variance, thereby improving the overall performance of the algorithm. It is:
[0160]
[0161] Where T is a given length, and λ is a hyperparameter. δ(t) is defined as:
[0162] δ(t) = r(t) + γV(s(t + 1)) - V(s(t)) (46)
[0163] Calculate the gradient of the target network according to Equation (21), and update the parameters θ and ξ through gradient descent to complete one round of iteration.
[0164] 3 Simulation Results and Analysis
[0165] 3.1 Parameter Settings
[0166] The system performance, feasibility, and efficiency of the algorithm are analyzed through Python simulation. Assume that there is a passive eavesdropper, 3 small base stations, and 10 user devices in the scenario. Each small base station is equipped with a MES, which can provide task offloading services for mobile users. The coverage radius of the base station is set to 500m. Additionally, for the implementation of the PPO-based solution, two fully connected neural networks are used for the Actor and Critic respectively. Each neural network has two hidden layers, and each hidden layer consists of 64 neurons. The parameter settings are shown in Table 1.
[0167] Table 1 Experimental Parameter Settings
[0168]
[0169] 3.2 Simulation Analysis
[0170] In this experiment, the performance of the proposed task offloading strategy optimization algorithm was compared with that of the Random strategy, the Full Local strategy, the Full MEC strategy, and the Deep Deterministic Policy Gradient (DDPG) algorithm. Among them, the Random strategy randomly generates offloading strategies and computing resource allocation strategies. The Full MEC strategy offloads all tasks to the edge server for processing, and selects edge nodes for offloading and allocates computing resources for the tasks of all edge nodes through the PPO algorithm. The DDPG algorithm is an advanced strategy that combines deep learning and reinforcement learning. This algorithm directly outputs an action vector instead of a probability distribution, requires a large replay buffer to learn the action value function, and is suitable for decision-making problems in continuous action spaces. It is the algorithm that is the focus of comparison.
[0171] Figure 3 This is the convergence of the system reward of the proposed algorithm at different learning rates, where the learning rate is used to adjust the update rate of the PPO network weights. When the learning rate is 0.001, its convergence rate is greater than the reward value when the learning rate is 0.0001, and the reward value is more stable. When the learning rate exceeds 0.01, the training falls into a local optimum and does not obtain the globally optimal reward. Therefore, the learning rate of the PPO network is selected as 0.001.
[0172] Figure 4 This is the influence of the mini-batch sample number (batch size) on the convergence of the system reward. The mini-batch sample number is the number of experience samples required for each training. After 300 training times, as the mini-batch sample number increases from 4 to 32, the system convergence rate is faster and the obtained reward is greater. Therefore, in terms of convergence speed and reward, the larger the mini-batch sample number, the better. When comparing and analyzing other performance indicators, select batch size as 32.
[0173] Figure 5 It can be seen that the selection of a higher maximum tolerable delay will allow the system to have greater flexibility in processing tasks, which usually leads to a higher total delay. A higher maximum delay allows the system to process tasks for a longer time, which may reduce the energy consumption burden at each moment, thereby reducing the overall energy consumption. Figure 5 It can be seen that after 300 training times, when the maximum delay constraint is 2.5 s, a balance point between delay and energy consumption is found, minimizing the average cost of the system.
[0174] In order to further understand the relationship between task delay and energy consumption, an experiment of adjusting the weight parameter (β) was carried out to illustrate their relationship, as Figure 6As shown in , the weighted sum of delay and energy consumption decreases as β increases, which numerically reflects that delay is greater than energy consumption. Figure 7 It can be seen that the final system average cost will eventually tend to be flat.
[0175] Figure 7 The convergence performance of the system rewards under the five algorithm strategies. It can be seen that with the increase of training rounds, the three algorithms of DDPG, Full MEC and PPO gradually reach convergence, the Random strategy has been fluctuating within a range, and the FullLocal strategy is fixed in each round, and the reward value remains unchanged. After convergence and stabilization, the PPO algorithm is better than the other four strategies in terms of convergence speed and reward.
[0176] In order to minimize the system's latency and energy consumption, the task offloading ratio, artificial noise power, user and MEC computing resource allocation are jointly optimized, and a mobile edge computing task offloading strategy considering secure transmission is proposed. Simulation results show that the trained agent can generate resource allocation and strategy optimization with low complexity. At the same time, the proposed scheme can significantly reduce the system's latency and energy consumption.
[0177] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention should be included in the scope of the present invention.
Claims
1. A mobile edge computing offloading strategy and resource optimization method based on physical layer security, characterized in that The specific steps are as follows: Step 1: Model the secure offloading of mobile edge computing tasks; Step 11): Build a task offloading model. The user equipment and the edge server are both randomly distributed in the same area. The user selects a single edge server for access and computing offloading. The user equipment offloads some or all of the tasks to the edge server for collaborative computing. The edge server emits interference signals to actively prevent eavesdropping, and the user achieves secure offloading; Step 12): Build a secure transmission channel model to obtain the channel gain between the user equipment and the MES, the channel gain between the user and the eavesdropper, and the channel gain between the MES and the eavesdropper; Step 13): Calculate the task offloading model; Step 131): Calculate the delay model; Select the corresponding offloading ratio to achieve partial offloading of the locally generated computing tasks. The generated tasks are divided into the part executed locally and the part offloaded to the MES for execution; Step 132): Build an energy consumption model. The total energy consumption of the energy consumption model consists of the energy consumption required for local data processing and the transmission energy consumption during task offloading; Step 14): Establish an optimization objective. For the mobile edge computing task offloading scenario, design a computing task offloading strategy based on physical layer security, introduce deep reinforcement learning, and through intelligent learning and decision-making, solve the offloading problem; Step 2: Solve the deep reinforcement learning problem based on multi-objective optimization; Step 21): MDP modeling; Sense the environmental state to obtain conditional information, model the offloading and resource allocation problems of computing tasks as MDP, and learn the optimal policy from the training environment through the DRL method; Step 22): Build a DRL training framework based on PPO. Use the PPO algorithm to achieve secure task transmission and optimize the task offloading decision and resource allocation decision. Use the PPO algorithm to solve this MDP problem, and select the corresponding offloading strategy under different states to achieve the optimization of system performance.
2. The mobile edge computing offloading strategy and resource optimization method based on physical layer security according to claim 1, characterized in that In step 11), the specific steps are as follows: Construct a task offloading model. There are N user devices, K small base stations, and L eavesdroppers in the system. Each small base station is equipped with an edge server. The set of user devices is denoted as U = {1, 2, 3,..., N}, the set of edge servers is denoted as M = {1, 2, 3,…, K}, and the set of eavesdroppers is denoted as E = {1, 2,…, l,…, L}. Assume that user device u n generates one or more computation-intensive and latency-sensitive tasks Q n =(d n , c n , t n ) in each time slot t ∈ T, where d n represents the number of generated computation tasks, c n represents the number of cycles to process each bit of data on user u n , t n represents the maximum tolerable latency to complete the task. Determine an offloading ratio λ n , and offload the task to the edge server for computation. (1 - λ n )d n of the data is computed locally, and the remaining λ n d n of the data is offloaded to the MES for computation.
3. The mobile edge computing offloading strategy and resource optimization method based on physical layer security according to claim 2, wherein In step 12), the channel gain between the user equipment and the MES is denoted as The channel gain between the user equipment and the eavesdropper is denoted as The channel gain between the MES and the eavesdropper is denoted as where M r represents the receiving antenna of the MES, and M t represents the transmitting antenna of the MES, and M e represents the receiving antenna of the eavesdropper. When the MES sends an interference signal, self-interference from the output to the input will be generated through the self-interference channel, denoted as The parameter ρ ∈ [0, 1] is introduced to represent the self-interference effect of the server full-duplex module. This system model has where the small-scale channel fading is generated by independent complex Gaussian distributions with unit variance and captured by H, and the large-scale channel fading is represented by denoted as represents the path loss constant represents the path loss exponent, and d UM represents the Euclidean distance between the user equipment and the MES, and d UE represents the Euclidean distance between the user equipment and the eavesdropper, and d ME represents the Euclidean distance between the MES and the eavesdropper; The user equipment sends a signal χ with a transmission power of p when unloading a task t The signal χ and the transmission power satisfy the constraint. The MES sends an interference signal n when exchanging data with the user equipment J The interference signal is obtained from a complex Gaussian distribution, denoted as where the covariance matrix Q of the interference signal is a Hermitian matrix and Q ≥ 0. For a given Q, the power of the corresponding interference signal is p J , that is, tr(Q) = p J Assume that the mobile device has only one antenna, and the edge server uses the maximum ratio combining receiver given by r M = h UM / ||h UM ||. Then the signal received by the edge server is as follows: wherein, is the channel noise, is the noise power of the antenna. Assume that the eavesdropper also uses a maximum ratio combining receiver r E = h UE / ||h UE ||, then the signal received by the eavesdropper is: wherein, is the corresponding channel noise, is the noise power; Secrecy rate It is expressed as: where, [v] + = max{v, 0}, represents the transmission rate of the task unloaded from the user equipment to the edge server, represents the task transmission rate between the user equipment and the eavesdropper, specifically expressed as: Among them, B represents the uplink bandwidth of the user equipment and the MES, and the secrecy rate is designed by adjusting the power of the AN, H M (Q) and H E (Q) are defined as: $t\in[0,\tau]$ represents the time elapsed after making the offloading decision, $\rho\in[0,1]$ represents the self-interference impact, $h$ UM represents the channel gain between the user equipment and the MES, $h$ UE represents the channel gain between the user and the eavesdropper, $h$ ME represents the channel gain between the MES and the eavesdropper.
4. The mobile edge computing offloading strategy and resource optimization method based on physical layer security according to claim 3, characterized in that In step 131), calculate the latency model, and the task proportion λ n ∈[0,1] determines how much task data is offloaded to the MES. The generated tasks are divided into two parts: one part is executed locally, and the other part is offloaded to the MES for execution; a. Local execution delay: When the mobile device processes computing tasks locally, the processing time depends on the computing resources of the mobile device. The local execution delay is expressed as: Among them, c n represents the number of CPU cycles required to process a piece of data, which is the CPU frequency processed by the local device; b. Task offloading delay: When the tasks of the terminal are offloaded to the MES for calculation, there are transmission delays for the tasks to be transmitted from the user equipment to the MES, execution delays for the MES to process the tasks, and transmission delays for the processing results to be transmitted to the terminal. Without considering the feedback transmission time of the downlink, the total delay of task offloading consists of data transmission delays Execution delay of the MES to process tasks and is expressed as: From the local computing delay, the execution delay of the MES processing tasks is: Among them, is the CPU frequency of the MEC in the t-th time slot; In step 132), is the power consumption of the CPU chip, with a coefficient of k, which depends on the CPU architecture. The energy consumption of local execution is expressed as: During the offloading process, the offloading transmission energy consumption is expressed as: Among them, p t is the transmission power of the user equipment when unloading tasks.
5. The mobile edge computing offloading strategy and resource optimization method based on physical layer security according to claim 4, characterized in that, In step 14), according to equations (1) to (10), the total delay of the system in t time slots is expressed as: The total energy consumption of the system in t time slots is expressed as: Optimize the delay and energy consumption simultaneously to minimize the cost of task offloading calculation, and establish the optimal decision model as: Among them, β is the weight factor, and λ in C1 n is the task offloading ratio, C2 is the delay constraint to ensure that the total delay for task completion does not exceed the maximum tolerable delay τ, C3 indicates that the computing resources allocated by the MES for the tasks offloaded to this server do not exceed the computing power of this edge server, and C4 indicates the constraint on the transmission power p AN when the MES exchanges data with the Internet of Things device, and C5 indicates the limitation of the transmission power of the user equipment.
6. The method for mobile edge computing offloading strategy and resource optimization based on physical layer security according to claim 5, characterized in that, In step 21), the MDP modeling process is defined as a triple M = (S, A, R), where S represents the state space, A represents the action space, and R represents the reward function. In MDP, the agent continuously interacts with the dynamic environment to optimize its own policy. The agent adjusts its own policy by observing the state s t+1 and the reward r t to maximize the cumulative reward. The basic three elements of MDP are defined as follows: Element 1. State space State. To comprehensively consider the characteristics between device tasks and MES resources in IIoT, at the beginning of each time slot, the agent observes the system state, and the system state includes the tasks Q to be executed in each time slot n =(d n ,c n ,t n ) and the task offloading ratio λ n , and represent the state space as: s t = {d n , c n , t n , λ n}} (16) Among them, d n represents the amount of computational task data generated, and c n represents the number of cycles per bit of data processed on user u n , and t n represents the maximum latency that can be tolerated to complete the task; Element 2, Action Space Action. The agent finds the optimal offloading strategy by allocating and offloading tasks for each time slot and outputs the action a according to the state space. t , at time slot t, the action a t ∈A t is defined as A t ={α, f, Q}, where α={x1(t), x2(t),…, x k (t)}, x n (t)∈{0, 1, …, K}. When x n (t)=0, it means that the task is offloaded and executed locally. When x n (t)=n, it means that the task is offloaded to the nth server for execution. The resource allocation decision represents the computing frequency allocated by the edge server to task Q n at time slot t. Element 3. Reward Reward, define the reward as r(t) = -Cost, where Cost is the weighted sum of the delay and energy consumption defined in formula (13).
7. The method for mobile edge computing offloading strategy and resource optimization based on physical layer security according to claim 6, characterized in that, In step 22), the agent observes the task offloading ratio and the tasks to be executed in the current time slot as the state, outputs the probability distribution of the action, and the agent selects its next action by sampling from this probability distribution. The Critic network is used to approximate the value function, defined as: where γ is the discount factor, denotes the expected value, and r(·) denotes the reward function with respect to the state and action; The updated neural network is used to collect new experiences and refill the memory. The gradient estimate is given by the following formula: where, π θ are the parameters of the policy neural network, is the estimated value of the advantage function, is defined as: Use the ratio of probabilities to control the update amplitude of the new policy. Let ρ t (θ) represent the probability ratio between the new policy π θ and the old policy : If ρ t The closer the ratio is to 1, the smaller the difference between the new and old strategies. The objective function is expressed as: Among them, ε c is a hyperparameter between 0 and 1, and clip is a truncation parameter that restricts the range of the policy update ratio to [1 - ∈ c , 1 + ∈ c ; The generalized advantage estimation method is adopted to adjust the parameter λ to find a balance between bias and variance. It is: where T is a given length, and λ is a hyperparameter, and δ(t) is defined as: δ(t) = r(t) + γV(s(t + 1)) - V(s(t)) (23) Calculate the gradient of the target network according to Equation (21), and update the parameters θ and ξ through gradient descent to complete one round of iteration.
Citation Information
Cited By
Industrial task unloading optimization method integrating edge computing and time-sensitive network
CN120892105A