Edge computing-oriented interpretable task unloading and position privacy protection method

By combining Monte Carlo tree search and deep reinforcement learning methods, Markov decision-making process is constructed and three-dimensional differential privacy technology is introduced, which solves the balance between task offload optimization and privacy protection in mobile edge computing, and achieves the protection of user location privacy and system performance improvement.

CN120343636APending Publication Date: 2025-07-18SHANXI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463783.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing task offloading methods are difficult to balance task offload optimization and privacy protection in mobile edge computing environments, and user location privacy faces the risk of leakage during data transmission.

Method used

Using a method of combining Monte Carlo tree search (MCTS) and deep reinforcement learning (DRL), we use the Markov decision-making process, design task offloading strategies, combine three-dimensional differential privacy technology to protect user location privacy, realize adaptive intelligent decision-making and location disturbance, and balance system performance and privacy protection.

Benefits of technology

In the dynamic multi-user and multi-edge node environment, the optimality of task offloading and the balance between privacy protection is achieved, theoretical interpretability and three-dimensional differential privacy guarantees are provided, and the difficulty of attackers inferring user locations is reduced, and system performance and privacy security are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343636A_ABST
    Figure CN120343636A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable task unloading and location privacy protection method (3D-LPPOS) for edge computing, and belongs to the technical field of mobile edge computing and task unloading. Aiming at the problem that an existing task unloading method is difficult to balance between task unloading optimality and privacy protection, the task unloading optimization problem of huge unloading decision space in a dynamic complex task unloading environment is solved by using a mode of combining Monte Carlo Tree Search (MCTS) and deep reinforcement learning (DRL). By introducing an optimal privacy parameter for balancing system performance and privacy protection requirements in task unloading, the problem of privacy leakage of user position information jointly inferred by unsafe edge nodes is solved by disturbing the real position of a mobile user based on the parameter; the task optimal unloading decision with theoretical interpretability in a dynamic multi-user multi-edge node scene is realized, three-dimensional differential privacy guarantee is provided for location privacy of mobile users, and system performance and privacy protection requirements are successfully balanced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mobile edge computing and task offloading, and particularly relates to an interpretable task offloading and location privacy protection method for edge computing. Background Art

[0002] As an innovative computing architecture, Mobile Edge Computing (MEC) aims to shift computing resources and data storage from traditional clouds to the network edge, closer to end-users. This technology effectively addresses the limitations of mobile devices in terms of computing resources and battery life, as well as the bottlenecks of cloud computing in terms of latency and security, greatly enhancing the response speed of real-time applications (such as face recognition, augmented reality, and virtual reality). By reducing data transmission latency and optimizing resource allocation, MEC not only enhances the user experience but also plays a crucial role in fields such as smart cities and industrial Internet, driving the revolutionary development of the next-generation network architecture.

[0003] The dynamic, multi-user, and multi-edge-node nature makes it difficult to find the optimal task offloading strategy in the mobile edge computing environment, and many offloading methods are prone to falling into local optimal solutions in such a complex environment. At the same time, there are potential security issues during the task offloading process. MEC systems usually rely on wireless networks for data transmission and computing task offloading, and user privacy faces the risk of leakage during the process of data transmission and computing task processing. For example, attackers can infer the real-time location of users by analyzing the offloading pattern, and then reveal their identities or daily activity patterns, and use this information to launch personalized attacks, thus threatening personal privacy and security. Optimizing the task offloading strategy and protecting user privacy are key objectives for improving the performance of MEC networks. In recent years, balancing offloading performance and security protection in mobile edge computing has become a research hotspot in the field. Summary of the Invention

[0004] Aiming at the problem that it is difficult to balance the optimality of task offloading and privacy protection in existing task offloading methods, to solve this problem, the following two important aspects are considered: one is how to design an efficient and accurate task offloading strategy, and the other is how to effectively protect user privacy while improving system performance, balancing the optimality of task offloading and user privacy protection; therefore, the present invention provides an interpretable task offloading and location privacy protection method for edge computing.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] Combining the mobile edge node task offloading problem with Monte Carlo Tree Search (MCTS), and realizing adaptive intelligent decision-making through Deep Reinforcement Learning (DRL), a three-dimensional differential privacy technology is proposed to interfere with the real location of mobile users to protect location privacy. The task offloading scenario is established as a complex environment with dynamic multi-users, multi-edge nodes and supporting partial user offloading. An offloading strategy consisting of two stages is formulated, and the problem is transformed into a Markov decision process for solution. The purpose is to minimize the task offloading cost, task discard rate and edge node load variability in mobile edge computing, while maximizing the privacy entropy that measures the privacy protection ability, in a mobile edge computing scenario facing dynamic changes in task characteristics, server computing capabilities, and the presence of multiple users and multiple edge nodes in the environment.

[0007] An interpretable task offloading and location privacy protection method for edge computing, the method comprising the following steps:

[0008] Step 1: Construct a dynamic simulation environment to simulate the real task offloading scenario in a mobile edge system, including complex characteristics of dynamic multi-users, multi-edge nodes and supporting partial user offloading. Model the task offloading of the mobile edge system, comprehensively consider the balance between task offloading energy consumption and location privacy, and transform it into a Markov decision process; In the present invention, the mobile edge node task offloading problem is modeled. Within each time slot, the computing capabilities of mobile users and edge nodes, the amount of tasks to be executed, and the computing density are changing in real time. Each mobile user can choose to execute tasks locally or choose to offload tasks to edge nodes. Edge nodes provide offloading services for multiple mobile users. The tasks generated by mobile users are detachable fine-grained, allowing users to partially offload.

[0009] Further, the specific operation of the step 1 is:

[0010] Establish the task offloading scenario as a complex environment with dynamic multi-users, multi-edge nodes and supporting partial user offloading. Formulate an offloading strategy consisting of two stages, and use the execution cost, task discard rate, edge node load variability and privacy entropy as the measurement metrics for task offloading;

[0011] Model the task offloading problem in mobile edge computing. The following is the formulaic description of the problem:

[0012]

[0013] 0≤λ≤1 (3)

[0014] Among them, P0 is the objective of the present invention, which means minimizing the cost of task offloading, task discard rate, and edge node load variability in mobile edge computing, while maximizing the privacy entropy that measures the privacy protection ability; T represents the number of time slots included; N represents the number of mobile users; M represents the number of edge nodes; C(t) represents the execution cost at time slot t; v m (t) represents the load variability of the edge node; qr n (t) represents the user UE n The task discard rate at time slot t, H n (t) represents the privacy entropy, which is used to quantify the privacy degree of the user UE n Offloading strategy; π n,0 (t) represents the proportion executed locally at the user UE n Proportion;

[0015] Constraint (1) means that the proportion of any user executing tasks locally and on all edge nodes is within the range of 0 - 1;

[0016] Constraint (2) means that the sum of the proportions of any user executing tasks locally and on all edge nodes is 1;

[0017] Constraint (3) means that the weight for balancing energy consumption and computing latency is within the range of 0 - 1;

[0018]

[0019] Among them, And Are the total latency and total energy consumption at time slot t respectively, and λ ∈ [0, 1] is the weight for balancing energy consumption and computing latency;

[0020] Represents the total latency at time slot t, which is the maximum value of the local computing latency And the mobile edge computing latency Of;

[0021] The local computing latency is expressed as the maximum latency of all mobile users computing locally;

[0022] The user UE n The latency formula for computing locally is:

[0023] Among them, l n (t) = π n,0 (t)d n (t) represents the amount of task size executed locally by the user UE n Proportion; d n (t) represents the task size, π n,0 (t) represents the proportion at the user UEn The proportion executed locally, b n (t) represents the number of CPUs required for the task of calculating each bit, represents the user UE n The local computing power;

[0024] The mobile edge computing latency is expressed as the maximum latency of all edge nodes;

[0025] The edge node MEC m Calculating the user UE n The latency of the offloading task consists of two parts, namely the transmission latency and the computing latency

[0026] The formula for the transmission latency is:

[0027] The formula for the computing latency is:

[0028] where, l n,m (t) = π n,m (t)d n (t) represents the user UE n Offloaded to the MEC m The amount of the task executed, π n,m (t) represents the proportion of the task offloaded to the MEC m Executed, r n,m (t) represents the user UE n To the MEC m The data transmission rate, Represents the MEC m The computing power.

[0029]

[0030] B represents the total available bandwidth of the user, P n,m (t) represents the user UE n To the MEC m The transmission power, N0 represents the channel noise, g n,m (t) = dist n,m (t) -α (t) represents the user UE n And the MEC m The channel power gain between them, dist n,m (t) represents the user UE n To the MEC m The distance, α represents the path loss exponent.

[0031] Denotes the total energy consumption at time slot t, which is obtained by adding the local computing energy consumption and the mobile edge computing energy consumption ;

[0032] Denoted as the sum of the local computing energy consumption of all mobile users;

[0033] Denoted as the sum of the transmission energy consumption of all edge nodes, and the calculation formula is: Considering that the edge nodes have a stable power supply, the computing energy consumption is ignored in this invention.

[0034] k represents the energy consumed per CPU cycle.

[0035] The load variability v m (t) of the edge nodes is measured by the standard deviation of the server load. The larger the standard deviation, the greater the load fluctuation and the higher the variability; conversely, the smaller the fluctuation, the lower the variability. A smaller load variability indicates that all edge nodes make full use of computing resources and reduce resource waste. The formula is:

[0036]

[0037] Where, represents the load of the edge node MEC m , represents the average load of all edge nodes.

[0038] The mobile edge node task offloading problem is transformed into a Markov decision process (MDP), which specifically includes:

[0039] State space: S(t) = [W(t), F L (t), F M (t), Π(t)], where W(t) is the task feature set of the mobile user, f L (t) is the computing ability set of the mobile user, F M (t) is the computing ability set of the edge node, and Π(t) is the offloading vector set of the mobile user;

[0040] Action space: A(t) = [a1(t), a2(t), …, a N (t)] T , where a n (t) = [β1, β2, …, β M , representing the decision vector of the user UE n on all edge nodes;

[0041] State transition probability: The state transition probability of the task offloading environment is set to 1;

[0042] Reward function: Set the reward function to where qr n (t) represents the task discard rate of user UE n at time slot t, and H n (t) represents the privacy entropy, which is used to quantify the privacy level of the offloading strategy of user UE n ;

[0043] Based on the task offloading preference, the present invention uses the privacy entropy H n (t) to quantify the privacy level of the offloading strategy of user UE n . The privacy entropy measures the preference uncertainty for different edge nodes in the offloading decision. The higher the entropy value, the more difficult it is to interpret the decision information, and the more difficult it is for the attacker to infer the user's location. When the user fully executes the task locally without offloading, the edge node cannot obtain the user information, and at this time, the privacy entropy is the largest. The calculation formula of the privacy entropy is:

[0044]

[0045] where represents the offloading preference of user UE n for MEC m , and is inferred according to the total offloading volume of user UE n and the partial offloading volume for MEC m ;

[0046] The calculation formula of the task discard rate of user UE n at time slot t is:

[0047]

[0048] where represents the amount of tasks discarded by user UE n at time slot t; τ represents the length of a time slot.

[0049] Step 2: For the mobile edge node task offloading problem, design an algorithm that combines Monte Carlo tree search and deep reinforcement learning. Select actions through an equilibrium strategy, obtain the most promising actions while avoiding falling into local optimal solutions; through the underlying feedback mechanism, the upper-layer nodes can provide accurate evaluations based on statistical characteristics, thereby effectively reducing the search space. Use this algorithm for dynamic decision-making of task offloading to achieve efficient allocation and execution, and has theoretical interpretability;

[0050] Furthermore, the specific operation of the said Step 2 is:

[0051] Set the initial offloading state as the root node. Starting from the root node, establish a task offloading search tree through the four steps of selection, expansion, simulation, and backpropagation of Monte Carlo tree search. The offloading strategy selected in each time slot corresponds to the action corresponding to the child node with the largest UCB value in each layer traversed from the root node to the leaf node of the search tree. At the same time, optimize the mapping relationship between the state and the action through the reward feedback in the environment interaction, gradually learn and approximate the optimal offloading strategy, and achieve adaptive intelligent decision-making.

[0052]

[0053] v′ is the child node of v. Ι and V represent the number of visits and the value respectively. c represents the exploration coefficient. By adjusting, the balance between exploitation (the first term of UCB) and exploration (the second term of UCB) of the algorithm is changed. When it is decreased, the algorithm is more inclined to exploit the known information; when it is increased, the algorithm is more inclined to explore.

[0054] Specifically, it includes:

[0055] Step 2.1: Model the task offloading problem and impose necessary constraints on the mobile edge system.

[0056] Step 2.2: Create a task offloading search tree

[0057] Step 2.2.1: Set the initial offloading state S0 as the root node Λ0, initialize the policy network Q with random weights θ, initialize the target network with weights Initialize the target network And initialize the experience replay pool à. Under a limited computing budget, loop through the following steps 2.2.2 - 2.2.4;

[0058] Step 2.2.2: Select an unexpanded node starting from Λ0 and expand it into Λ′;

[0059] Step 2.2.2.1: When Λ0 is in a non - terminal state and has been fully expanded, use the above UCB(v,c) formula to select the child node with the largest UCB value in Λ0 as the new Λ0;

[0060] Step 2.2.2.2: When Λ0 is in a non - terminal state and has not been fully expanded, then Λ0 expands to a child node Λ′. Select an unvisited action a from the possible action space of Λ0 and use it as the action of Λ′, and update the depth of Λ′;

[0061] Step 2.2.2.3: Use the state transition function to find the state S(Λ′) and the reward value r of Λ′, store them together with the action a in the experience replay pool à, and update the network parameters regularly;

[0062] Step 2.2.3: Starting from the extended node Λ′, simulate until the termination state;

[0063] Step 2.2.3.1: Λ′ adopts the action a with the maximum value in the policy network * = argmaxQ(S(Λ′), Α(Λ′); θ), and uses the state transition function to obtain the next state S′(Λ′) and the reward value r, and updates the current state s(Λ′) = s′(Λ′);

[0064] Step 2.2.3.2: Store the sequence [s(Λ′), a * , r, s′(Λ′)] into the experience replay pool à, and update the network parameters regularly;

[0065] Step 2.2.4: Update the values and visit counts of all nodes on the Λ′ path in reverse;

[0066] Step 2.3: Let node be the root node of the task offloading search tree When node does not meet the termination state condition, obtain the optimal offloading strategy for each time slot;

[0067] Step 2.3.1: Update node to the child node corresponding to the maximum UCB(node, 0);

[0068] Step 2.3.2: Obtain the time slot t = deep(node), and the optimal offloading strategy A * (t) = a(node);

[0069] Step 2.4: Calculate the task according to the obtained optimal offloading strategy A * (t) for each time slot.

[0070] Step 3: Introduce the optimal privacy parameter that balances system performance and privacy protection requirements in task offloading. By dynamically adjusting the privacy parameter, it not only meets the efficiency of task offloading but also ensures the security of user location privacy, achieving a balance between task offloading performance and privacy protection requirements. Based on this parameter, the true location of the user is perturbed to achieve three-dimensional geographical indistinguishability, ensuring that even if an attacker can obtain some information, they cannot accurately infer the exact location of the user, providing three-dimensional differential privacy guarantee for the location privacy of mobile users;

[0071] Furthermore, the specific operation of Step 3 is as follows:

[0072] Step 3.1: Conduct z interference experiments for different privacy coefficients ε, generate the perturbation radius R = Γ(3, 1 / ε), randomly select the unit vector U = (ξ, ψ), and the perturbed location of the user UE n The perturbed location

[0073] Step 3.2: Record the average reward of the perturbation position generated by ε in Table Ο;

[0074] Step 3.3: Obtain the optimal privacy coefficient ε according to the maximum average reward in Table Ο * ;

[0075] Step 3.4: Use ε * to calculate the perturbed position of the mobile user;

[0076] Compared with the prior art, the present invention has the following advantages:

[0077] 1) The present invention proposes a task offloading method. Through the feedback mechanism of Monte Carlo Tree Search (MCTS) and the adaptive learning ability of Deep Reinforcement Learning (DRL), while obtaining the most promising actions, it avoids falling into local optimal solutions. The upper-level nodes provide accurate evaluations based on statistical characteristics, thereby effectively reducing the search space. It realizes the optimal task offloading decision in a dynamic multi-user multi-edge node environment and has theoretical interpretability.

[0078] 2) The present invention proposes a location privacy protection mechanism that satisfies three-dimensional differential privacy guarantees. It introduces three-dimensional location privacy protection into task offloading decisions, perturbs the user's true location with optimal privacy parameters, realizes three-dimensional geographical indistinguishability, and provides three-dimensional differential privacy guarantees for the location privacy of mobile users.

[0079] 3) The present invention effectively realizes the balance between the performance and privacy protection requirements of task offloading in a mobile edge computing system. Specifically, it includes two aspects: when modeling the task offloading problem as a Markov decision process, a reward function that comprehensively considers the balance between task offloading energy consumption and location privacy is set; an optimal privacy parameter is innovatively introduced, and by optimizing the configuration of the privacy parameter, both the efficiency of task offloading and the security of user location privacy are ensured. Brief Description of the Drawings

[0080] Figure 1 is the overall framework diagram of an interpretable task offloading and location privacy protection method for edge computing provided by an embodiment of the present invention;

[0081] Figure 2 is the two-stage diagram of task calculation of an interpretable task offloading and location privacy protection method for edge computing provided by an embodiment of the present invention; Detailed Embodiments

[0082] To deeply understand the present invention, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.

[0083] An interpretable task offloading and location privacy protection method for edge computing, the method comprising the following steps:

[0084] Step 1: Construct a dynamic simulation environment to simulate the real task offloading scenario in a mobile edge system, which includes complex features of dynamic multi-users, multi-edge nodes and supports partial offloading of users. Model the task offloading of the mobile edge system, comprehensively consider the balance between task offloading energy consumption and location privacy, and transform it into a Markov decision process; In this invention, the task offloading problem of mobile edge nodes is modeled. Within each time slot, the computing capabilities, the amount of tasks to be executed, and the computing density of mobile users and edge nodes change in real time. Each mobile user can choose to execute tasks locally or choose to offload tasks to edge nodes. Edge nodes provide offloading services for multiple mobile users. The tasks generated by mobile users are detachable fine-grained particles, allowing users to partially offload;

[0085] Further, the specific operation of step 1 is as follows:

[0086] Establish the task offloading scenario as a complex environment of dynamic multi-users, multi-edge nodes and supports partial offloading of users, formulate an offloading strategy including two stages, and use the execution cost, task discard rate, edge node load variability and privacy entropy as metrics for task offloading;

[0087] Model the task offloading problem in mobile edge computing. The following is the formulation description of the problem:

[0088]

[0089] 0 ≤ λ ≤ 1 (3)

[0090] Among them, P0 is the goal of this invention, which means to minimize the cost, task discard rate and edge node load variability of task offloading in mobile edge computing, while maximizing the privacy entropy that measures the privacy protection ability; T represents the number of time slots included; N represents the number of mobile users; M represents the number of edge nodes; C(t) represents the execution cost in time slot t; v m (t) represents the load variability of the edge node; qr n (t) represents user UE n The task discard rate in time slot t, H n (t) represents the privacy entropy, which is used to quantify the privacy degree of the offloading strategy of user UE n π n,0 (t) represents the proportion executed locally by user UE n The proportion executed locally;

[0091] Constraint (1) means that the proportion of tasks executed by any user locally and on all edge nodes is within the range of 0-1;

[0092] Constraint (2) indicates that the sum of the proportions of any user executing tasks on the local and all edge nodes is 1;

[0093] Constraint (3) indicates that the weight for balancing energy consumption and computing latency is within the range of 0 - 1;

[0094] represents the execution cost at time slot t;

[0095] where, and are the total latency and total energy consumption at time slot t respectively, and λ ∈ [0, 1] is the weight for balancing energy consumption and computing latency;

[0096] represents the total latency at time slot t, which is the maximum of the local computing latency and the mobile edge computing latency ;

[0097] The local computing latency is expressed as the maximum latency of all mobile users computing locally;

[0098] User UE n The latency formula for computing locally is:

[0099] where, l n (t) = π n,0 (t)d n (t) represents the amount of task executed locally by user UE n d n (t) represents the amount of task, and π n,0 (t) represents the proportion of execution on the local of user UE n b n (t) represents the number of CPUs required to compute each bit of the task, represents the computing power of user UE n locally;

[0100] The mobile edge computing latency is expressed as the maximum latency of all edge nodes;

[0101] Edge node MEC m The latency for computing the offloaded task of user UE n consists of two parts, namely the transmission latency and the computing latency

[0102] The formula for the transmission latency is:

[0103] The formula for the computing latency is:

[0104] Among them, l n,m (t) = π n,m (t)d n (t) represents the amount of tasks executed by the user UE n unloaded to the MEC m executed, π n,m (t) represents the proportion of tasks unloaded to the MEC m executed, r n,m (t) represents the user UE n to the MEC m data transmission rate, represents the MEC m computing power.

[0105]

[0106] B represents the total available bandwidth of the user, P n,m (t) represents the user UE n to the MEC m transmission power, N0 represents the channel noise, g n,m (t) = dist n,m (t) -α represents the user UE n and the MEC m channel power gain between, dist n,m (t) represents the user UE n to the MEC m distance, α represents the path loss exponent.

[0107] represents the total energy consumption in time slot t, which is obtained by adding the local computing energy consumption and the mobile edge computing energy consumption added together;

[0108] represents the sum of the local computing energy consumption of all mobile users;

[0109] represents the sum of the transmission energy consumption of all edge nodes, and the calculation formula is: Considering that the edge nodes have a stable power supply, the present invention ignores the computing energy consumption.

[0110] k represents the energy consumed per CPU cycle.

[0111] The load variability v of the edge nodes m(t) is measured by the standard deviation of the server load. The larger the standard deviation, the greater the load fluctuation and the higher the variability. Conversely, the smaller the fluctuation, the lower the variability. A smaller load variability indicates that all edge nodes are fully utilizing computing resources, reducing resource waste. The formula is:

[0112]

[0113] where represents the load of the edge node MEC m and represents the average load of all edge nodes.

[0114] The mobile edge node task offloading problem is transformed into a Markov decision process (MDP), which specifically includes:

[0115] State space: S(t) = [W(t), F L (t), F M (t), Π(t)], where W(t) is the task feature set of the mobile user, f L (t) is the computing ability set of the mobile user, F M (t) is the computing ability set of the edge node, and Π(t) is the offloading vector set of the mobile user;

[0116] Action space: A(t) = [a1(t), a2(t), …, a N (t)] T , where a n (t) = [β1, β2, …, β M , representing the decision vector of the user UE n on all edge nodes;

[0117] State transition probability: The state transition probability of the task offloading environment is set to 1;

[0118] Reward function: The reward function is set to where qr n (t) represents the task discard rate of the user UE n at time slot t, and H n (t) represents the privacy entropy, which quantifies the privacy degree of the offloading strategy of the user UE n ;

[0119] Based on the task offloading preference, the present invention uses the privacy entropy H n (t) to quantify the user UE nPrivacy level of the offloading policy. Privacy entropy measures the uncertainty of preferences for different edge nodes in the offloading decision. The higher the entropy value, the more difficult it is to interpret the decision information, increasing the difficulty for attackers to infer the user's location. When the user executes tasks locally completely without offloading, the edge node cannot obtain user information, and at this time, the privacy entropy is the largest. The calculation formula for privacy entropy is:

[0120]

[0121] where represents the offloading preference of user UE n for MEC m and is inferred based on the total offloading volume of user UE n and the partial offloading volume for MEC m ;

[0122] The calculation formula for the task discard rate of user UE n at time slot t is:

[0123]

[0124] where represents the amount of tasks discarded by user UE n at time slot t; τ represents the length of one time slot.

[0125] Step 2: For the mobile edge node task offloading problem, design an algorithm that combines Monte Carlo tree search and deep reinforcement learning. Select actions through an equilibrium strategy, obtain the most promising actions while avoiding falling into local optimal solutions; through the underlying feedback mechanism, the upper-layer nodes can provide accurate evaluations based on statistical characteristics, thereby effectively reducing the search space. Use this algorithm to dynamically make decisions on task offloading, achieve efficient allocation and execution, and have theoretical interpretability;

[0126] Furthermore, the specific operation of Step 2 is as follows:

[0127] Set the initial offloading state as the root node. Starting from the root node, establish a task offloading search tree through the four steps of selection, expansion, simulation, and backpropagation of Monte Carlo tree search The offloading strategy selected in each time slot corresponds to the action corresponding to the child node with the largest UCB value for each layer traversed from the root node to the leaf node of the search tree. At the same time, optimize the mapping relationship between the state and the action through the reward feedback in the environment interaction, gradually learn and approximate the optimal offloading strategy, and achieve adaptive intelligent decision-making.

[0128]

[0129] v' is a child node of v. Ι and V represent the number of visits and value respectively, and c represents the exploration coefficient. By adjusting, the balance between exploitation (the first term of UCB) and exploration (the second term of UCB) of the algorithm is changed. When it decreases, the algorithm tends to exploit the known information more; when it increases, the algorithm tends to explore more.

[0130] Specifically, it includes:

[0131] Step 2.1: Model the task offloading problem and impose necessary constraints on the mobile edge system.

[0132] Step 2.2: Create a task offloading search tree

[0133] Step 2.2.1: Set the initial offloading state S0 as the root node Λ0, initialize the policy network Q with a random weight θ, initialize the target network with the weight Initialize the target network and initialize the experience replay pool à. Under a limited computing budget, loop through the following steps 2.2.2 - 2.2.4;

[0134] Step 2.2.2: Starting from Λ0, select an unexpanded node and expand it into Λ';

[0135] Step 2.2.2.1: When Λ0 is in a non - terminal state and has been fully expanded, use the above UCB(v, c) formula to select the child node with the largest UCB value in Λ0 as the new Λ0;

[0136] Step 2.2.2.2: When Λ0 is in a non - terminal state and has not been fully expanded, then Λ0 expands to a child node Λ'. Select an unvisited action a from the possible action space of Λ0 and use it as the action of Λ', and update the depth of Λ';

[0137] Step 2.2.2.3: Use the state transition function to find the state S(Λ') and reward value r of Λ', store them together with the action a in the experience replay pool à, and update the network parameters regularly;

[0138] Step 2.2.3: Starting from the expanded node Λ', simulate until the terminal state;

[0139] Step 2.2.3.1: Λ' adopts the action a with the largest value in the policy network * = argmaxQ(S(Λ'), Α(Λ'); θ), use the state transition function to obtain the next state S'(Λ') and reward value r, and update the current state s(Λ') = s'(Λ');

[0140] Step 2.2.3.2: Store the sequence [s(Λ'), a * , r, s'(Λ')] in the experience replay pool à, and update the network parameters regularly;

[0141] Step 2.2.4: Reverse-update the values and visit counts of all nodes on the Λ′ path;

[0142] Step 2.3: Let node be the root node of the task offloading search tree When node does not satisfy the termination condition, obtain the optimal offloading strategy for each time slot;

[0143] Step 2.3.1: Update node to the child node corresponding to the maximum UCB(node,0);

[0144] Step 2.3.2: Obtain the time slot t = deep(node), and the optimal offloading strategy A * (t) = a(node);

[0145] Step 2.4: Calculate tasks according to the obtained optimal offloading strategy A * (t) for each time slot.

[0146] Step 3: Introduce the optimal privacy parameter that balances system performance and privacy protection requirements in task offloading. By dynamically adjusting the privacy parameter, both the efficiency of task offloading is satisfied and the security of user location privacy is guaranteed, achieving a balance between task offloading performance and privacy protection requirements. Based on this parameter, perturb the true location of the user to achieve three-dimensional geographical indistinguishability, ensuring that even if an attacker can obtain some information, they cannot accurately infer the exact location of the user, providing three-dimensional differential privacy guarantee for the location privacy of mobile users;

[0147] Furthermore, the specific operation of Step 3 is as follows:

[0148] Step 3.1: Conduct z interference experiments for different privacy coefficients ε, generate the perturbation radius R = Γ(3, 1 / ε), randomly select the unit vector U = (ξ, ψ), and the perturbed location of the user UE n The perturbed location

[0149] Step 3.2: Record the reward mean of the perturbed locations generated by ε in Table Ο;

[0150] Step 3.3: Obtain the optimal privacy coefficient ε according to the maximum reward mean in Table Ο * ;

[0151] Step 3.4: Use ε * to calculate the perturbed location of the mobile user;

[0152] The hardware platform used in this invention is a computer equipped with an Intel(R) Core(TM) i9-14900KF processor and an NVIDIA GeForce RTX 4090 GPU. Our software platform is based on Python 3.8. Each experiment is run 1000 times, and the average results are reported.

[0153] In the embodiments provided by this invention, the main parameters of the mobile users and the mobile edge network are shown in Table 1.

[0154] Table 1 Main simulation parameters

[0155]

[0156]

[0157] In the ablation experiment for comparison in this invention, the reward functions are set as shown in Table 2.

[0158] In this invention, reward functions related to the computing cost, privacy entropy, task discard rate, and MEC load are set respectively, that is, r1 = -C(t).

[0159] Table 2 Comparison of offloading performance with different reward functions

[0160]

[0161] After comparison, (3) has the best effect. In this invention, the reward function is set as r(t) = r1 + r2 + r3.

[0162] This invention uses the execution cost, task discard rate, edge node load variability, and privacy entropy as the metrics for task offloading.

[0163] This invention is compared with the following 3 offloading algorithms:

[0164] Local computing (LC): All users compute tasks locally without considering task offloading.

[0165] Edge node computing (EC): All users offload all tasks to the edge node for computing.

[0166] Deep Q-Network (DQN): Use the DQN algorithm to train the offloading problem to obtain the offloading decisions of all users.

[0167] Table 3 Comparison of offloading performance under different task sizes

[0168]

[0169]

[0170] Table 4 Comparison of Offloading Performance under Different Numbers of Users

[0171]

[0172]

[0173] Table 5 Comparison of Offloading Performance under Different Weights of Latency and Energy Consumption

[0174]

[0175] The experimental results of the method proposed above and the comparison with other task offloading methods are shown in Tables 3 - 5. The best performance results are shown in bold. The experimental results show that the present invention makes the weighted cost of task calculation latency and energy consumption lower, the task discard rate lower, the variability of the edge node load lower, and more secure privacy protection is achieved. In most cases, the offloading performance of each item is better than other methods.

[0176] The content not described in detail in the specification of the present invention belongs to the prior art well known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. An interpretable task offloading and location privacy protection method for edge computing, characterized in that, The method includes the following steps: Step 1: Build a dynamic simulation environment to simulate the real task offloading scenario in the mobile edge system, model the task offloading of the mobile edge system, comprehensively consider the balance between task offloading energy consumption and location privacy, and transform it into a Markov decision process; Step 2: Design an algorithm that combines Monte Carlo tree search and deep reinforcement learning for the task offloading problem of mobile edge nodes; Step 3: Introduce the optimal privacy parameter that balances system performance and privacy protection requirements in task offloading. By dynamically adjusting the privacy parameter, achieve the balance between task offloading performance and privacy protection requirements, and perturb the real location of the user based on this parameter.

2. The interpretable task offloading and location privacy protection method for edge computing according to claim 1, wherein The specific operation of Step 1 is as follows: Establish the task offloading scenario as a dynamic multi-user, multi-edge node and complex environment that supports partial offloading of users. Develop an offloading strategy that includes two stages, and use execution cost, task discard rate, edge node load variability, and privacy entropy as metrics for task offloading; Model the task offloading problem in mobile edge computing. The following is the formulation of the problem: 0≤λ≤1 (3) Among them, P0 is the objective of the present invention, which means minimizing the cost of task offloading, task discard rate, and edge node load variability in mobile edge computing, while maximizing the privacy entropy that measures the privacy protection ability; T represents the number of included time slots; N represents the number of mobile users; M represents the number of edge nodes; C(t) represents the execution cost at time slot t; v m (t) represents the load variability of the edge node; qr n (t) represents the user UE n 's task discard rate at time slot t, H n (t) represents the privacy entropy, which is used to quantify the privacy degree of the user UE n 's offloading strategy; π n,0 (t) represents the proportion of local execution at the user UE n ; Constraint (1) means that the proportion of any user executing tasks locally and on all edge nodes is within the range of 0 - 1; Constraint (2) means that the sum of the proportions of any user executing tasks locally and on all edge nodes is 1; Constraint (3) means that the weight for balancing energy consumption and computing latency is within the range of 0 - 1; Transform the task offloading problem of mobile edge nodes into a Markov decision process, specifically including: State space: S(t) = [W(t), F L (t), F M (t), Π(t)], where W(t) is the task feature set of the mobile user, f L (t) is the computing power set of the mobile user, F M (t) is the computing power set of the edge node, and Π(t) is the offloading vector set of the mobile user; Action space: A(t) = [a1(t), a2(t), …, a N (t)] T , where a n (t) = [β1, β2, …, β M , representing the decision vector of the user UE n on all edge nodes; State transition probability: Set the state transition probability of the task offloading environment to 1; The calculation formula of the reward function is: The calculation formula of privacy entropy is: Among them, represents the offloading preference of the user UE n for MEC m which is inferred based on the total offloading volume of the user UE n and the partial offloading volume for MEC m ; l n,m l(t) represents the amount of tasks executed by the user UE n offloaded to MEC m ; User UE n The calculation formula for the task discard rate in time slot t is as follows: Among them, represents the amount of tasks discarded by the user UE n in time slot t; τ represents the length of a time slot; r n,m (t) represents the user UE n to the MEC m data transmission rate; l n (t) represents the user UE n the amount of tasks executed locally, b n (t) represents the number of CPUs required to calculate each bit of the task, represents the user UE n local computing power; Load variability v of the edge node m (t) is measured by the standard deviation of the server load, and the formula is: Among them, represents the load of the edge node MEC m , represents the average load of all edge nodes.

3. The interpretable task offloading and location privacy protection method for edge computing according to claim 2, characterized in that The specific operation of Step 2 is as follows: Set the initial offloading state as the root node, and starting from the root node, establish a task offloading search tree through the four steps of selection, expansion, simulation, and backpropagation of Monte Carlo tree search Specifically include: Step 2.1: Model the task offloading problem and impose necessary constraints on the mobile edge system; Step 2.2: Create a task offloading search tree Step 2.2.1: Set the initial uninstalled state S0 as the root node Λ0, initialize the policy network Q with random weights θ, and initialize the target network and initialize the experience buffer Under a limited computing budget, loop through the following steps 2.2.2 - 2.2.4; Step 2.2.2: Starting from the root node Λ0, select an unexpanded node and expand it into Λ′; Step 2.2.2.1: When the root node Λ0 is non-terminal and has been fully expanded, use Select the child node with the largest UCB value in the root node Λ0 as the new root node Λ0, where v′ is the child node of v, Ι and V represent the number of visits and the value respectively, and c represents the exploration coefficient; Step 2.2.2.2: When the root node Λ0 is non-terminal and unexpanded, the root node Λ0 expands to generate a child node Λ′. Select an unvisited action a from the action space of the root node Λ0 and use it as the action of Λ′, and update the depth of Λ′; Step 2.2.2.3: Use the state transition function to find the state S(Λ′) and reward value r of Λ′, and store them together with the action a in the experience buffer Update the network parameters regularly; Step 2.2.3: Starting from the expanded node Λ′, simulate until the terminal state; Step 2.2.3.1: Λ' adopts the action a with the maximum value in the policy network * = argmaxQ(S(Λ'), Α(Λ'); θ), and uses the state transition function to obtain the next state S'(Λ') and the reward value r, and updates the current state s(Λ') = s'(Λ'); Step 2.2.3.2: Store the sequence [s(Λ′), a * , r, s′(Λ′)] in the experience buffer Update the network parameters regularly; Step 2.2.4: Update the value and visit count of all nodes on the path of Λ′ in reverse; Step 2.3: Let node be the root node of the task offloading search tree When node does not meet the termination condition, obtain the optimal offloading strategy for each time slot; Step 2.3.1: Update the node to the child node corresponding to the maximum UCB(node,0); Step 2.3.2: Obtain the time slot t = deep(node), and the optimal offloading strategy A * (t) = a(node); Step 2.4: Each time slot calculates tasks according to the obtained optimal offloading policy A * (t).

4. An interpretable task offloading and location privacy protection method for edge computing according to claim 3, characterized in that, The specific operation of Step 3 is as follows: Step 3.1: Conduct z interference experiments for different privacy parameters ε, generate a perturbation radius R = Γ(3, 1 / ε), randomly select a unit vector U = (ξ, ψ), and user UE n Perturbed location Step 3.2: Record the mean reward of the perturbed locations generated by different privacy parameters ε in table Ο; Step 3.3: Derive the optimal privacy parameter ε based on the maximum reward mean in Table Ο * ; Step 3.4: Use the optimal privacy parameter ε * Calculate the perturbed location of the mobile user.