Multi-User Multi-Access Point Based Task Offloading Privacy Protection System and Method

Through the task offloading of the privacy protection system with multiple users and multiple access points, and the reinforcement learning neural network training decision-maker is used to solve the privacy leakage problem caused by uninstalling preferences in mobile edge computing, achieving a balance between user privacy and experience quality.

CN115913712BActive Publication Date: 2025-07-22HUZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211431934.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-07-22
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

In mobile edge computing, the offload preference of multiple access points leads to user privacy leakage, and existing methods are difficult to balance user experience quality and privacy protection, especially in dynamic environments.

Method used

A task offloading privacy protection system based on multi-user and multi-access points is adopted, and a training module of a trusted third-party server and a feedback module of a mobile device is used to train a decision-maker through a reinforced learning neural network (such as DQN), which comprehensively considers user privacy, energy consumption, delay and task loss, and formulates the optimal offloading strategy.

Benefits of technology

Effectively protect users' real-time location privacy, taking into account the quality of user experience during the uninstallation process, including computing energy consumption and delay, and achieving a balance between user privacy and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115913712B_ABST
    Figure CN115913712B_ABST
Patent Text Reader

Abstract

The present invention discloses a task offloading privacy protection system and method based on multi-user and multi-access points, which is applied to a mobile edge computing network; it includes a training module arranged in a trusted third-party server and a feedback module arranged in a mobile device; the training module is used to train a decision maker according to the empirical data provided by the feedback module and provide the decision maker with converged training to the feedback module; the feedback module is used to adopt the decision maker provided by the training module in each time slot to determine the optimal offloading strategy for offloading tasks to multiple available edge nodes according to the current state it observes. The present invention considers the privacy leakage caused by users' offloading preferences in a multi-access point environment, formulates a privacy evaluation index based on information entropy, and comprehensively considers the privacy, energy consumption, delay, and task loss of users as optimization objectives, so as to achieve a balance between user privacy and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet of Things security, and more specifically, relates to a task offloading privacy protection system and method based on multiple users and multiple access points. Background Art

[0002] With the rapid development of Internet of Things technology and the popularity of mobile devices, portable mobile devices are embedded with technologies such as face recognition and augmented reality, and these technology applications have enriched the quality of user experience. However, due to the size limitation of mobile devices, their computing power and battery power are difficult to meet the growing computing demands. In mobile edge computing, operators transfer cloud computing centers with sufficient computing resources and storage capabilities to edge nodes closer to users. These edge nodes have good computing resources and can provide computing resources for mobile devices to reduce the computing latency and power consumption on the local mobile devices. Usually, there are multiple edge nodes available for offloading in a region, and there is also a problem of resource competition when multiple mobile devices offload to the same edge node. Therefore, it is very important to select a suitable offloading strategy to enable each user to obtain the optimal quality of user experience.

[0003] Currently, the offloading preferences of mobile devices for different edge nodes will expose the real-time location of users. Specifically, when mobile devices only focus on optimizing latency and energy consumption, since mobile devices tend to offload tasks to the nearest edge node for computing in order to reduce energy consumption and latency (the closer the distance, the better the corresponding channel gain), this offloading preference may lead to location leakage. If multiple edge nodes cooperate, the channel conditions between the user and each edge node can be inferred based on the amount of tasks offloaded from the same mobile device to each edge node, thereby obtaining the real-time location of the user.

[0004] Existing traditional privacy protection solutions, such as authentication, secure and private data storage and computing, intrusion detection, etc., are difficult to solve the above privacy problems exposed by offloading decisions. In addition, offloading decisions that overly pursue privacy protection will also lead to an increase in the computing latency and energy consumption of users, thereby affecting the quality of user experience. Therefore, the biggest challenge in protecting privacy in mobile edge computing is to find the best offloading strategy to balance the quality of user experience and privacy protection. Existing task offloading methods include Lyapunov optimization, linear programming, game theory, etc. However, most of these methods consider the instantaneous optimization of the environment and do not take into account the dynamic changes of the environment. Moreover, traditional methods are difficult to solve the curse of dimensionality problem and require prior knowledge, while the system state is difficult to describe using certain specific distributions. Summary of the Invention

[0005] The present invention aims to solve the problem of privacy leakage caused by offloading preferences in a multi-access point environment in existing mobile edge computing, and proposes a task offloading privacy protection system and method based on multi-user multi-access points.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided a task offloading privacy protection system based on multi-user multi-access points, which is applied to a mobile edge computing network; the mobile edge computing network includes: a plurality of edge nodes for providing mobile services nearby and accepting tasks for offloading; and mobile devices for performing multi-to-multi communication requests for task offloading services with the edge nodes.

[0007] It includes a training module provided in a trusted third-party server and a feedback module provided in a mobile device;

[0008] The training module is used to train a decision maker according to the experience data provided by the feedback module, and provide the decision maker with converged training to the feedback module;

[0009] The feedback module is used to adopt the decision maker provided by the training module in each time slot to make an optimal offloading strategy for offloading tasks to multiple available edge nodes according to the current state it observes, and after performing the corresponding actions of the optimal offloading strategy, evaluate the action reward and observe the state of the next time slot, forming local experience including the current state, action and reward and the state of the next time slot; and is used to provide the local experience for a period of time to the training module.

[0010] Preferably, in the task offloading privacy protection system based on multi-user multi-access points, the training module integrates the local experience provided by the feedback modules of multiple mobile devices into global experience as experience data.

[0011] Preferably, in the task offloading privacy protection system based on multi-user multi-access points, the decision maker is a reinforcement learning neural network based on Markov, preferably a model-free reinforcement learning structure, specifically a DQN reinforcement learning.

[0012] Preferably, in the task offloading privacy protection system based on multi-user multi-access points, the decision maker of mobile device l:

[0013] State at the current time slot t where is the task that mobile device l needs to offload at the current time slot t, is the location where mobile device l is located at the current time slot t;

[0014] Offloading decision at the current time slot t is that mobile device l transmits the task volume at power is The task to the edge node m, and the computing power required for local computing and the task volume Denoted as:

[0015]

[0016] The execution offloading decision in the current time slot t The obtained reward Is: the weighted sum of the quality of user experience and the privacy level, determined according to the principle that the higher the quality of user experience and the higher the privacy level, the greater the reward value; where the quality of user experience includes two aspects: computing delay and task loss volume, and the quality of user experience is determined according to the principle that the longer the computing delay and the more the task loss volume, the lower the quality of user experience.

[0017] Preferably, in the task offloading privacy protection system based on multi-user and multi-access point, the computing delay is the larger value of the local computing delay and the offloading delay, and the task loss volume is the size of the task that is lost because the calculation is not completed in one time slot; the privacy level is determined according to the entropy value of the offloading preference of the mobile device l for each edge node, according to the principle that the larger the value, the higher the privacy level.

[0018] Preferably, in the task offloading privacy protection system based on multi-user and multi-access point, the training module set in the trusted third-party server trains the decision maker according to the following method:

[0019] S1. Experience data collection: The training module collects the local experiences provided by multiple mobile device feedback modules and integrates them into global experience; the global experience is the set of local experiences provided by multiple mobile device feedback modules, including the global state in the current time slot t The global offloading decision in the current time slot t The global reward in the current time slot t

[0020] S2. Decision maker independent training: For each mobile device, independently use its offloading decision in the previous time slot t The offloading decisions of all other mobile devices Global state S t and the global state S in the next time slot t t+1 as samples, and use gradient update to update the decision maker with the goal of maximizing the objective function, and the objective function represents the action reward of this mobile device;

[0021] Preferably, the task offloading privacy protection system based on multi-user and multi-access point uses the deep Q-network (DQN) network Actor in reinforcement learning as the decision maker, with parameters as the objective function J(π of the reinforcement learning neural network with network parametersl ) is as follows:

[0022]

[0023] Update the Actor network π based on the gradient l parameters The gradient of the above objective function can be expressed as:

[0024]

[0025] Preferably, the task offloading privacy protection system based on multi-user and multi-access point updates in a soft update manner.

[0026] Preferably, for the task offloading privacy protection system based on multi-user and multi-access point, the state-action Q function of mobile device l is expressed as:

[0027]

[0028] Wherein, represents the expected value, S t represents the global observation, represents the action set of mobile devices other than mobile device l, and γ is the discount factor of the long-term reward;

[0029] Adopt the Critic neural network Q l to approximate the state-action Q function of mobile device l, and the network parameters corresponding to the neural network are Update the parameters by minimizing the loss function of this mobile device l Loss function is defined as:

[0030]

[0031] Wherein, represents taking the mathematical expectation for the samples in the experience pool, Preferably, use the Critic neural network Q' l to calculate the value of y l value.

[0032] According to another aspect of the present invention, a task offloading privacy protection method based on multi-user and multi-access point is provided. Applying the task offloading privacy protection system provided by the present invention, it includes the following steps:

[0033] Set up a training module on the trusted third-party server to create or train a decision maker for all mobile devices; the mobile devices download the decision maker from the third-party server;

[0034] For any mobile device l, when the mobile device l needs to offload tasks in time slot t, the following steps are executed:

[0035] (1) Detect the current location of the mobile device and the tasks that need to be offloaded in the current time slot Obtain the state of the current time slot t Input to the decision maker to obtain the offloading decision The offloading decision packet moves the device l with power The amount of data transmitted for the task is Transmit the task of to the edge node m, as well as the computing power required for local computing and the amount of data

[0036] (2) The mobile device l makes task offloading according to the offloading decision obtained in step (2) And observe the state of the next time slot And evaluate the execution of the offloading decision in the current time slot t The obtained reward Construct a local experience data of the mobile device l

[0037] After a preset period of time, the feedback modules of multiple mobile devices collect local experiences and submit them to the training module of the trusted third-party server. The training module integrates the local experiences of multiple mobile devices into global experiences and updates the decision maker for the multiple mobile devices accordingly.

[0038] Generally speaking, compared with the prior art, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0039] 1. The present invention considers the privacy leakage caused by users' offloading preferences in a multi-access point environment, formulates a privacy evaluation index based on information entropy, and comprehensively considers users' privacy, energy consumption, delay, and task loss as optimization objectives, so as to achieve a balance between user privacy and user experience.

[0040] 2. The present invention uses multi-agent deep reinforcement learning to learn strategies. Compared with traditional single-agent reinforcement learning, it considers the game between multiple users and the environmental changes caused by the changes in the offloading strategies of other users, and establishes a centralized training and distributed execution architecture with a trusted third party.

[0041] 3. The present invention considers the mobile edge computing environment of multiple users and multiple access points, while most current studies focus on the single-access point environment; in addition, it also considers the impact of user mobility on offloading decisions, and establishes a strategy for intelligent offloading based on the amount of data and the physical location of users. Description of the Drawings

[0042] Figure 1 is a schematic diagram of the scenario of the embodiment of the present invention;

[0043] Figure 2 is a structural diagram of the task offloading privacy protection system based on multi-user and multi-access points provided by the embodiment. Detailed implementation manners

[0044] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0045] The task offloading privacy protection system based on multi-user and multi-access points provided by the present invention is applied to a mobile edge computing network; the mobile edge computing network includes: a plurality of edge nodes for providing mobile services nearby and receiving tasks for offloading; and mobile devices for making many-to-many communication requests for task offloading services with the edge nodes.

[0046] The task offloading privacy protection system provided by the present invention includes a training module provided in a trusted third-party server and a feedback module provided in a mobile device;

[0047] The training module is used to train a decision maker according to the experience data provided by the feedback module and provide the decision maker with converged training to the feedback module; in a preferred solution, the training module integrates the local experiences provided by the feedback modules of multiple mobile devices into global experience as experience data.

[0048] The feedback module is used to adopt the decision maker provided by the training module in each time slot to decide the optimal offloading strategy for offloading tasks to multiple available edge nodes according to the currently observed state, and after performing the corresponding actions of the optimal offloading strategy, evaluate the action reward and observe the state of the next time slot to form local experience including the current state, action, reward, and the state of the next time slot; and is used to provide the local experience for a period of time to the training module.

[0049] The decision maker is a Markov-based reinforcement learning neural network, preferably a model-free reinforcement learning structure, specifically a DQN reinforcement learning; the decision maker of mobile device l:

[0050] The state at the current time slot t wherein is the task that mobile device l needs to offload at the current time slot t, is the location where mobile device l is located at the current time slot t;

[0051] The offloading decision for the current time slot t Transmit the task volume of for mobile device l with power to edge node m, and the computing power required for local computing and the task volume are denoted as:

[0052]

[0053] Execute the offloading decision for the current time slot t The obtained reward is: the weighted sum of the quality of user experience and the privacy level, determined according to the principle that the higher the quality of user experience and the higher the privacy level, the greater the reward value; where the quality of user experience includes two aspects: computing delay and task loss volume, and the quality of user experience is determined according to the principle that the longer the computing delay and the more the task loss volume, the lower the quality of user experience. The computing delay is the larger value of the local computing delay and the offloading delay, and the task loss volume is the size of the task that is not completed and lost in a time slot; the privacy level is determined according to the entropy value of the offloading preference of mobile device l for each edge node, according to the principle that the larger the value, the higher the privacy level.

[0054] In a preferred solution, the training module set in the trusted third-party server trains the decision maker according to the following method:

[0055] S1. Experience data collection: The training module collects the local experiences provided by multiple mobile device feedback modules and integrates them into global experiences; the global experiences are the set of local experiences provided by multiple mobile device feedback modules, including the global state of the current time slot t The global offloading decision for the current time slot t The global reward for the current time slot t

[0056] S2. Independent training of the decision maker: For each mobile device, independently use its offloading decision in the previous time slot t The offloading decisions of all other mobile devices The global state S t The global state S of the next time slot t t+1 as samples, and use gradient update to update the decision maker with the goal of maximizing the objective function, where the objective function represents the action reward of this mobile device; in a preferred solution, use the deep Q-network (DQN) network Actor in reinforcement learning as the decision maker, and use the parameter The objective function J(π l ) of the reinforcement learning neural network with network parameters is:

[0057]

[0058] Update the Actor network π based on the gradient l parameters The gradient of the above objective function can be expressed as:

[0059]

[0060] Preferably, soft update is used for the update.

[0061] In a preferred solution, a DQN network is used as the decision maker; the state-action Q function of the mobile device l is expressed as:

[0062]

[0063] wherein, represents the expected value, S t represents the global observation, represents the action set of mobile devices other than the mobile device l, and γ is the discount factor of the long-term reward.

[0064] In a preferred solution, a Critic neural network Q l is used to approximate the state-action Q function of the mobile device l, and the network parameters corresponding to the neural network are The parameters are updated by minimizing the loss function of the mobile device l Loss function is defined as:

[0065]

[0066] wherein, represents the mathematical expectation when taking for the samples in the experience pool, Preferably, a Critic neural network Q' l is used to calculate the value of y l value.

[0067] The task offloading privacy protection method provided by the present invention includes the following steps:

[0068] Set up a training module on a trusted third-party server to create or train a decision maker for all mobile devices; the mobile devices download the decision maker from the third-party server;

[0069] For any mobile device l, when the mobile device l needs to perform task offloading at time slot t, the following steps are executed:

[0070] (1) Detect the current location of the mobile device and the task to be offloaded in the current time slot Obtain the state at the current time slot t Input decision maker to obtain offloading decisions The offloading decision packet moves device l with power The transmission task volume is Tasks of to edge node m, as well as the computing power required for local computing And task volume

[0071] (2) Mobile device l makes task offloading according to the offloading decision obtained in step (2) And observe the state of the next time slot And evaluate the execution of the offloading decision in the current time slot t The obtained reward Construct a local experience data of mobile device l

[0072] After a preset time period, the feedback modules of multiple mobile devices collect local experiences and submit them to the training module of the trusted third-party server. The training module integrates the local experiences of multiple mobile devices into global experiences and updates the decision maker for the multiple mobile devices accordingly.

[0073] This patent designs a task offloading strategy considering privacy protection and resource allocation for a multi-node edge computing environment, which not only effectively protects the real-time location of users, but also takes into account the quality of user experience during the offloading process, including computing energy consumption, computing latency, and task loss volume. By comprehensively considering user privacy and experience, a balance is achieved between the two.

[0074] The following is an embodiment:

[0075] As Figure 1 shown, the mobile edge computing scenario of this embodiment is an Internet of Things with three layers of nodes. The first layer is the cloud computing center, which migrates some services to the edge nodes so that it can serve mobile users nearby; the second layer is the edge nodes, which can accept the offloading tasks of users to reduce the energy consumption and computing latency of users; the third layer is the mobile terminals. These mobile devices move with the users continuously, and the channel state will also change continuously. Therefore, a fixed offloading strategy cannot be adopted. It should be noted that in this offloading scenario, a cell with multiple access points is considered. Users can offload computing tasks to multiple edge nodes, and the edge nodes (access points) can be defined as {M1, M2, M3, …, M m}, and mobile users can be defined as {N1, N2, N3, …, N n}

[0076] During the task offloading process, the interaction among mobile users, edge nodes, and the trusted third party is as Figure 2As shown. Since the task offloading experience data of each user contains the user's location information, and the offloading decision also reveals the user's location privacy, the common agent information sharing in multi-agent reinforcement learning cannot be applied to this scenario. Therefore, we consider setting up a trusted third party to centrally train the decision-making device of each mobile device, so as to achieve task offloading privacy protection.

[0077] The task offloading privacy protection system provided in this embodiment includes a training module set in a trusted third-party server and a feedback module set in a mobile device.

[0078] The training module is used to train the decision-making device according to the experience data provided by the feedback module and provide the decision-making device with converged training to the feedback module. Specifically, the training module integrates the local experiences provided by the feedback modules of multiple mobile devices into global experience as experience data.

[0079] The feedback module is used to adopt the decision-making device provided by the training module in each time slot to decide the optimal offloading strategy for offloading tasks to multiple available edge nodes according to the currently observed state, and after performing the corresponding actions of the optimal offloading strategy, evaluate the action reward and observe the state of the next time slot to form local experience including the current state, action and reward, and the state of the next time slot; and is used to provide the local experience of a period of time to the training module.

[0080] The decision-making device adopts a DQN reinforcement learning neural network; the decision-making device of mobile device l:

[0081] The state at the current time slot t where is the task that mobile device l needs to offload at the current time slot t, is the location where mobile device; is located at the current time slot t;

[0082] The offloading decision at the current time slot t is that mobile device l transmits the task volume of at power to edge node m, and the computing power required for local computing and the task volume are denoted as:

[0083]

[0084] The execution offloading decision at the current time slot t The obtained reward It is the weighted sum of the quality of user experience and the privacy level, determined according to the principle that the higher the quality of user experience and the higher the privacy level, the greater the reward value; where the quality of user experience includes two aspects: computing latency and task loss, and the quality of user experience is determined according to the principle that the longer the computing latency and the more the task loss, the lower the quality of user experience. The computing latency is the larger value of the local computing latency and the offloading latency, and the task loss is the size of the task that is not completed and lost in one time slot; the privacy level is determined according to the entropy value of the offloading preference of the mobile device l for each edge node, following the principle that the larger the value, the higher the privacy level.

[0085] Specifically in this embodiment, the reward in time slot t is calculated as follows:

[0086] Obtaining the computing latency:

[0087] 1. Calculate the CPU frequency of the mobile device according to the local computing power and the factor k determined by the mobile device chip structure:

[0088]

[0089] 2. The local computing delay of the mobile device l can be expressed as L represents the number of CPU computing cycles required for 1 bit of data; is the local computing task volume.

[0090] 3. The computing energy consumption of the mobile device l is calculated as follows:

[0091]

[0092] 4. The system uses code division multiple access and considers the interference caused by other users offloading to the same edge node. The signal-to-noise ratio between the mobile device l and the edge node v is as follows, σ 2 is the channel noise.

[0093]

[0094] where represents the channel gain between the mobile device l and the edge node v.

[0095] 5. The channel gain v , y v ) between the mobile device l and the edge node v with coordinates (x is as follows. g0 represents the reference channel gain at a distance of 1 meter from the edge node v, and the mobile device coordinates are represented

[0096]

[0097] 6. Calculate the transmission rate r between the mobile device l and the edge node v according to the channel gain and the bandwidth B l,v :

[0098]

[0099] 7. Calculate the transmission delay of the mobile device l offloaded to the edge node v according to the transmission rate r l,v and the offloading amount

[0100]

[0101] 8. Calculate the transmission energy consumption of the mobile device l according to the calculation delay of the mobile device l offloaded to each edge node

[0102]

[0103] 9. The edge node evenly distributes the computing resources according to the offloading amount of the mobile device, and according to the computing frequency of the edge node The delay for the edge node v to complete the computing task can be expressed as:

[0104]

[0105] 10. The energy consumption of the user side includes the energy consumption generated by local computing and the energy consumption generated by transmitting the offloading task

[0106]

[0107] 11. The computing delay is the larger value of the local computing delay and the offloading delay where the offloading delay is

[0108]

[0109] Considering that edge computing can transmit at high power and the calculation result is small, the delay of the returned result is ignored.

[0110] Therefore, the total computing delay can be expressed as

[0111] Obtaining the task loss amount:

[0112] The task loss amount This is because the system requires tasks to be computed within one time slot, and tasks that are not completed will be lost. The amount of lost tasks can be expressed as:

[0113]

[0114] where ζ represents the length of one time slot, and the custom function f(·) represents:

[0115]

[0116] Privacy level acquisition:

[0117] Based on the offloading amount, infer the offloading preference of mobile device l for each edge node, and thus evaluate the overall privacy level. The specific process is as follows:

[0118] 1. Calculate the total amount of tasks offloaded to the edge node according to the offloading decision:

[0119]

[0120] 2. Infer the offloading preference of mobile device l for edge node v from the offloading amount

[0121]

[0122] 3. Calculate the privacy entropy of mobile device l at time slot t according to the offloading preference of each edge node

[0123]

[0124] Calculate the reward

[0125] The weighted sum of the quality of user experience and the privacy level is used as the reward function, that is

[0126]

[0127] where ω i , i ∈ {1, 2, 3, 4} belongs to the weight factor.

[0128] Each piece of experience will be stored locally on mobile device l, and mobile device l uploads its local experience to the trusted third-party server at regular intervals.

[0129] The training module set in the trusted third-party server adopts the following method to train the decision maker:

[0130] ​​S1. Experience data collection: The training module collects the local experiences provided by multiple mobile device feedback modules and integrates them into global experience; the global experience is the set of local experiences provided by multiple mobile device feedback modules, including the global state at the current time slot t. The global offloading decision at the current time slot t The global reward at the current time slot t

[0131] S2. Independent training of the decision maker: For each mobile device, the offloading decision at its previous time slot t The offloading decisions of all other mobile devices The global state S t and the global state S at the next time slot t t+1 are used as samples, and gradient update is adopted to update the decision maker with the goal of maximizing the objective function, where the objective function represents the action reward of this mobile device; in this embodiment, the deep Q-network (DQN) network Actor in reinforcement learning is used as the decision maker, and with the parameter as the network parameter, the objective function J(π l ) of the reinforcement learning neural network is:

[0132]

[0133] Based on the gradient, update the parameters of the Actor network l l The gradient of the above objective function can be expressed as:

[0134]

[0135] In this embodiment, soft update is adopted for the update, specifically as follows:

[0136] The Actor network also adopts an online network π l and a target network π′ l . For smoother update, both the Actor and Critic networks adopt the soft update method, and the specific update is as follows: where δ is the soft update parameter.

[0137] Continuously update each network until the decision network of each agent converges, and then the mobile device downloads the latest Actor network to the local of the mobile device. After that, the offloading strategy can be calculated locally.

[0138] In this embodiment, the DQN network is used as the decision maker; the state-action Q function of mobile device l is expressed as:

[0139]

[0140] ​Among them, represents the expected value, S t represents the global observation, represents the set of actions of mobile devices other than mobile device l, and γ is the discount factor of the long-term reward.

[0141] The Critic neural network Q l is used to approximate the state-action Q function of mobile device l, and the network parameters corresponding to the neural network are The parameters are updated by minimizing the loss function of this mobile device l The loss function is defined as:

[0142]

[0143] Among them, represents taking for the samples in the experience pool, the mathematical expectation, In this embodiment, the Critic neural network Q′ l is used to calculate y l In the above formula and y l in need to be updated simultaneously. Therefore, to avoid algorithm divergence, two Critic neural networks are set separately. The online neural network Q l is used to calculate The target neural network Q′ l is used to calculate the value of y l .

[0144] The task offloading privacy protection method provided by this embodiment includes the following steps:

[0145] Set up a training module on the trusted third-party server to create or train a decision maker for all mobile devices; the mobile devices download the decision maker from the third-party server;

[0146] For any mobile device l, when the mobile device l needs to perform task offloading at time slot t, the following steps are executed:

[0147] (1) Detect the current location of this mobile device and the task to be offloaded in the current time slot Obtain the state at the current time slot t Input the decision maker to obtain the offloading decision The offloading decision packet moves mobile device l to transmit a task with a task volume of at power to the edge node m, as well as the computing power required for local computing and the task volume

[0148] (2) The mobile device l makes a task offloading decision obtained according to step (2). Perform task offloading and observe the state of the next time slot. And evaluate the execution of the offloading decision for the current time slot t. The obtained reward Construct a local experience data of the mobile device l.

[0149] After a preset time period, the feedback modules of multiple mobile devices collect local experiences and submit them to the training module of the trusted third-party server. The training module integrates the local experiences of multiple mobile devices into global experiences and updates the decision maker for the multiple mobile devices accordingly.

[0150] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A task offloading privacy protection system based on multi - user and multi - access point, characterized in that, Applied to a mobile edge computing network; The mobile edge computing network includes: a plurality of edge nodes for providing mobile services in the vicinity and accepting task offloading; and mobile devices that perform many-to-many communication requests for task offloading services with the edge nodes; It includes a training module provided in a trusted third-party server and a feedback module provided in a mobile device; The training module is used to train the decision maker according to the experience data provided by the feedback module and provide the decision maker with converged training to the feedback module; the training module integrates the local experiences provided by the feedback modules of multiple mobile devices into global experience as experience data; the decision maker is a reinforcement learning neural network based on Markov; Mobile device Decision maker of: Current time slot status of where is the current time slot task to be offloaded is the current time slot location where Current time slot offloading decision for a mobile device to transmit a task volume of at power to an edge node , and the computing power and task volume required for local computing are denoted as: Denoted as: ; Current time slot for performing offloading decision The obtained reward is: the weighted sum of the quality of user experience and the privacy level, determined according to the principle that the higher the quality of user experience and the higher the privacy level, the greater the reward value; where the quality of user experience includes two aspects: computing latency and task loss, and the quality of user experience is determined according to the principle that the longer the computing latency and the more the task loss, the lower the quality of user experience; The calculated delay is the larger value of the local calculation delay and the offloading delay, and the task loss amount is the size of the tasks that are lost because the calculation is not completed in a time slot; the privacy level is determined according to the mobile device For the entropy value of the offloading preference of each edge node, it is determined according to the principle that the larger the value, the higher the privacy level; The feedback module is used to adopt the decision maker provided by the training module in each time slot to decide the optimal offloading strategy for offloading tasks to multiple available edge nodes according to the currently observed state, and after performing the corresponding actions of the optimal offloading strategy, evaluate the action reward and observe the state of the next time slot to form local experience including the current state, action, reward, and the state of the next time slot; and is used to provide the local experience for a period of time to the training module.

2. The task offloading privacy protection system based on multi-user and multi-access point according to claim 1, characterized in that, The decision maker is a model-free reinforcement learning structure.

3. The task offloading privacy protection system based on multi-user and multi-access point according to claim 2, characterized in that, The decision maker is a DQN reinforcement learning structure.

4. The task offloading privacy protection system based on multi-user and multi-access points according to claim 1, characterized in that, The training module provided in the trusted third-party server adopts the following method to train the decision maker: S1. Empirical data collection: The training module collects local experiences provided by multiple mobile device feedback modules and integrates them into global experiences; the global experiences are a set of local experiences provided by multiple mobile device feedback modules, including the global state of the current time slot the global state of the current time slot the global offloading decision of the current time slot the global reward ; S2, decision maker independent training: for each mobile device, independently use its previous time slot Uninstall decision , uninstall decisions for all other mobile devices , Global State , next time slot The global state of For samples, gradient update is adopted to update the decision maker with the goal of maximizing the objective function, where the objective function represents the action reward of the mobile device.

5. The task offloading privacy protection system based on multi-user and multi-access point according to claim 4, characterized in that, Using the reinforcement learning neural network DQN network Actor as the decision maker, with parameters as the network parameters of the reinforcement learning neural network objective function is as follows: ; Update the parameters of the Actor network based on the gradient of the parameter , the gradient of the above objective function can be expressed as: 。 6. The task offloading privacy protection system based on multi-user and multi-access point according to claim 5, wherein Update in a soft update manner.

7. The task offloading privacy protection system based on multi-user and multi-access points according to claim 5, characterized in that Mobile device Status behavior The function is expressed as: ; Among them, represents the expected value, represents the global observation, represents the set of actions of mobile devices other than the mobile device, is the discount factor of the long-term reward; Adopt a Critic neural network to approximate the state behavior of a mobile device The network parameters corresponding to the neural network are , and the parameters are updated by minimizing the loss function of this mobile device The loss function is defined as: ; Among them, represents taking the mathematical expectation of the samples in the experience pool 。 8. The task offloading privacy protection system based on multi-user and multi-access point according to claim 7, characterized in that, Use the Critic neural network to calculate the value of 9. A task offloading privacy protection method based on multi-user and multi-access points, applied to the task offloading privacy protection system according to any one of claims 1 to 8, including the following steps: The training module provided in the trusted third-party server creates or trains a decision maker for all mobile devices; The mobile device downloads the decision maker from the third-party server; For any mobile device , in time slot when the mobile device needs to offload tasks, perform the following steps: (1)Detect the current location of the mobile device , and the tasks to be offloaded in the current time slot , obtain the state of the current time slot , input it to the decision maker, and obtain the offloading decision , the offloading decision includes that the mobile device transmits tasks with a task volume of at a power to the edge node , as well as the computing power and task volume required for local computing ;​​ (2) Mobile device The offloading decision obtained in step (2) Perform task offloading and observe the state of the next time slot And evaluate the current time slot The executed offloading decision The obtained reward to construct a local experience data of the mobile device ; ; After a preset time period, the feedback modules of multiple mobile devices collect local experiences and submit them to the training module of the trusted third-party server, and the training module integrates the local experiences of multiple mobile devices into global experience and updates the decision maker for the multiple mobile devices accordingly.

Citation Information

Patent Citations

  • Task unloading optimization method for mobile edge computing user privacy protection

    CN114528081A

  • Method for offloading computing task of mobile user

    WO2022121097A1